Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial Training
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.35890v1 Announce Type: new Abstract: Sparse-autoencoder decomposability and concentrated feature attribution are increasingly treated as evidence that a model's computation is easier to reverse-engineer. Whether this representational and attributional cleanliness actually predicts a smaller or more tractable causal circuit remains an open question. We test this directly using adversarial training as a controlled instrument: it reliably reshapes internal representations, but this alone does not constitute a test of circuit size. We investigate this question through reverse-engineering complexity: the causal structure required to recover a model's behavior at a fixed level of faithf
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Chinese firms trail global peers on profits, but AI power boom offers bright spot: NatixisSCMP Tech · September 30, 2026
- The Layer Mystery of VLA: An Information-Theoretical Analysis of VLA Latent InterfacearXiv cs.AI · September 30, 2026
- An Exact Generate - Transform Decomposition of Small-LLM Team Scaling Across Orchestration ArchitecturesarXiv cs.AI · September 30, 2026
- Memory Is a Derivation: The Distributed-Evidence Paradox in Long-Term AgentsarXiv cs.AI · September 30, 2026
- More Features Are Not More Evidence: Limits of Training-Free Human Activity Recognition with JevarXiv cs.AI · September 30, 2026
- An Empirical Study and Assessment of EU AI Act Compliance CheckersarXiv cs.AI · September 30, 2026