Open-Weight Masked Introspection: Measuring What Language Models Can Report About Their Own Computation
arXiv cs.AIen
arXiv:2608.20569v1 Announce Type: new Abstract: Are frontier models able to introspect about their internal states? Recent work suggests that under certain conditions a complex enough model can audit its own internals, call out what changed, and report back confidently about it. We tested that claim on eight open-weight models from seven families and found no such ability: asked whether their own computation had been altered, none answered better than chance. To test it we built Open-Weight Masked Introspection (OWMI), a framework that intervenes on residual-stream sites, attention heads and sparse-autoencoder features, then interrogates the model about the change against the null conditions
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- ChatGPT får nye funktioner og ændringer: En af dem er gigantisk bagdør til din iPhoneIngeniøren · August 24, 2026
- Source: AI researcher Luke Metz, who returned to OpenAI from TML earlier this year, joins Meta's Superintelligence Labs and will report to Alexandr Wang (Ina Fried/Axios)Techmeme · August 24, 2026
- Analysis: SK Hynix pushes beyond HBM with HBF and CPODIGITIMES · August 24, 2026
- Terminal Agents: A Survey of AI Agents in Command-Line EnvironmentsarXiv cs.AI · August 24, 2026
- Difficulty-Aware Semantic-ID Optimization for Generative RecommendationarXiv cs.AI · August 24, 2026
- Environmental Slow AI: Design Principles for Generative SystemsarXiv cs.AI · August 24, 2026