The Off-Support Barrier: Why Semantic Safety Constraints Are Not Learning-Problem Invariants, and What Follows for Prior Design, Containment, and Verification
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2608.11243v1 Announce Type: new Abstract: We argue that a single structural fact organizes a wide range of phenomena in contemporary AI safety: a semantic safety constraint (e.g., the agent does not escape its sandbox) is an off-support object. Formally, if q is the data distribution and \(p(\cdot\mid w)\) the model, the safety predicate B is not measurable with respect to \(\sigma(\text{model}, q)\), whereas the real log-canonical threshold (RLCT) of singular learning theory (SLT) is. From this non-invariance we derive, as corollaries rather than independent observations: (i) why reward hacking and sandbox escape arise under outcome-based optimization; (ii) why encoding such constrain
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
Related AI news
- Rogue AI aren’t science fiction anymoreThe Verge AI · August 16, 2026
- When AI models aren't allowed to reflect on themselves, it changes their entire worldviewThe Decoder · August 16, 2026
- I gave Tencent’s WeChat AI agent control for 24 hours: where it excelled – and stumbledSCMP Tech · August 16, 2026
- Optima tackles AI benchmarking's biggest flaw by letting users test models against their own dataThe Decoder · August 16, 2026
- Pathway, which is developing AI models based on what it calls its "Post-Transformer" BDH architecture, raised a $30M seed at a $500M valuation (Antoine Tardif/Unite.AI)Techmeme · August 16, 2026
- Faire ressusciter les défunts grâce à l’intelligence artificielleIntelligence artificielle (FR) · August 16, 2026