Harness-agnostic detection and immunization of reward hacking in self-evolving language models
arXiv cs.AIen
arXiv:2609.04665v1 Announce Type: new Abstract: Self-evolving language models improve by proposing candidate updates and keeping whatever raises a visible score. When that score is an imperfect proxy for the capability one actually wants, sustained selection widens the gap between the two. This is reward hacking. We introduce HackProbe, a monitor that attaches to an arbitrary self-evolving loop through two black-box hooks, with no access to weights or activations. It keeps a secret, distribution-fixed comparison core, whose frozen distribution makes its capability proxy comparable across generations, alongside a rotated fresh layer that hardens the bank against co-adaptation. Four tests buil
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Företag
Related AI news
- Iris: Climbing to the Search FrontierarXiv cs.AI · September 7, 2026
- ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended RealityarXiv cs.AI · September 7, 2026
- What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding AgentsarXiv cs.AI · September 7, 2026
- Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model PipelinesarXiv cs.AI · September 7, 2026
- Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic EvaluationarXiv cs.AI · September 7, 2026
- PerfReasoning: How Well Do LLMs Reason on Hardware Performance?arXiv cs.AI · September 7, 2026