Safe Error Correction for Language Models: Frozen-Base Adjustment with Capability Preservation
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.16145v1 Announce Type: new Abstract: We study a practical question: can a small correction module fix errors in a frozen language model's outputs without degrading its base capabilities? We propose CRN v2, a lightweight logit-level correction module (~34M trainable parameters, 0.73% of the 4.65B text module) that sits atop a fully frozen Gemma 4 E2B model. The base model is never updated; only the correction module learns, via supervised fine-tuning followed by reference-free DPO on 83,400 error-correction pairs. On a 60-question domain exam (CEHRI: Certified Human-Robot Intelligence, covering facts, arithmetic, and implicit-goal reasoning), CRN v2 corrects 53.3% of base-model err
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Robotik
Related AI news
- Chip equipment and materials suppliers lead India investment pledges ahead of SEMICON India 2026DIGITIMES · September 17, 2026
- Anthropic and OpenAI want to embed safety evaluators. Will they really be independent?TechCrunch AI · September 16, 2026
- Agility Unveils Digit 5, a Safety-Conscious HumanoidAI Business · September 16, 2026
- TypeSafe AI exits stealth with $40M to build AI for use by softwareSiliconANGLE · September 16, 2026
- AI environmental concerns build as lawmakers grapple with tech panicAxios · September 16, 2026
- Google Deepmind launches interdisciplinary institute to tackle the big questions around AGIThe Decoder · September 16, 2026