OBC-Prune: Outcome-Based Calibration for Large Reasoning Model Pruning
arXiv cs.AIen
arXiv:2609.17890v1 Announce Type: new Abstract: Large reasoning models (LRMs) generate long chain-of-thought traces before answering, creating significant inference overhead. Pruning can reduce this cost, but its effectiveness depends on the calibration data used to estimate parameter importance. Recent work calibrates on the model's own rollouts instead of generic dataset, but treats all reasoning tokens uniformly, regardless of whether they contribute to successful reasoning. As a result, pruning protects weights by statistical salience rather than by their contribution to correct reasoning, so weights behind erroneous computation survive as readily as those behind correct computation. The
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Sources: Emulate, a month-old UK AI startup founded by former Google DeepMind researchers, is in advanced talks to raise as much as $700M at a $3.7B valuation (Financial Times)Techmeme · September 17, 2026
- The Other Half of the Memory Wall: Serving 35B MoEs from SSD with Trained Routing PredictionarXiv cs.AI · September 17, 2026
- EvolveTrade: Experience-Driven Policy Refinement for Self-Evolving LLM Trading AgentsarXiv cs.AI · September 17, 2026
- SNOMED CT Concept Recommendation from Masked Clinical ContextarXiv cs.AI · September 17, 2026
- Memory Has Geometry: Non-Uniform Geometric Memory for Long-Horizon Personalized AIarXiv cs.AI · September 17, 2026
- When to Call an LLM: A Confidence-Gated Hybrid for Cost-Effective Emotion Recognition in Conversational AIarXiv cs.AI · September 17, 2026