Extremely Sparse Supervision Incentivizes Reasoning Ability
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.04565v1 Announce Type: new Abstract: Large language models demonstrate increasingly strong reasoning capabilities through effective post-training. Yet, prevailing post-training methods optimize over massive numbers of tokens, implicitly assuming that effective learning must be token-intensive. We revisit this assumption in the on-policy distillation (OPD) setting, which naturally admits dense teacher supervision at every generated token. Using the Qwen3 family, we discover a counter-intuitive phenomenon: reasoning can be effectively incentivized by an extremely small fraction of generated tokens--as few as one or two tokens per reasoning trajectory, corresponding to only 0.05% of
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Reglering
Related AI news
- Iris: Climbing to the Search FrontierarXiv cs.AI · September 7, 2026
- ResLearn-XR: Residual Learning for Network Traffic and Quality-of-Experience-Aware Modeling in Extended RealityarXiv cs.AI · September 7, 2026
- What Does Multi-Harness RL Learn? Credit Assignment and Portability in Coding AgentsarXiv cs.AI · September 7, 2026
- Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model PipelinesarXiv cs.AI · September 7, 2026
- Harbor Adapters and Harbor-Index: Infrastructure and a Curated Meta-Dataset for Large-Scale Agentic EvaluationarXiv cs.AI · September 7, 2026
- PerfReasoning: How Well Do LLMs Reason on Hardware Performance?arXiv cs.AI · September 7, 2026