From Self-Distillation to Self-Practice: Privileged Information for Multi-Turn Agents
arXiv cs.AIen
arXiv:2609.29051v1 Announce Type: new Abstract: On-policy self-distillation (OPSD) has become a popular recipe for post-training LLM agents. It supervises the agent model at the token level with a stronger teacher view of the same model, obtained by conditioning on privileged information (PI). In this work, we show that in multi-turn agents, this paradigm teaches the student to act with confidence but without the information behind it. The trained agent behaves as if it had privileged information it never observed, and its performance falls well short of plain RL, in the worst case below the untrained base model. Therefore, we propose Privileged Self-Practice (PSP), which keeps the PI and mo
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
- Reglering
Related AI news
- "One small step for TPUs": Google CEO Sundar Pichai announces Project Suncatcher to test AI compute in SpaceEconomic Times Tech · September 25, 2026
- PAWS: Policy-driven Agentic World SimulationarXiv cs.AI · September 25, 2026
- Functional Architecture of European Electricity Trading Markets: Requirements for AI Supported Trading Systems under Regulatory ConstraintsarXiv cs.AI · September 25, 2026
- When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability RoutingarXiv cs.AI · September 25, 2026
- TW3Cast: A Frozen Router of Lightly Fine-Tuned Foundation Models for Time-Series Forecasting on GIFT-Eval, Selected Entirely on the Training SplitarXiv cs.AI · September 25, 2026
- BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data PipelinesarXiv cs.AI · September 25, 2026