GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation
arXiv cs.AIen
arXiv:2609.21432v1 Announce Type: new Abstract: Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Despite recent advances in post-training methods, such as Group Relative Policy Optimization (GRPO), their practical deployment remains impeded by training instability arising from the reliance on importance sampling. We introduce Group Variance Policy Optimization (GVPO), a novel post-training method that integrates the analytical solution of KL-constrained reward maximization into its gradient weighting scheme. This formulation provides an intuitive interpretation: GVPO's gradient corresponds to the mean squa
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Reglering
Related AI news
- A researcher used GPT-6 Astra to decipher a WWI German radio transmission from 1918, one of the 50 famous unsolved ciphers listed on a German science blog (prinz)Techmeme · September 21, 2026
- DENSE: Distilling Agent Trajectories into Evidence-Grounded Shortcut Trees for Self-RefinementarXiv cs.AI · September 21, 2026
- Driving on Registers, Reasoning on Risk: Risk-Aware Occupancy for Register-Based End-to-End Autonomous DrivingarXiv cs.AI · September 21, 2026
- RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language ModelsarXiv cs.AI · September 21, 2026
- Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language ModelsarXiv cs.AI · September 21, 2026
- Ability-Residual Decoupled Modeling for Affective Cognitive DiagnosisarXiv cs.AI · September 21, 2026