GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation

arXiv cs.AIen

GVPO++: Group Variance Policy Optimization for LLM Post-Training and On-Policy Distillation

arXiv:2609.21432v1 Announce Type: new Abstract: Post-training plays a pivotal role in enhancing the reasoning capabilities and task-specific expertise of large language models (LLMs). Despite recent advances in post-training methods, such as Group Relative Policy Optimization (GRPO), their practical deployment remains impeded by training instability arising from the reliance on importance sampling. We introduce Group Variance Policy Optimization (GVPO), a novel post-training method that integrates the analytical solution of KL-constrained reward maximization into its gradient weighting scheme. This formulation provides an intuitive interpretation: GVPO's gradient corresponds to the mean squa

This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.

Read the full story at arXiv cs.AI
  • Forskning
  • Reglering

Related AI news