Don't Peek at the Answer: Outcome-Masked Group Relative Policy Optimization for Label-Free RLVR
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2608.03119v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) improves LLM reasoning but typically relies on ground-truth (GT) answers, limiting scalability. Voting-based label-free RLVR replace gold supervision with answer-level consensus from model samples. However, collapse arises when the same answer-level signal is used both to estimate rewards and to drive token-level policy optimization, encouraging the model to directly reinforce answer tokens rather than improve reasoning. We propose OM-GRPO, a label-free RLVR framework that decouples reward estimation from policy optimization. OM-GRPO masks gradients on the answer span while retaining answer-
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Reglering
Related AI news
- Google targets AI startup Mechanize’s technology and talent in proposed $1.5B dealSiliconANGLE · August 6, 2026
- ByteDance's new "watch and listen" AI signals a broader Chinese push beyond chatbotsDIGITIMES · August 6, 2026
- Meta takes on Anthropic and OpenAI with its first AI coding agent, Muse CodeSiliconANGLE · August 6, 2026
- Confirming rumors, Anthropic reveals plan to develop custom chipSiliconANGLE · August 6, 2026
- How OpenAI's agents broke out of testing to hack Hugging FaceAxios · August 6, 2026
- Google's AI leadership shake-up puts Gemini execution and research retention under pressureDIGITIMES · August 6, 2026