GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings
Apple Machine Learningen
Apple Machine Learning
AI Global WireReinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric. We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of base models, training languages, and different reasoning language rewards. We find that training to reason in the native language often leaves only a small gap to training for English reasoning. We further observe strong…
This is a short summary published by AI Global Wire. The full article is owned and hosted by Apple Machine Learning — open it there to read it in full.
Read the full story at Apple Machine Learning- Forskning
- Reglering
Related AI news
- Anthropic CEO says AI centralizes by nature and open models just shift power to whoever owns the chipsThe Decoder · August 18, 2026
- As AI beats doctors, regulators shouldn't force a human into the loop, JAMA piece saysThe Decoder · August 18, 2026
- DOJ probes Andreessen Horowitz over partners sitting on competing AI boardsThe Decoder · August 18, 2026
- OpenAI debuts ChatGPT for Teens, a mode that limits high-risk chats about self-harm, eating disorders, violence, and more, and has studying tools and guardrails (Cecilia Kang/New York Times)Techmeme · August 18, 2026
- OpenAI launches ChatGPT for Teens, a new mode that limits high-risk chats around self-harm, eating disorders, and other topics, adds studying tools, and more (Cecilia Kang/New York Times)Techmeme · August 18, 2026
- We still don’t know how people are really using AIMIT Technology Review · August 18, 2026