Reinforcement Learning with Verifiable Rewards for Small Search Agents
arXiv cs.AIen
arXiv:2609.28765v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) performs well on problems with clear rewards, such as mathematics and coding, but whether it also works where the reward is less clear remains open. The reason-over-search recipe applies RLVR to open-domain question answering, where retrieval grounds the answer and a match against the reference supplies the reward. So far it has been demonstrated on large models, and below one billion parameters only with distillation from a larger teacher. We test the recipe on a small model. We train Qwen3.5-0.8B with Group Relative Policy Optimization (GRPO) and an interleaved Wikipedia-search tool on MuS
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
- Reglering
Related AI news
- "One small step for TPUs": Google CEO Sundar Pichai announces Project Suncatcher to test AI compute in SpaceEconomic Times Tech · September 25, 2026
- PAWS: Policy-driven Agentic World SimulationarXiv cs.AI · September 25, 2026
- Functional Architecture of European Electricity Trading Markets: Requirements for AI Supported Trading Systems under Regulatory ConstraintsarXiv cs.AI · September 25, 2026
- AI-satsningar hindrar viktiga moderniseringsprojektComputer Sweden · September 25, 2026
- When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability RoutingarXiv cs.AI · September 25, 2026
- TW3Cast: A Frozen Router of Lightly Fine-Tuned Foundation Models for Time-Series Forecasting on GIFT-Eval, Selected Entirely on the Training SplitarXiv cs.AI · September 25, 2026