SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

arXiv cs.AIen

SLCA-GRPO: Resolving Cross-Segment Credit Misattribution in Tool-Calling RL

arXiv:2609.29050v1 Announce Type: new Abstract: Tool-calling agents produce heterogeneous outputs, interleaving structured tool invocations with user-facing natural language summaries. This output heterogeneity presents a structural failure mode in standard on-policy Reinforcement Learning (RL): algorithms like GRPO indiscriminately broadcast a homogeneous trajectory-level scalar advantage to all tokens. Consequently, gradient noise from summary generation leaks into tool-decision tokens, causing cross-segment credit misattribution and brittle optimization. In this work, we propose SLCA-GRPO, a framework incorporating Segment-Locked Credit Assignment (SLCA). To enable scalable exploration wi

This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.

Read the full story at arXiv cs.AI
  • Verktyg
  • Forskning
  • Agenter
  • Reglering

Related AI news