ContractRL: Shielded Group-Relative Policy Optimization for Auditable Tool-Call Repair
arXiv cs.AIen
arXiv:2610.00328v1 Announce Type: new Abstract: Structured tool calls often fail after only a small number of fields violate a schema or an execution contract. Regenerating the complete object enlarges the action surface and makes repeated repair difficult to audit. We introduce ContractRL, a contract-constrained sequential repair protocol that models verifier-guided JSON repair as a bounded decision process. At each step the policy observes the candidate, typed verifier feedback, JSON Pointer, immutable repair history, and remaining budget; a contract-derived action mask filters malformed or prohibited RFC-6902 operations before a deterministic validator performs the transition. We specify
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Reglering
Related AI news
- ArXiv limits preprint submissions to two per month per submitter, as AI access fuels a record 40,363 submissions in September 2026, vs. 20,569 in September 2024 (Kat Boboris/arXiv)Techmeme · October 2, 2026
- LG Uplus, OptAI team up on AI token optimizationTech in Asia · October 2, 2026
- Google's first Suncatcher satellite reaches orbit to test AI chips in spaceDIGITIMES · October 2, 2026
- Before Agents Decide: Epistemic Action in LLM-Based SystemsarXiv cs.AI · October 2, 2026
- Heavy-Tailed Memory Traces in Long-Horizon Language AgentsarXiv cs.AI · October 2, 2026
- Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?arXiv cs.AI · October 2, 2026