Rules to Tools: Executable Checks for LLM Agents in Scientific Computing
arXiv cs.AIen
arXiv:2610.00313v1 Announce Type: new Abstract: Scientific coding agents receive equations, boundary conditions, and output requirements in writing, then must assess the programs they revise. Rules to Tools (R2T) supplies prepared executable checks of public scientific requirements. Matched SciCode repair groups share written checks, starting programs, model, and budgets; the tool group receives a callable implementation. Across two task-ID cohorts, complete repair is 26/30 with text and 29/30 with the prepared checks. Three task IDs favor tools, one favors text, and eleven tie. The eight-ID cohort scores 13/16 versus 15/16, with a task-cluster bootstrap 95% interval of [-12.5, 43.75] percen
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
Related AI news
- LG Uplus, OptAI team up on AI token optimizationTech in Asia · October 2, 2026
- AutoSynthData: Generating Training Data for Enterprise AgentsHugging Face · October 2, 2026
- Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMsarXiv cs.AI · October 2, 2026
- Benchmarking Prompt Optimization of Large Language Models With ChessarXiv cs.AI · October 2, 2026
- Measuring the Microtask Eligibility Gap: When Is an Off-the-Shelf SLM Enough for an Agent Harness?arXiv cs.AI · October 2, 2026
- Comedic Fool's Gold: Reward Exploits and Countermeasures in Conversational HumorarXiv cs.AI · October 2, 2026