A Jagged Frontier: Evaluating Robustness of Code Agents to Semantics-Preserving Transformations
arXiv cs.AIen
arXiv:2608.18389v1 Announce Type: new Abstract: AI code agents are increasingly deployed to resolve real software issues, yet their reliability under superficial code variations remains poorly understood. We evaluate whether coding agents that repair repository-level issues remain reliable when the surrounding codebase is rewritten into a semantically equivalent form. We introduce a random variant sampler that applies common semantics-preserving transformations (SPTs) - spanning control-flow rewrites, dead-code injection, and identifier renaming - to produce perturbed variants. We evaluate two agentic scaffolds (mini-SWE agent and OpenCode) each backed by one of four frontier models (Claude
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Anthropic
- Verktyg
- Forskning
- Agenter
Related AI news
- How Unitree's Go series, which helped the company dominate the quadruped robot market, drew on openly published US university research funded by the US military (Michael Martina/Reuters)Techmeme · August 20, 2026
- Optimized Fuzzy Logic Approach with the IEEE Key Gas Method for Diagnosing Power Transformer Faults Using Dissolved Gas AnalysisarXiv cs.AI · August 20, 2026
- Position: Profiling Game Worlds by Transition ComplexityarXiv cs.AI · August 20, 2026
- SESSE: Sketch, Expand, Sort, Summarize, Evaluate -- LLM-as-Judge Evaluation via Structured DecompositionarXiv cs.AI · August 20, 2026
- Measuring the Partial-Credit Gap: A Strict Benchmark on Vietnam's 2025 Convex Marking SchemearXiv cs.AI · August 20, 2026
- FinSkillBench: Evaluating AI Agents and Domain Skills for Investment ManagementarXiv cs.AI · August 20, 2026