RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

arXiv cs.AIen

RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests

arXiv:2608.27831v1 Announce Type: new Abstract: Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues--long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured. To characterize this gap, we define a six-category information taxonomy and four dimensions of linguistic style, and apply them to real user prompts from SWE-chat and problem statements from SWE-bench Verified and Pro. We find that requests carrying only a problem statement, alone or with limited additional context, account for 88% of real prompts but just 7% of benchmark problems. Furthermore, 87% of r

This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.

Read the full story at arXiv cs.AI
  • Verktyg
  • Forskning
  • Agenter
  • Företag

Related AI news