RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests
arXiv cs.AIen
arXiv:2608.27831v1 Announce Type: new Abstract: Coding agents are now commonly evaluated on the SWE-bench family of benchmarks, whose tasks are built from curated GitHub issues--long, structured, and information-rich. Real user requests, however, are typically far shorter and less structured. To characterize this gap, we define a six-category information taxonomy and four dimensions of linguistic style, and apply them to real user prompts from SWE-chat and problem statements from SWE-bench Verified and Pro. We find that requests carrying only a problem statement, alone or with limited additional context, account for 88% of real prompts but just 7% of benchmark problems. Furthermore, 87% of r
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
- Företag
Related AI news
- 엘리스그룹, IPO 자금 활용 '모듈 데이터센터' 전국 확대한다ETNews (KR) · August 31, 2026
- Urheberrechtsklage gegen KI-Entwickler: Sony und Warner verklagen AnthropicGolem.de · August 31, 2026
- Big Tech reported Q2 "other income" rose significantly to $160B+, driven by investments in AI companies, raising concerns of paper gains overstating the AI boom (Financial Times)Techmeme · August 31, 2026
- IPO documents: SoftBank's SB Energy plans to file for an IPO as soon as this week, aiming to raise $5B to $7B, and has awarded OpenAI warrants worth ~$5.5B (Anissa Gardizy/Wall Street Journal)Techmeme · August 31, 2026
- OpenAI issued warrants worth $5.5 billion in SB Energy, WSJ reportsEconomic Times Tech · August 31, 2026
- AI and robotics drive an IPO boom in China as Shein lists in Hong KongEconomic Times Tech · August 31, 2026