SemVerBench: Benchmarking LLM Comprehension of Version-Constraint Resolution Semantics
arXiv cs.AIen
arXiv:2609.11180v1 Announce Type: new Abstract: Large language model (LLM) coding agents constantly decide whether a version satisfies a constraint such as ^1.2.3 or >=2.0, 1.2 means >=1.3.0) traps every model on Cargo (near 60%), and although standard PEP 440 prefix matching is universal, on zero-pad/post-release corner cases GPT-5.1 collapses (0/26) while Claude stays at 97-100% (verified on a 67-item oracle-validated set). Opus significantly outperforms all other models, and Sonnet outperforms the OpenAI models (McNemar). The failures look more like an activation/application gap than a knowledge gap: injecting the rule or a light correct hint recovers most errors, whereas interval decompo
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- OpenAI
- Anthropic
- Verktyg
- Forskning
Related AI news
- Nvidia in talks to invest $10b in Anthropic IPO: sourcesTech in Asia · September 12, 2026
- A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive ReasoningarXiv cs.AI · September 12, 2026
- CryptoL: Towards Scale Dominance and Physics Constraints Mitigation in Financial Multivariate Time Series ForecastingarXiv cs.AI · September 12, 2026
- Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM WorkflowsarXiv cs.AI · September 12, 2026
- Agentic Share-of-Search: A Multi-Agent AI System for Competitive Decision-Making in LLM-Mediated E-CommercearXiv cs.AI · September 12, 2026
- Debate-to-Skill: Capability-Bound Process Supervision for Industrial Query-to-Agent AnnotationarXiv cs.AI · September 12, 2026