Harness or Model? Isolating the Harness Effect in Agentic Coding with a Contamination-Controlled Private Suite
arXiv cs.AIen
arXiv:2609.11987v1 Announce Type: new Abstract: An agentic coding system couples a language model to a harness: the tools, prompts and control flow that turn a chat model into an autonomous software engineer. Vendors ship harnesses tuned to their own models, and practitioners assume the vendor-native pairing solves more tasks. We measure that assumption with paired same-model contrasts on a private, contamination-controlled suite of 256 repository and post-cutoff contest tasks. The same 80 tasks ran under claude-agent-sdk and under deepagents on claude-opus-4-8, and under the openai-codex SDK and deepagents on gpt-5.5, with gemini-3.5-flash and deepseek-v3.2 as side cells. 792 of 800 planned
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- OpenAI
- Anthropic
- DeepSeek
Related AI news
- Noch vor Hugging-Face-Hack: KI-Agenten von OpenAI haben Rubygems attackiertGolem.de · September 14, 2026
- 打破入門機框架!Googlebook 規格曝光,搭 15.3 吋 OLED 螢幕與 Core Ultra 5 處理器TechNews (TW) · September 14, 2026
- AI stocks slide after major CEOs unite to urge slowdownCNBC Technology · September 14, 2026
- SoftBank Group plunges 11% after OpenAI says no IPO this yearNikkei Asia · September 14, 2026
- How hyperscalers like Amazon, Microsoft, and Google are siding with consumers on data center power costs and sweetening offers for communities to gain support (Ann Davis Vaughan/The Information)Techmeme · September 14, 2026
- 美政府 AI 沙皇嗆 Anthropic 與 OpenAI 想放緩就放緩,別假裝需要他人同意TechNews (TW) · September 14, 2026