More Programs or More Rolls? Separating Coverage from Specialization in LLM Harnesses
arXiv cs.AIen
arXiv:2609.35873v1 Announce Type: new Abstract: Automated generation of LLM harnesses promises to improve inference through task specialization. Yet additional answer coverage can arise from repeated execution of the same program, making specialization difficult to identify. We introduce a controlled evaluation that separates answer coverage, repeatable task advantages, and gains from pre-execution selection. On 386 MATH-500 tasks, we compare eight generated harnesses plus a baseline with nine byte-identical baseline copies, using three executions per member. Identical programs yield 2.16 percentage points of repeat-averaged oracle headroom. Generated programs exhibit substantially more repe
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Företag
Related AI news
- Risk-Averse Online POMDP Planning via CVaR of the Immediate Cost with Performance GuaranteesarXiv cs.AI · September 30, 2026
- OpenAI-HuggingFace: A Reproduction & Lessons for Alignment TestingarXiv cs.AI · September 30, 2026
- Neurosymbolic Routing for Reliable Reasoning on Resource-Constrained Edge DevicesarXiv cs.AI · September 30, 2026
- Beyond Symmetric Agents: Cognitive Diversity and Multi-Agent Debate in Small Language ModelsarXiv cs.AI · September 30, 2026
- Representational Simplicity and Circuit Size Dissociate in a Threshold-Dependent Way: A Controlled Test via Adversarial TrainingarXiv cs.AI · September 30, 2026
- Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled ObjectivesarXiv cs.AI · September 30, 2026