There Is No Neutral Harness: Modern LLM Leaderboards Are Manufactured by Config-Fragile Items
arXiv cs.AIen
arXiv:2608.21382v1 Announce Type: new Abstract: Multiple-choice benchmarks fix the questions and the correct answers, but not the harness: the order of the options, the wording of the prompt, and whether a language model's answer is read from generated text or from per-option likelihoods. Work on this harness sensitivity reports it as aggregate score variance, leaving unexamined which items the variance falls on and whether they are the items that separate one model from the next. We treat the evaluation harness of large language models (LLMs) as an independent variable and resolve its effect to single items. We introduce the \textit{fragility grid}: 12 open-weight instruction-tuned LLMs fro
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Företag
Related AI news
- AI chipmaker Enflame sets subscription date for near $900 million Shanghai IPOEconomic Times Tech · August 25, 2026
- Measuring Activation Control in Large Language ModelsarXiv cs.AI · August 25, 2026
- KVBoost: Chunk-Level Key-Value Cache Reuse with Deviation-Guided Recomputation for Efficient Large Language Model InferencearXiv cs.AI · August 25, 2026
- AIREP: A Protocol for Per-Decision Evidence in AI Runtime GovernancearXiv cs.AI · August 25, 2026
- LitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review PlatformarXiv cs.AI · August 25, 2026
- SchemaRouter: Field-Aware Tool Routing for Efficient Heterogeneous Agentic RAGarXiv cs.AI · August 25, 2026