Agreement Overstates Evidence: Error Dependence in LLM Judge Consensus
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.22512v1 Announce Type: new Abstract: Consensus among LLM judges is often taken as strong evidence that a decision is correct. This assumes that judges make their errors independently. In practice, LLM judges are often trained and evaluated in similar ways, so they can make the same mistakes. We study how this dependency affects the reliability of consensus. We find substantial error correlation across both open-weight and frontier LLM judges. In our main bank of ten judges, the average pairwise correlation between judge errors is 0.21. As a result, the ten judges only provide roughly as much statistical information as 3.5 independent judges. The dependency is even stronger among t
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Do Existing Preconditioners Improve Biomedical Tabular Foundation Learning? An Empirical Study on TabPFN OptimizationarXiv cs.AI · September 23, 2026
- 4DGS-JEPA: Temporally Compositional Joint-Embedding Prediction for Dynamic Gaussian SplattingarXiv cs.AI · September 23, 2026
- An Accurate and Interpretable Hyper Graph Neural Network for GBM Survival PredictionarXiv cs.AI · September 23, 2026
- Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal EmbeddingsarXiv cs.AI · September 23, 2026
- X-Planner: Event-Structured Task Planning for Embodied IntelligencearXiv cs.AI · September 23, 2026
- Lean Pool: An AI-Maintained Archive of Formalized MathematicsarXiv cs.AI · September 23, 2026