Right Order, Wrong Scale: Auditing LLM Judges for Occupational AI Measurement
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2610.02492v1 Announce Type: new Abstract: LLM judges are increasingly used to assess whether AI outputs meet workplace requirements, but agreement on response rankings does not establish agreement on acceptance rates or occupational aggregates. We introduce O*NET-BENCH, an audit suite derived from an existing survey of 45,796 worker ratings, and evaluate 33 pre-existing judge configurations across six model families on 4,501 test ratings. Twenty-five configurations achieve tie-aware pair accuracy of at least 0.60, although a train-fitted response-only TF-IDF baseline nearly matches the strongest judge. Despite this ordering agreement, judges estimate that 3.0%-97.9% of responses are ac
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- World Action Modeling with Progressive Visual PlanningarXiv cs.AI · October 5, 2026
- How to Have a Sensitive Debate: An Instance-Optimal Protocol for AI DebatearXiv cs.AI · October 5, 2026
- Choosing Before Acting: Comparative Value Estimation for Long-Horizon Tool-Use AgentsarXiv cs.AI · October 5, 2026
- A Multi Method Importance and Performance Efficiency Analysis of Topological Metrics for Natural Visibility Graph Based Cyber Attack DetectionarXiv cs.AI · October 5, 2026
- MintFlow: Minimal Trajectory Intervention for Constrained Flow MatchingarXiv cs.AI · October 5, 2026
- Fast Models, Slow Evidence: A Paired and Self-Audited Evaluation of System-1 Decision Models for LLM Agent HarnessesarXiv cs.AI · October 5, 2026