Human-AI-Powered Hypothesis Testing: Cost-Aware Selective AI Scoring and Sequential Human Escalation
arXiv cs.AIen
arXiv:2609.28859v1 Announce Type: new Abstract: Large language models are increasingly used as inexpensive judges to evaluate outputs, label data, and assess whether a system meets a desired quality standard. Yet using AI judgments for formal statistical inference is fundamentally different from simply treating them as ground-truth labels: AI evaluations can be biased or noisy, and rigorous hypothesis testing requires explicit control of type-I and type-II errors. We study how to use AI judgments, together with selective human verification, to conduct a valid hypothesis test at minimum cost. We consider a population of items with hidden binary labels. After choosing a fixed pool of items, th
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Företag
Related AI news
- "One small step for TPUs": Google CEO Sundar Pichai announces Project Suncatcher to test AI compute in SpaceEconomic Times Tech · September 25, 2026
- PAWS: Policy-driven Agentic World SimulationarXiv cs.AI · September 25, 2026
- Functional Architecture of European Electricity Trading Markets: Requirements for AI Supported Trading Systems under Regulatory ConstraintsarXiv cs.AI · September 25, 2026
- When Should Forecasting Agents Reason? Behavioral Stress Tests for Reliability RoutingarXiv cs.AI · September 25, 2026
- TW3Cast: A Frozen Router of Lightly Fine-Tuned Foundation Models for Time-Series Forecasting on GIFT-Eval, Selected Entirely on the Training SplitarXiv cs.AI · September 25, 2026
- BaseCamp --- An Agentic AI Framework for Automating DNA Sequencing Data PipelinesarXiv cs.AI · September 25, 2026