SearchAuditor: Auditing and Attributing Failures in Long-Horizon Search Agents
arXiv cs.AIen
arXiv:2608.05212v1 Announce Type: new Abstract: Deep search agents tackle challenging questions through long-horizon web interactions, a process that is both complex and fragile: small reasoning errors may propagate through long, noisy trajectories into fluent but incorrect answers. Diagnosing such failures is difficult, requiring the manual inspection of extremely long execution traces, which could be beyond human capacity. We therefore introduce SearchAuditBench, a benchmark that evaluates whether LLM auditors can localize, attribute, and repair these failures, thereby reducing the human burden. SearchAuditBench comprises 1,243 failed trajectories, averaging 73.1 messages and 65.1K tokens,
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
Related AI news
- The Ignition Index: Measuring Global Workspace Dynamics in Language ModelsarXiv cs.AI · August 7, 2026
- Otter: A Time-Aware, History-Conditioned Human Chess AIarXiv cs.AI · August 7, 2026
- Project2Task: Graph-Guided Project-Level Planning for Autonomous ResearcharXiv cs.AI · August 7, 2026
- C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal ModelsarXiv cs.AI · August 7, 2026
- Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasksarXiv cs.AI · August 7, 2026
- CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation PredictionarXiv cs.AI · August 7, 2026