Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks
arXiv cs.AIen
arXiv:2607.28685v1 Announce Type: new Abstract: Agent-safety benchmarks measure different behaviors, and their scores get quoted interchangeably as an agent's safety. We treat four of them (R-Judge, InjecAgent, AgentHarm, AgentDojo) as measurements to be validated, running each under its official implementation and author-provided scorer on up to 22 models, with MMLU and GPQA measured by us under one protocol as a capability composite. The metric is the first problem. On any binary trace-judgment benchmark scored by $F_1$, an ``always positive'' policy attains $F_1 = 2\pi/(1+\pi)$; on R-Judge that is $0.690$, above five of the 21 models that actually discriminate. The three broad-coverage be
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
- Reglering
Related AI news
- Europe starts enforcing AI Act rulesSifted · August 3, 2026
- An Ontology-Guided, Deduplication-Aware Extraction Layer for Knowledge Graph Construction from Heterogeneous DocumentsarXiv cs.AI · August 3, 2026
- EarlyDx: An Admission-Anchored Benchmark for Open-Ended Generation of Evidence-Supported ED-Encounter DiagnosesarXiv cs.AI · August 3, 2026
- LLM Framework for Discovering Major Mathematical Conjectures: AI's Quest for the Next Riemann HypothesisarXiv cs.AI · August 3, 2026
- Library Reachability in LSR-Synth: How Anti-Memorization Design Changes the Measurement of Symbolic DiscoveryarXiv cs.AI · August 3, 2026
- MMShopBench: A Real-Log Benchmark for Multimodal, Multi-Turn Shopping AgentsarXiv cs.AI · August 3, 2026