ClaimReceipt: Verifying Evidence Sufficiency and Coverage in Agent Evaluations
arXiv cs.AIen
arXiv:2609.01992v1 Announce Type: new Abstract: Agent evaluations face two distinct evidentiary questions: whether a reported claim is recomputable from retained evidence (sufficiency), and whether the retained records cover the committed experiment set (coverage). Generic logs and hash-linked transcripts answer neither reliably. We introduce ClaimReceipt, a claim-relative receipt specification and selective verifier that binds typed transaction evidence to a signed experiment manifest and returns PASS, INVALID, or INCONCLUSIVE per claim. We freeze the specification before implementation (SHA-256 18d109...b81). On 1,392 historical buyer--seller records, a CR-2 verifier reproduces all five ma
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
- Företag
Related AI news
- OpenAI is building 'automated shutdown' capabilities for AI tools, letter to lawmakers saysEconomic Times Tech · September 3, 2026
- Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy PatternarXiv cs.AI · September 3, 2026
- FUSE: An Evaluating Framework for Dangerous Capabilities of LLMsarXiv cs.AI · September 3, 2026
- READY or Not: Reliable Enterprise Agent DeploymentarXiv cs.AI · September 3, 2026
- When Can a Machine Trust a Statute? A Survival Certificate for Machine-Extracted Legal LogicarXiv cs.AI · September 3, 2026
- Benchmarking Language Models for Statistical Problem FormulationarXiv cs.AI · September 3, 2026