ContractEval: Query-Conditioned Execution Matching for Procedural Instruction Conformance
arXiv cs.AIen
arXiv:2609.09458v1 Announce Type: new Abstract: As LLM agents move from answering questions to carrying out procedures, failures can be unwarranted rather than visibly wrong: the final response looks acceptable even though the system skipped the check, branch, dependency, or invariant that made the answer justified. Output-only evaluation sees the answer, and trace-aware judging sees activity, but neither identifies which obligations were active for the query. We introduce CONTRACTEVAL, a diagnostic framework for making those active obligations explicit. It represents procedural instructions as query-active obligations and matches them against response or trace evidence, turning omissions, w
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
- Företag
Related AI news
- Meta share price surges after personal AI agent Muse releaseEconomic Times Tech · September 10, 2026
- Exclusive: Bynario raises €2.1m to tackle cybersecurity’s AI slop problemSifted · September 10, 2026
- Donnerstag: Apples neue iPhones auch aufklappbar, KI-Agenten weiter ungezügeltheise online – KI · September 10, 2026
- Generative AI a new tool in Mali's information war: studyEconomic Times Tech · September 10, 2026
- Sources: Alibaba is set to lead a $300M round in AI model testing startup UniPat AI at a $2.5B valuation; UniPat founder Li Kuan worked at Alibaba's Tongyi Lab (Bloomberg)Techmeme · September 10, 2026
- OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist WorkflowsarXiv cs.AI · September 10, 2026