OpenDiscoveryTrace: Process Traces for Evaluating AI Scientist Workflows
arXiv cs.AIen
arXiv:2609.09203v1 Announce Type: new Abstract: Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained. This makes it impossible to audit scientific methodology, diagnose failure modes, or distinguish systematic reasoning from fortunate guessing. We present \textbf{OpenDiscoveryTrace}, a public dataset of 558 complete AI scientific agent trajectories that captures how models reason, not just what they produce. Each trajectory records a structured 9-field-per-step trace---including thoughts, tool calls, observations, errors, revision triggers, and self-reporte
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Agenter
Related AI news
- Meta share price surges after personal AI agent Muse releaseEconomic Times Tech · September 10, 2026
- Donnerstag: Apples neue iPhones auch aufklappbar, KI-Agenten weiter ungezügeltheise online – KI · September 10, 2026
- Generative AI a new tool in Mali's information war: studyEconomic Times Tech · September 10, 2026
- Adaptive Entangled Game Modules in Artificial General IntelligencearXiv cs.AI · September 10, 2026
- An Autonomous GeoAI Agent for Arctic Eco-NavigationarXiv cs.AI · September 10, 2026
- The Menu Is an Execution Prior: State-Path Tool Menus for Online AgentsarXiv cs.AI · September 10, 2026