DI-Bench: Systematically Generating In-Domain Data Intelligence Benchmarks for Enterprise Agents
arXiv cs.AIen
arXiv:2609.05776v1 Announce Type: new Abstract: Evaluating enterprise agents on domain-specific benchmarks is critical, yet public benchmarks rarely evaluate whether agents can integrate business knowledge with analytical computation, and constructing such benchmarks manually is costly. We present DI-Bench, a pipeline for generating realistic benchmarks for data intelligence (DI), the practice of extracting insights from large volumes of enterprise data. To emulate realistic DI tasks that require both computation and knowledge retrieval, DI-Bench builds an artifact linkage graph over data tables, dimensions, metrics, and documents to form questions involving structured data and associated kn
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
Related AI news
- BharatPe launches Gemini-powered AI assistant for merchantsTech in Asia · September 9, 2026
- Mittwoch: Huawei-Verstöße gegen US-Sanktionen, Metas privater KI-Agent für alleheise online – KI · September 9, 2026
- Planning and Scheduling Business Processes under Control-Flow UncertaintyarXiv cs.AI · September 9, 2026
- Agents Trust Tools Too Much: Measuring Reliance on Unreliable ToolsarXiv cs.AI · September 9, 2026
- Distilling Vision-Language Models for On-Device Fire UnderstandingarXiv cs.AI · September 9, 2026
- When and What to Teach: Budget-Aware Online Adaptation for Web AgentsarXiv cs.AI · September 9, 2026