DocHop: Benchmarking Out-of-domain Multi-hop Reasoning in Information-Dense Documents
arXiv cs.AIen
arXiv:2609.02059v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) have achieved strong performance on structured visual understanding tasks such as chart and document question answering. However, existing benchmarks typically evaluate these domains in isolation, leaving underexplored a key capability: whether models can use textual context to determine how chart evidence should be selected, interpreted, and aggregated. We introduce DocHop, a benchmark for integrated chart--context reasoning in document-style images. In DocHop, the document narrative specifies multi-step compositional constraints, while charts provide the corresponding data values. Questions are grounde
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Bild
Related AI news
- Architecting Conversational Data Systems for Stateless LLM APIs: The Hydration Proxy PatternarXiv cs.AI · September 3, 2026
- FUSE: An Evaluating Framework for Dangerous Capabilities of LLMsarXiv cs.AI · September 3, 2026
- Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?arXiv cs.AI · September 3, 2026
- When Does Information Sharing Improve Decentralized Discovery? Aggregation, Independent Rescue, and Equilibrium SelectionarXiv cs.AI · September 3, 2026
- READY or Not: Reliable Enterprise Agent DeploymentarXiv cs.AI · September 3, 2026
- When Can a Machine Trust a Statute? A Survival Certificate for Machine-Extracted Legal LogicarXiv cs.AI · September 3, 2026