C$^3$PO: Evaluating Cross-Modal Composition and Counterfactual Performance in Omnimodal Models
arXiv cs.AIen
arXiv:2608.05381v1 Announce Type: new Abstract: Current Multimodal Large Language Models (MLLMs) can process diverse sensory inputs, yet their reasoning remains heavily biased toward a dominant modality, resulting in brittle cross-modal reasoning. We introduce C$^3$PO, a benchmark of 3,404 samples spanning video, audio, image, and text, evaluating two abilities: information composition (fusing dispersed evidence) and counterfactual conflict (resolving deliberate contradictions). C$^3$PO's paired IC/CC structure and four-tier design enable targeted diagnosis of when and why cross-modal reasoning fails. Built through a fully automatic pipeline using 25 logically grounded templates, C$^3$PO rev
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Bild
Related AI news
- The Ignition Index: Measuring Global Workspace Dynamics in Language ModelsarXiv cs.AI · August 7, 2026
- Otter: A Time-Aware, History-Conditioned Human Chess AIarXiv cs.AI · August 7, 2026
- Project2Task: Graph-Guided Project-Level Planning for Autonomous ResearcharXiv cs.AI · August 7, 2026
- Agentic self-driving microscopy benchmarks support qualification but do not necessarily generalize to unseen tasksarXiv cs.AI · August 7, 2026
- CASCADE: An Agentic Regulatory Network Framework for Patient-Data-Validated Downstream Perturbation PredictionarXiv cs.AI · August 7, 2026
- DoctorAgents: an agentic framework to iteratively refine AutoML pipeline for small clinical temporal dataarXiv cs.AI · August 7, 2026