CARGO-VL: Counterfactual Arbitration with Risk-Constrained Group Optimization for Vision-Language Models
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2608.04509v1 Announce Type: new Abstract: Vision-language systems combine images with retrieved text, but these sources can disagree or jointly fail to support an answer. Reliable models must identify the trustworthy source and abstain when neither is adequate. Existing post-training objectives score instances independently and therefore do not enforce coherent behavior under counterfactual evidence changes. We introduce CARGO-VL, a group-relative framework that optimizes matched variants covering aligned, image-correct, text-correct, and both-wrong (A/V/T/N) evidence states as one bundle. Its objective couples condition-wise correctness with transition rewards for answer invariance, s
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Bild
Related AI news
- DeepSeek signals ‘significant’ price hike amid surge in demand for low-cost AI modelsSCMP Tech · August 6, 2026
- Donnerstag: Aus für Google Assistant, Snapchat-Verbot von KI-Videosheise online – KI · August 6, 2026
- Improving Auto-Design of Neural PDE Solvers with a Domain-Specific LanguagearXiv cs.AI · August 6, 2026
- FinProBench: Evaluating Financial AI Agents with Role-Grounded Rubrics Derived from Professional DeliverablesarXiv cs.AI · August 6, 2026
- Architectural Implications of Agentic AI WorkflowsarXiv cs.AI · August 6, 2026
- What Is a Skill Worth? Structure-Aware Shapley Valuation of Agent SkillsarXiv cs.AI · August 6, 2026