Can LLMs Follow Medical Expert Logic? A Benchmark for Hierarchical Logical Consistency in Risk-of-Bias Assessment
arXiv cs.AIen
arXiv:2609.11185v1 Announce Type: new Abstract: Evidence-based medicine demands strict logical consistency, yet current evaluations of large language models (LLMs) prioritize superficial label matching over genuine reasoning. We introduce LogiMed-RoB, a benchmark grounded in Cochrane Risk of Bias (RoB) 2.0 expert logic, comprising 860 randomized controlled trials (RCTs) and 14,820 queries. It evaluates models under the Hierarchical Logical Consistency (HLC) framework across four dimensions: Atomic Consistency, Domain Consistency, Aggregation Consistency, and Evidential Faithfulness. Experiments on 10 state-of-the-art LLMs reveal a catastrophic Error Compounding Effect: despite the top model
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
- Företag
Related AI news
- NYC-based Luminary, which develops AI-powered workflow tools for estate planning and wealth transfer management, raised a $22M Series A led by Ten Coves Capital (Davis Janowski/Wealth Management)Techmeme · September 12, 2026
- Nvidia in talks to invest $10b in Anthropic IPO: sourcesTech in Asia · September 12, 2026
- Sequoia leads Mecka AI round at nearly $500m valuationTech in Asia · September 12, 2026
- A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive ReasoningarXiv cs.AI · September 12, 2026
- The Oligarch Barely Steers Model Collapse in Multi-Model EcosystemsarXiv cs.AI · September 12, 2026
- CryptoL: Towards Scale Dominance and Physics Constraints Mitigation in Financial Multivariate Time Series ForecastingarXiv cs.AI · September 12, 2026