A collective capability boundary in frontier large language models on guideline-conformant and case-specific oncology decision-making
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2608.28592v1 Announce Type: new Abstract: Large language models (LLMs) achieve high scores on medical knowledge examinations, yet real-world oncology is not a knowledge test--it is a sequence of guideline-pathway choices, escalation judgments, and commitments under uncertainty. Existing benchmarks largely measure factual recall, leaving open whether frontier LLMs share decision-path blind spots that combining models cannot fix. We built the Oncology Decision Boundary Benchmark (ODBB)--2,005 oncology decision points across NCCN guidelines and colorectal cancer cases--and evaluated nine frontier LLMs (four closed-source, five open-weight families) released between June 2025 and April 202
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Anthropic launches Claude Fable 5.1 and Mythos 5.1, cuts agentic-task costs by up to 45%DIGITIMES · September 2, 2026
- 先進封裝邁向「化圓為方」!美商 ACM Research 卡位 FOPLP,電鍍、清洗、濕式蝕刻「三箭齊發」TechNews (TW) · September 2, 2026
- CrowdStrike builds security frontier models with Nvidia and opens an AI labSiliconANGLE · September 1, 2026
- Anthropic opens Claude AI text detection to regulators, media, fact-checkers, and othersThe Decoder · September 1, 2026
- Anthropic's Claude Fable 5.1 promises better coding and research at up to 45 percent lessThe Decoder · September 1, 2026
- AI labs are facing an agent control problemAxios · September 1, 2026