Sense and Sensitivity: Benchmarking LLM Clinical Triage Recommendations with Physician Experts
arXiv cs.AIen
arXiv:2609.38600v1 Announce Type: new Abstract: As large language models (LLMs) are increasingly used in clinical settings, it is critical to evaluate their reliability under realistic variation in clinical text. We study this question in clinical triage, comparing LLMs to practicing physicians under text perturbations that preserve the underlying clinical setting. We introduce a benchmark of over 6,000 clinical scenarios, 7,000 physician annotations, and 225,000 model responses. Using this benchmark, we make two key observations. First, LLMs are more likely than physicians to recommend unnecessary care at baseline, and this tendency increases under perturbed inputs. Further, we find that LL
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Kaikki noudattivat ohjeita – Kukaan ei ollut vastuussaTivi · October 1, 2026
- China's chip-tool localization accelerates: ACM Research backlog jumps 88%DIGITIMES · October 1, 2026
- Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization LimitsarXiv cs.AI · October 1, 2026
- ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via CodearXiv cs.AI · October 1, 2026
- Can an AI Agent Rediscover a Blaschke-Curve Invariant?arXiv cs.AI · October 1, 2026
- SimTrace: Grounded Multimodal User Trajectories Generation for Online User ModelingarXiv cs.AI · October 1, 2026