Asclepius: An Adaptive Harness for Long-Horizon Clinical Agents
arXiv cs.AIen
arXiv:2609.13543v1 Announce Type: new Abstract: LLM agents are predominantly benchmarked on short, single-task trajectories, yet real deployments run for hours under contention, surfacing a different class of failures. We use the Clinical Environment Simulator (CES), in which an agent manages an entire emergency-department shift under continuous time and resource pressure, as a testbed: long-horizon execution failures manifest measurably in a single rollout under structured, multi-dimensional grading. On CES, current agents reach the correct diagnosis in most cases yet fail to deliver complete and timely critical actions, revealing an execution gap. We attribute this gap to three long-horizo
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Agenter
Related AI news
- AI agents are moving faster than Europe can regulateSifted · September 15, 2026
- Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting AgentsarXiv cs.AI · September 15, 2026
- Token Efficient Task Execution via Application Behavior Modeling for Web AgentsarXiv cs.AI · September 15, 2026
- Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment GenerationarXiv cs.AI · September 15, 2026
- AutoTailor: Automatic, User-Aligned Capability Selection and Adaptation for Web AgentsarXiv cs.AI · September 15, 2026
- Toward a Decision-Assurance Layer for AI-Assisted Flight Planning in Air Traffic ManagementarXiv cs.AI · September 15, 2026