Grounded Adjudication of Variations across Extracted TimeLines (GAVEL): Comparing Clinical Timelines Against Their Case Reports
arXiv cs.AIen
arXiv:2609.13475v1 Announce Type: new Abstract: Existing pipelines for clinical timeline extraction from case reports are evaluated using an expert reference and are limited by imperfect reference annotations and imprecise event alignment. We developed GAVEL, an LLM judge protocol that compares two timelines with the case report and returns a discrepancy type, verdict, and report passage for each difference. We evaluated the event matcher, reviewed 2,738 findings from GPT5.6sol and DeepSeek V3.2, ranked six LLM extractors and two human annotators, and tested GAVEL guided merging. True match rates were 60% immediately below and 48% immediately above the 0.10 cutoff. Manual review confirmed 89
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- DeepSeek
- Forskning
Related AI news
- Microsoft commits to sweeping AI privacy rules for students. Will other tech giants follow?Economic Times Tech · September 15, 2026
- OrchSLM: Probing the Dynamics of Small Language Model OrchestrationarXiv cs.AI · September 15, 2026
- Asclepius: An Adaptive Harness for Long-Horizon Clinical AgentsarXiv cs.AI · September 15, 2026
- Vibe Patenting: Evaluating LLM Judges for Professional Patent-Drafting AgentsarXiv cs.AI · September 15, 2026
- Token Efficient Task Execution via Application Behavior Modeling for Web AgentsarXiv cs.AI · September 15, 2026
- Fraglingo: Molecular Design via Attachment-Aware Autoregressive Fragment GenerationarXiv cs.AI · September 15, 2026