Towards a Deterministic Math Solver for Clinical Language Models
arXiv cs.AIen
arXiv:2609.10728v1 Announce Type: new Abstract: Large language models are unreliable at arithmetic, which is a problem for clinical calculators where a single numerical error changes the recommendation. The standard response is to hardcode each calculator as a validated function, one at a time. We test an alternative: the model does not calculate. Instead, it writes case-specific Python that a restricted local executor runs as a deterministic solver, and the model's task reduces to deciding how to use it. We evaluate this Program-Solve interface on MedCalc-Bench Verified (1,100 cases, 55 calculators) against direct model arithmetic and a hand-written 22-calculator library, using Qwen2.5-7B a
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- The Oligarch Barely Steers Model Collapse in Multi-Model EcosystemsarXiv cs.AI · September 12, 2026
- CryptoL: Towards Scale Dominance and Physics Constraints Mitigation in Financial Multivariate Time Series ForecastingarXiv cs.AI · September 12, 2026
- Decoupling Readiness from Release for Tail-Aware Scheduling of Agentic LLM WorkflowsarXiv cs.AI · September 12, 2026
- Defining AI Agents: A Compendium of Criteria, Metrics, and BenchmarksarXiv cs.AI · September 12, 2026
- MOSAIC: Query-Aware Exploration Policy Adaptation for GraphRAGarXiv cs.AI · September 12, 2026
- A Multi-Stage Rule-Chaining Framework for Compositional and Interpretable Cognitive ReasoningarXiv cs.AI · September 12, 2026