Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.22161v1 Announce Type: new Abstract: Medical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how these data types differentially shape model capabilities remains unclear. We address this issue with token-matched experiments that vary the didactic-to-clinical ratio and analyze how data composition affects performance, capability profiles, and error patterns across knowledge-intensive and clinic-oriented tasks. We uncover an asymmetric transfer across task types: clinical data improves clinic-oriented tasks while remaining competitive on knowledge-intensive ones, whereas didactic data mainly imp
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Do Existing Preconditioners Improve Biomedical Tabular Foundation Learning? An Empirical Study on TabPFN OptimizationarXiv cs.AI · September 23, 2026
- 4DGS-JEPA: Temporally Compositional Joint-Embedding Prediction for Dynamic Gaussian SplattingarXiv cs.AI · September 23, 2026
- An Accurate and Interpretable Hyper Graph Neural Network for GBM Survival PredictionarXiv cs.AI · September 23, 2026
- Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal EmbeddingsarXiv cs.AI · September 23, 2026
- X-Planner: Event-Structured Task Planning for Embodied IntelligencearXiv cs.AI · September 23, 2026
- Lean Pool: An AI-Maintained Archive of Formalized MathematicsarXiv cs.AI · September 23, 2026