CARAT: Do Materials LLMs Reason or Recite?
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.38340v1 Announce Type: new Abstract: When a materials LLM answers a question about crystal structure, does it reason from the structure or copy an answer already printed in its input? Accuracy cannot tell: a structural description often prints the very field it is scored against. CARAT holds question and gold answer fixed across eight matched views, names each structural relation separately in GraphSpace, and adds matched fine-tuning, answer masking, evidence injection, paired inference, and a rule that can withhold claims. First, on the benchmark's hardest families the grounded view is worth 17.3 points over formula inputs. Second, we turn that scrutiny on ourselves. GraphSpace b
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Kaikki noudattivat ohjeita – Kukaan ei ollut vastuussaTivi · October 1, 2026
- China's chip-tool localization accelerates: ACM Research backlog jumps 88%DIGITIMES · October 1, 2026
- Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization LimitsarXiv cs.AI · October 1, 2026
- ChartRevise: A Dataset and Evaluation Protocol for Exact Chart Editing via CodearXiv cs.AI · October 1, 2026
- Can an AI Agent Rediscover a Blaschke-Curve Invariant?arXiv cs.AI · October 1, 2026
- SimTrace: Grounded Multimodal User Trajectories Generation for Online User ModelingarXiv cs.AI · October 1, 2026