A Benchmark for LLM's Understanding of Middle School and High School Science Topics
arXiv cs.AIen
arXiv:2609.32020v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly integrated into educational settings, yet educators lack robust, standards-aligned tools to evaluate their effectiveness in K-12 science contexts. Existing benchmarks predominantly assess general language or advanced scientific reasoning, leaving a critical gap in understanding LLMs' performance on content directly relevant to secondary science curricula. To address this gap, we developed a comprehensive NGSS-aligned benchmark for both middle and high school science using a rigorous synthetic data pipeline, multi-judge validation, and item-level psychometric analysis. Nine open-weight LLMs were syst
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Verktyg
- Forskning
Related AI news
- SG startup Ropedia launches academic program for physical AITech in Asia · September 29, 2026
- Reuters: Anthropicin pörssilistautumisen tiedot julki – Yhtiö tekee jättitappiotaTivi · September 29, 2026
- Google appeals against EU order to share data, open Android to AI rivalsEconomic Times Tech · September 29, 2026
- AMD acquires ‘godmother of AI’ Li Fei-Fei’s start-up as battle with Nvidia intensifiesSCMP Tech · September 29, 2026
- Receiver-Conditioned Latent Communication gives 94% CacheBackarXiv cs.AI · September 29, 2026
- EngramRAG: Dynamic Usage-Weighted Topology and Synaptic Consolidation for Multi-Hop Agentic MemoryarXiv cs.AI · September 29, 2026