SciLitBench: Benchmark and Design Principles for LLM-Powered Systematic Literature Reviews
arXiv cs.AIen
arXiv:2609.05505v1 Announce Type: new Abstract: Systematic reviews require sustained human judgment across thousands of records, yet existing evaluations of large language models (LLMs) typically examine review stages in isolation. We introduce SciLitBench, a multi-stage benchmark spanning title and abstract screening, full-text screening, and schema-guided data extraction, with 42,981 retrieved records, 1,012 full texts, and annotations for 888 included papers. Across 22 open-weight LLMs from six model families, explicit inclusion and exclusion criteria improve title and abstract screening $F_2$ by 28.8\%, while researcher-authored rationales improve full-text screening by 15\%. Data extrac
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
- Företag
Related AI news
- China's DeepSeek taps CITIC Securities for domestic IPOEconomic Times Tech · September 9, 2026
- Oracle plans HPE networking rollout for AI data centersTech in Asia · September 9, 2026
- Sources: China Securities Regulatory Commission is informally tightening IPO approvals for humanoid startups after a volatile debut by industry leader Unitree (The Information)Techmeme · September 9, 2026
- Planning and Scheduling Business Processes under Control-Flow UncertaintyarXiv cs.AI · September 9, 2026
- Agents Trust Tools Too Much: Measuring Reliance on Unreliable ToolsarXiv cs.AI · September 9, 2026
- Distilling Vision-Language Models for On-Device Fire UnderstandingarXiv cs.AI · September 9, 2026