QVAC Genesis III: A Large-Scale, High-Quality Open Synthetic STEM Corpus for Efficient Language Model Pre-Training
arXiv cs.AIen
arXiv cs.AI
AI Global WirearXiv:2609.19513v1 Announce Type: new Abstract: High-quality pre-training data is a critical bottleneck for educational and STEM-specific language models targeting edge AI and on-device deployment where token budgets are tightly constrained. While major organizations train ever-larger models on private corpora, the open ecosystem lacks STEM-focused synthetic datasets that deliver high per-token learning value efficiently for small models. To address this gap, we introduce QVAC Genesis III, a 191.43B-token, STEM-focused multi-domain synthetic corpus covering 19 domains across several difficulty levels and different educational styles. QVAC Genesis III is built via a dual generation strategy t
This is a short summary published by AI Global Wire. The full article is owned and hosted by arXiv cs.AI — open it there to read it in full.
Read the full story at arXiv cs.AI- Forskning
Related AI news
- Security researchers in an OpenAI bug bounty program hacked OpenAI, accessing its "monorepo" on GitHub, using a cybersecurity version of Opus 4.8 and Opus 5 (Robert McMillan/Wall Street Journal)Techmeme · September 18, 2026
- What Do Current Systematic Generalization Tasks Miss? A Reasoning-Centered AnalysisarXiv cs.AI · September 18, 2026
- Do AI Agents Understand Computer Architecture?arXiv cs.AI · September 18, 2026
- Self Improvement via Fast Tree-searcharXiv cs.AI · September 18, 2026
- When Hiring Becomes Agent-Mediated: Evaluating Access and Recurrence in Two-Agent R\'esum\'e ScreeningarXiv cs.AI · September 18, 2026
- Reach or Solve? Attributing Agentic RL Gains with Checkpoint HandoffsarXiv cs.AI · September 18, 2026