Scaling Laws for Mixture Pretraining Under Data Constraints

Apple Machine Learningen

Apple Machine Learning

AI Global Wire

As language models scale, the amount of data they require grows – yet many target data sources, such as low-resource languages or specialized domains, are inherently limited in size. A common strategy is to mix this scarce but valuable target data with abundant generic data, which presents a fundamental trade-off: too little target data in the mixture underexposes the model to the target domain, while too much target data repeats the same examples excessively, yielding diminishing returns and eventual overfitting. We study this trade-off across more than 2,000 language-model training runs…

This is a short summary published by AI Global Wire. The full article is owned and hosted by Apple Machine Learning — open it there to read it in full.

Read the full story at Apple Machine Learning
  • Forskning
  • Reglering

Related AI news