Skip to content
StrataHub

Moments · 2020

Scaling Laws

A 2020 paper found that intelligence follows a power law, and you could budget for it.

In January 2020, Jared Kaplan, Sam McCandlish, and colleagues at OpenAI published 'Scaling Laws for Neural Language Models'. The paper reported something almost suspiciously tidy: a language model's loss falls as a smooth power law as you increase model size, dataset size, or training compute, holding across many orders of magnitude.

Just as striking was what did not matter much. Architectural details like network width versus depth had only modest effects compared to raw scale. The finding recast progress in language AI as, to a first approximation, a resource allocation problem: decide your compute budget, and the laws tell you how big a model to train and roughly how good it will be.

That predictability changed how AI was funded and built. Training a frontier model costs enormous sums, and scaling laws let organizations forecast returns before spending. GPT-3, released later in 2020 with 175 billion parameters, was in large part a bet that the curves would keep holding. They did.

In 2022, DeepMind's Chinchilla paper by Jordan Hoffmann and colleagues delivered an important correction: the field had been building models too large for the data they were fed. Compute-optimal training wanted roughly twenty tokens of data per parameter, and the 70-billion-parameter Chinchilla, trained on far more data, outperformed models several times its size.

Scaling laws turned a research field into an industrial discipline and gave empirical teeth to Richard Sutton's bitter lesson, the observation that general methods riding more computation beat human-crafted cleverness. The debates they opened, about data running out and about whether the curves eventually bend, still shape frontier AI strategy today.

From history to production

We turn these ideas into working systems

The same techniques, shipped into your stack with evals, observability, and measurable ROI.