Dashboard

What Is a Scaling Law in AI?

Scaling laws are curves fitted to hundreds of training runs. They are why the model you use costs what it costs, and why making models bigger stopped being the answer in 2022.

Cecilia Iona
Cecilia Iona
Senior Editor, AI & Product
4 September 20261 min read

A scaling law in AI is an empirical formula that predicts how much better a model gets when you give it more of something: more parameters, more training data, or more compute. The laws are not physics. They are curves fitted to hundreds of training runs, and they hold well enough that labs use them to decide how to spend nine-figure training budgets before writing a line of the training script.

The reason a non-researcher should care is narrower and more useful: scaling laws are why the model you use costs what it costs, and why "just make it bigger" stopped being the answer around 2022.

The basic shape

Train a family of models at different sizes on the same data and measure prediction error. Plot error against compute on log axes and you get something close to a straight line. Extend the line and you have a prediction for a model you have not trained yet.

Three quantities move together:

  • Parameters, the learned weights inside the model

  • Training tokens, how much text the model sees during training

  • Compute, roughly parameters multiplied by tokens

Scaling laws describe how to spend a fixed compute budget across the first two. That question has a right answer, and for several years the industry had the wrong one.

The Chinchilla correction

The influential early result, Scaling Laws for Neural Language Models from Kaplan and colleagues in 2020, suggested that when compute grows, parameters should grow faster than data. The field followed it. GPT-3 had 175 billion parameters and saw roughly 300 billion training tokens, under two tokens per parameter.

In 2022, Hoffmann and colleagues at DeepMind published Training Compute-Optimal Large Language Models, which trained over 400 models to test that prescription and found it wrong. For compute-optimal training, parameters and training tokens should scale at equal rates: double the model, double the data. The ratio they landed on was roughly 20 tokens per parameter, ten times more data per parameter than the prevailing practice.

They demonstrated it by training a 70 billion parameter model called Chinchilla on 1.4 trillion tokens. It outperformed models several times its size, including the 175 billion parameter GPT-3 and the 280 billion parameter Gopher.

That result reshaped the industry, and the effects are visible in what you use today.

Era

Prescription

Result

2020 to 2022

Grow parameters faster than data

Very large models, undertrained

2022 onward

Grow parameters and data together

Smaller models, far more data, cheaper to serve

Why this shows up in your bill

A model's parameter count drives its serving cost. Every token generated requires a pass through the weights, so a smaller model is cheaper and faster to run for every request, forever. Training cost is paid once. Inference cost is paid every time anyone uses the thing.

Chinchilla scaling said you could get the same quality from a smaller model by feeding it more data during training. That is an enormous economic win, because it moves cost from the recurring side of the ledger to the one-time side. It is a large part of why capable models became cheap enough to put in products, and it is what sits underneath the per-token price of a larger model.

It is also why small language models got good. A three billion parameter model trained on trillions of tokens is a very different object from a three billion parameter model from 2021, and the gap is scaling laws rather than architecture magic.

Where scaling laws stop being useful

Scaling laws predict one number: how well the model predicts the next token on held-out data. That number correlates with usefulness, imperfectly.

They say nothing about whether the model follows instructions, refuses appropriately, uses tools correctly, or holds a plan together across an hour of work. Those come from post-training, and post-training does not have a clean scaling law. This is why two models with similar training compute can feel completely different in practice, and why model parameter counts are a poor proxy for capability across vendors.

They also assume data you do not have. High-quality text is finite. Once a lab has used most of the usable internet, "just add more tokens" stops being an available move, which is why synthetic data, longer training on repeated data, and spending compute at inference time instead all became active areas. Test-time compute is in large part an answer to a scaling law running out of road.

What to take from it

If you build with AI rather than train models, the practical takeaway is one sentence: parameter count is not the number to shop on. A well-trained smaller model frequently beats a larger, undertrained one on your task, at a fraction of the price, and the only way to know is to test both on your own data.

That is the same conclusion the field reached by fitting curves to 400 training runs. You can reach it with an afternoon and an evaluation set. Our overview of how AI models work covers the surrounding machinery.

FAQ

Are scaling laws still holding?

The loss curves have continued to behave predictably, which is why labs keep spending against them. What has changed is that raw pretraining scale is no longer the only lever, and post-training plus inference-time compute now account for a large share of the capability gains between releases.

Does a scaling law tell you how smart a model will be?

No. It predicts prediction error on held-out text. Capabilities that matter to users, such as instruction following and reliable tool use, emerge unevenly and are not directly predicted by the curve.

Is 20 tokens per parameter still the right ratio?

It was the compute-optimal ratio for the training budget in that study. Labs routinely train well past it now, because overtraining a smaller model produces something more expensive to train and much cheaper to serve. When you serve billions of requests, that trade is worth making.

What is the difference between a scaling law and a benchmark?

A scaling law predicts a model's performance from its inputs before it exists. A benchmark measures a model that already exists on a fixed test. One is a forecast, the other is a measurement, and only the second can be gamed by training on the test set.

Do scaling laws apply to fine-tuning?

Loosely. Fine-tuning has its own curves relating data volume to task performance, and they flatten much faster. A few thousand well-chosen examples usually gets most of the available gain, which is a different regime from pretraining entirely.

How did this land?

About the author

Cecilia Iona
Cecilia Iona

Senior Editor, AI & Product

Cecilia leads the Swarmz editorial desk. She has spent a decade turning complex AI and product topics into writing people actually finish, and she owns the blog's quality bar.

Share

Get the next post in your inbox

One email a month. Product updates, engineering posts, and the best of Built with Swarmz.

I agree to receive emails about AI building tips and Swarmz product news. Unsubscribe any time.