Clicked Gallery

What are scaling laws?

Highlighted from a real engineering doc. Explained by Clicked.

Used in a sentence

Engineering Notes · AI Systems

The paper argued that scaling laws justified the lab's next training run, projecting performance before a single chip was rented.

The reader highlighted one word in the docs. Clicked explained the technical term “scaling laws” in simple terms:

Explained in three depths

Same facts, different vibe — Slang mode 😎

The Clicked way

●○○

Overview

Scaling laws are the observed pattern that AI models improve in a smooth, predictable way as three things grow together: their size, their data, and the computing power spent training them. The pattern is steady enough that labs train small cheap models, plot the curve, and forecast how good a far bigger one will be before spending the money. That predictability is what set off the race to build ever larger models.
●○○

Overview

Scaling laws are the reason AI labs buy chips like the factories might stop making them. Someone plotted model size against performance and got a curve so smooth you can rule where it goes next: make it bigger, it gets better, by roughly this much. No genius insight required at the next step, just more of everything. When improvement is that bookable in advance, the argument inside every lab collapses to one word: more. 😎

A quick take — often all you need.

●●○

Detail

Scaling laws describe how much better a model gets as more is spent on it. Make the model bigger, feed it more text, give it more computing time, and its errors shrink along a smooth curve, from small experimental models up to the largest ones built. That regularity is the valuable part. Training a frontier model costs hundreds of millions of dollars, and the laws turn the bet into a forecast: train a family of small models, measure the curve, and read off how good the expensive one should be. Two caveats keep the word "law" honest. Each equal step up in ability costs roughly ten times the spend of the last one, which is why budgets grew from racks of chips to entire data centres. And size, data and computing time only pay together: a huge model given too little text wastes most of its size, and finding the right balance among the three is its own branch of the science. The laws also predict a narrow thing, accuracy at guessing the next word. Which abilities arrive at which size, arithmetic, translation, code, is exactly what they stay silent about.
●●○

Detail

Scaling laws turned AI progress from gambling into shopping. Before them, nobody knew whether the next huge system would be smarter or merely pricier. Then researchers plotted performance against everything spent on a run, and the dots fell on a line so tidy you can extend it with a pencil. That pencil line is the whole business case: test a handful of cheap prototypes, extend the line, and read off what the nine-figure version will score before paying for it. The fine print comes in two parts. Each step along the line runs about ten times the price of the previous one, so the hardware bill went from a rack to a campus. And the mix matters as much as the total, because a colossal network without enough reading material squanders most of its capacity; finding the right recipe is half the craft. The line also tracks exactly one skill, predicting whatever comes next in a sentence. The headline talents, solving maths problems or working in a language nobody taught it on purpose, appear along the way unannounced. The line has always known how good the guessing gets. What the guessing unlocks, it finds out when we do. 😎

Want more? One click digs deeper.

●●●

Analogy

Scaling laws work like the height charts in a doctor's office. Measure a child at two, three and four, and the points trace a curve steady enough to predict their height at ten, without waiting six years to find out. Growth this regular makes forecasting possible, and that is the whole trick of it. The chart predicts height, and only height. Milestones follow age too, but loosely: the reference tables give ranges rather than dates, and nothing on the height curve narrows the range. Knowing a child tracks the 60th percentile says nothing about the week the bicycle finally stays upright.
●●●

Analogy

Scaling laws are a wartime codebreaking operation. Staff more analysts, tap more radio traffic, build more decoding machines, and the daily count of enemy messages read climbs steadily enough to plan a war around. The catch is the price of steady, because each bump in the count costs multiples of what the last one did. The inputs only pay together too: analysts without fresh intercepts just drink tea. The picture sharpens on schedule. The morning the enemy's cipher finally cracks open picks itself. 😎

Unfamiliar concept? A real-world example makes it click — fresh analogies on tap.

AI explanations may contain errors · Not professional advice

Formal definition — The same term, explained the usual way

Scaling laws are empirical power-law relationships between a model's test loss and the scale of its training, expressed in parameter count, dataset size and total compute. Formulated for neural language models by Kaplan et al. in 2020 and refined by the Chinchilla results of 2022, which showed loss is minimized when parameters and training tokens grow in roughly constant proportion, they permit extrapolation of large-model performance from smaller training runs and underpin compute-budget planning for frontier systems. The relationships concern aggregate loss rather than specific capabilities, and whether they persist at further scale remains an open empirical question.

Want Clicked to explain terms like “scaling laws” directly in your browser — including on PDFs?

Add to Chrome — Free

50 free Explanations · No credit card required