Clicked Gallery

What are AI Benchmarks?

Highlighted from a real engineering doc. Explained by Clicked.

Used in a sentence

Engineering Notes · AI Systems

The company led with benchmark scores, though independent testers reported a far narrower gap.

The reader highlighted one word in the docs. Clicked broke down the technical term “benchmark” into plain English:

Explained in three depths

Same facts, different vibe — Slang mode 😎

The Clicked way

●○○

Overview

AI benchmarks are set lists of questions that every model is asked, so the scores can be compared like for like. They are where the numbers in every launch announcement come from. A model can top the table and still disappoint you, which is why who ran the test matters as much as the score.
●○○

Overview

A benchmark is one fixed list of questions, and every model gets asked that exact same list, which is the only thing making launch-day scores comparable. The snag: most of those lists sat on the open internet where the models could read them first. So the winner might have recognised the paper rather than worked anything out. 😎

A quick take — often all you need.

●●○

Detail

You could just try a model yourself, but you would be testing whatever questions occurred to you that afternoon, with no way to tell whether it beat a rival or simply had an easy run. A benchmark is one specific list of questions with known answers, and every model is asked that identical list, which is the only reason scores can be compared at all. Most such lists were made public so anyone could check the results, and that is where the trouble starts. Models are trained on the open internet, so a model may well have read both the questions and the answers before it ever sat the test. The field calls this contamination, and it means a high score can simply mean the model recognised the paper. A second problem follows from competition itself, because once everyone chases one number, teams tune for the number rather than the skill beneath it. So the more trusted benchmarks are now held by independent organisations that keep the list private, put each model through it themselves, and release only the marks. Every model still faces identical questions, and nobody can revise in advance.
●●○

Detail

Testing a model yourself tells you how it handled the six things you thought to ask on a Tuesday. Fine for you, useless for comparing anything to anything. So the industry standardised: one fixed list of questions, the same list for everybody, marks you can actually line up side by side. Then everybody put their lists online, and the models read what is online, and you can see precisely where this is heading. Half a leaderboard might merely recognise the paper, which the field calls contamination because calling it cheating would be impolite. Worse, when a single number is the whole competition, teams start polishing the number instead of the thing the number was supposed to measure. The unglamorous fix is outside referees who keep their questions locked in a drawer, sit every model down themselves, and release nothing but the marks. 😎

Want more? One click digs deeper.

●●●

Analogy

A restaurant guide's star rating. It is a real assessment, and no kitchen collects three stars by accident, but it reflects one set of judges applying one set of standards on the nights they happened to visit. Kitchens also work out what those judges reward and cook toward it, so the stars climb while the food stays exactly where it was. The rating is honest, and it still cannot tell you whether you will enjoy dinner there on a wet Tuesday.
●●●

Analogy

A school where every past paper from the last decade is online. Some pupils learn the subject and some memorise the papers, and on results day both of them collect an A. Nothing in those marks tells you which pupil is which, however hard you squint at them. Set a fresh paper nobody has seen and the two groups stop looking remotely alike.

Unfamiliar concept? A real-world example makes it click — fresh analogies on tap.

AI explanations may contain errors · Not professional advice

Formal definition — The same term, explained the usual way

AI benchmarks are standardized evaluation suites of tasks with known ground truth, administered identically across models to permit comparative measurement of capability. Their validity is limited by training-data contamination, where benchmark items appear in pretraining corpora, and by optimization pressure toward the metric itself; mitigations include held-out and private test sets, and independent third-party evaluation.

Want Clicked to explain terms like “benchmark” directly in your browser — including on PDFs?

Add to Chrome — Free

50 free Explanations · No credit card required