All lessons
CHAPTER 11 · BUILD YOUR OWN EVALUATION

Read a benchmark without being fooled

Read a benchmark without being fooled

A benchmark is an argument, not a fact. Someone chose the task, the scoring and the conditions, and those choices decide the result as much as the model does. Learning to read one is the difference between "this model scored 92" and "this number means something for my work."

The five questions to ask any benchmark

Before you believe a number, ask:

  1. What exactly was measured? Speed (prefill, decode, time-to-first-token) and quality (correctness, task success) are different questions. A headline that mixes them is hiding one behind the other.
  2. What was held out? If the model could see the test data, the score is memorisation, not capability. The strongest benchmarks keep the scoring suite out of the model's reach, so it cannot be gamed or trained on.
  3. What varied? "Model A beats model B" is usually a claim about setups: quantisation, runtime, hardware, context and sampling all move at once. Change two things and you tested two things.
  4. How many runs? A single run is an anecdote. Local inference is noisy: thermal state, background load and caching all drift. Without repeats, a small difference is noise, not a result.
  5. What was reported, and what was left out? A benchmark that only shows the wins is marketing. The failures, the OOMs and the abandoned configurations are the useful part.

Beware the noise floor

Here is the mistake almost everyone makes. You run the same prompt twice, get 40 tok/s and then 44 tok/s, and you announce a 10% speedup. You changed nothing. You measured noise.

Before you call any difference real, you need a drift floor: run the identical setup twice, and see how far the numbers wander on their own. If identical runs drift 20%, then a 3% "win" is a sign on a coin flip. Write the floor down, and refuse to report anything under it. Refusing a result is a result.

The same discipline applies to scale: a cell whose spread across repeats is huge is not a datum, it is unusable. Report the median and the spread, or report nothing.

A worked example worth studying

One local-AI project, boxabirds/awesome-local-ai, does this well enough to learn from. Its installers ship tested model/hardware/stack pairings, and every published number was measured on real hardware. Two things stand out:

  • Its Vidi build benchmark runs a coding setup through a real spec (17 stories, 3 epics) one story at a time, and scores the result with a held-out acceptance suite the agent cannot read. The same spec and the same sandbox are used for the reference stack (a frontier cloud model) and for every local setup, so the comparison is like-for-like.
  • Its speed harnesses ship a drift floor: an explicit rule that a delta below ~20% between two identical runs is not a change, plus a spread threshold above which a measurement is thrown away rather than reported.

That is the posture to copy. Not the specific numbers (they belong to that author's hardware), but the honesty: measured vs extrapolated, held-out vs visible, real vs noise.

The output

Two habits, and you are ahead of most published numbers:

  1. Record your own baseline under identical conditions, so you know your noise floor. A speed claim without it is a guess.
  2. State the caveats with the number. "12% faster in one run, on the same machine, same quant" is honest. "12% faster" is not.

A benchmark you cannot qualify is a benchmark you cannot trust. The next lesson puts this to work: change one thing, run again, and compare without fooling yourself.