What the scores mean.
WHAT THEY ARE
Standardised tests measuring model performance on defined tasks.
WHAT THEY ARE USEFUL FOR
Rough comparison between models Tracking progress over time
WHAT THEY ARE NOT USEFUL FOR
Predicting performance on your specific task
WHY
A model scoring well on a reasoning benchmark may perform poorly on your writing, your documents, your domain.
THE CONTAMINATION PROBLEM
Benchmark questions appear on the internet, which is where models learn.
A model may have seen the test.
That inflates scores without indicating capability.
WHAT VENDORS DO
Report benchmarks favourable to their model.
Every model is the best at something.
WHAT TO DO INSTEAD
Test on your own work.
Build a small set of representative tasks and compare models on those.
THAT IS THE ONLY RELIABLE COMPARISON
Twenty examples from your actual work tells you more than any published score.