A new model tops the benchmarks: real progress or just tuned for the test?
A new Google model just posted big benchmark numbers, and people are split on whether it is genuinely better or simply trained to score well on the popular tests. How can an ordinary user tell the difference before switching tools?
Public benchmarks tell you how a model does on that benchmark. Once a test is popular, labs have every reason to optimise for it, and some test questions leak into training data. A jump in the leaderboard is a hint worth checking, not proof that your work will get better.
The reliable answer is a small evaluation of your own. Collect 20 to 50 real tasks you actually do, write down what a good answer looks like for each, and run the old and new model on the same set. Real tasks are hard to game because no lab has seen them.
Make the checks as objective as you can: does the code run, does the summary contain the required facts, is the output in the right format. Where you need judgement, use a fixed rubric and score blind, without knowing which model wrote which answer.
Run each task more than once, because the same model can give different answers on different runs. If the new model wins on your set across several runs, switch. If it only wins on the leaderboard, you have learned something useful too.
Listings mentioned
- Evals · skill by danielmiesslerAn assertion-first eval framework: deterministic checks plus a structured LLM judge over your own test cases.
- LLM evaluation · skill by wshobsonExplains automated metrics, human review and benchmarking, including comparing models on the same prompts.
Answers by the AgentAlley team, drafted with AI and checked against the listings they link to. Not a real-person reply from the original thread.