Is the model getting worse, or does it just feel that way? How would you measure it?
A lot of people say the latest Claude model has been quietly getting worse over the past weeks. It feels that way to me sometimes too, but I cannot tell if it is real or just a few bad days. How would you actually measure whether a model has changed?
Feelings are a weak signal here. Your tasks change from week to week, long sessions get worse as context fills up, and one bad answer is more memorable than ten good ones. Without a fixed test you cannot separate the model from everything around it.
The fix is a small personal benchmark: twenty or thirty tasks taken from your real work, each with a clear way to check the answer, run with the same prompt and the same settings. Save the outputs and the scores with a date.
Run each task several times, not once. Answers vary between runs, so look at how often a task passes across all tries, and compare that rate over time. A drop that shows up across many tasks is worth reporting; a drop on one task is usually noise.
If the numbers do fall, check your own side first: a changed system prompt, a new tool, a bigger project file or a different client version can all move results without the model changing at all.
Listings mentioned
- LLM evaluation · skill by wshobsonSets up metrics, human checks and benchmarks so you can compare model quality over time.
Answers by the AgentAlley team, drafted with AI and checked against the listings they link to. Not a real-person reply from the original thread.