How StormEval tells whether an AI model got worse
Fresh tasks every day, a fixed difficulty mix, a program as the grader, and each model compared with its own past. What the test does, and what it cannot tell you.
People often feel that a model got worse while its name stayed the same. A feeling is hard to check. StormEval turns the question into a measurement that can be repeated every day. This post explains how it works and where it stops.
The question
For each model we ask one thing: is it worse now than it was before?
We do not ask which model is best. Every model is compared with its own past and with nothing else.
Fresh tasks every day
Each model gets 30 new problems every day. No problem is ever sent twice, to any model.
The reason is simple. A task that repeats can be learned. If the same 30 prompts arrived every day, a vendor could cache them or tune on them, and the score would hold up while the model got worse everywhere else. A task nobody has seen before cannot be prepared for.
Each task is generated from a seed made of the day, the slot and the model's surface. Before a task is sent, we check that the same prompt was never sent before.
A fixed difficulty mix
The tasks change. The difficulty does not.
A day's set is a list of slots. A slot is a task family at a difficulty level. The families are search and counting puzzles with one whole-number answer: the shortest round trip through a set of cities, the number of ways to give each worker one job, the number of task orders that obey a set of rules, the number of paths through a grid with walls, and the number of ways to place queens on a board with blocked squares. The list of slots is frozen, so every day has the same number of problems of each kind at each level.
The levels were picked near the point where a strong model solves about half of them. A problem a model always solves, or never solves, says nothing about change.
Graded by a program
Every answer is checked by a program against the one correct answer. It is right or it is wrong.
No model judges another model. A judge that drifts would look like every model drifting at once.
A rating with an error band
Passes and fails become a rating, in the way a chess rating is built from wins and losses. Solving a hard problem counts for more than solving an easy one.
Thirty problems a day is a small sample, so the rating comes with an error band of one standard error. Movement inside the band means nothing. This is the most important thing to keep in mind when you look at a single day.
Compared with its own past
The first 14 days of a model's results are its baseline. Until those 14 days exist, and at least one newer day after them, there is no verdict. The page says "building baseline" and shows which day it is.
After that, the most recent days, up to 7 of them, are compared with the baseline. The comparison is made slot by slot: each family and level is compared with the same family and level. A hard slot is never compared with an easy one.
Timeouts count as fails
A run that gives no answer within 8 minutes is graded as a fail. The model did not solve the task in the time allowed. Leaving such runs out would raise the pass rate of exactly the model that struggles.
Errors on our side, such as a network failure, are different. They are retried, and if they still fail they are left out of the pass rate.
What the statuses mean
| Status | What it means |
|---|---|
| Stable | The recent pass rate is in line with the baseline. |
| Watch | The pass rate dipped. It is an early warning, not yet enough to call. |
| Degraded | The pass rate fell by more than chance explains. |
| Changed | The pass rate held, but the model now spends a clearly different number of output tokens on problems of the same difficulty. |
The line between them is a z-score: how far the recent result is from what the baseline predicts, measured in standard deviations. Watch starts at 2 below. Degraded starts at 3 below. Changed needs the token shift to be at least 3 in either direction.
Degraded and Changed are kept apart on purpose. A model can change how it works without getting worse.
What it cannot tell you
- It is not a leaderboard. Models run at different effort settings and through different tools, and the difficulty mix was tuned to one model. A rating is only comparable with the same model's earlier ratings.
- It cannot say why. A drop shows that our results changed. It does not show what a vendor did or intended.
- The puzzles are narrow. They are search and counting problems with one exact answer. They are not coding, writing or long conversations. A model can get worse at something these puzzles do not touch.
- It measures one way of running a model. A model reached through a coding tool and the same model reached through an API are different surfaces and can behave differently.
- Small drops take time. With 30 problems a day, a large drop shows up within days. A small one needs weeks, or is not seen at all.
The method is the same every day, and every number on the site comes from it. See today's results or the method summary.