Same name. Fresh tasks. Every day.
Is your AI model getting worse?
StormEval gives each model fresh tasks of the same difficulty every day and compares it with its own past. A model that gets worse under the same name shows up here, with a date.
Baselines are still building. No verdicts yet.
- Models watched
- 5
- With a verdict
- 0 of 5
- Degraded now
- –
- Answers graded
- 270
02 · Station reportsEach model against its own past
- Stable
- Watch
- Changed
- Degraded
| Model | Status | Trend | Rating | Change | Pass rate | Days |
|---|---|---|---|---|---|---|
| Claude Opus 5.5Claude Code · effort mediumShow details | Building baseline, day 2 of 14 | 1773 | +55 vs baseline | 83% | 2 | |
| GPT-6 AstraCodex · effort mediumShow details | Building baseline, day 2 of 14 | 1628 | +103 vs baseline | 73% | 2 | |
| GPT-6.1 SolCodex · effort mediumShow details | Building baseline, day 2 of 14 | 1591 | −127 vs baseline | 70% | 2 | |
| Claude Sonnet 5.5Claude Code · effort highShow details | Building baseline, day 2 of 14 | 1466 | −91 vs baseline | 57% | 2 | |
| DeepSeek V4.1 FlashAPI · effort defaultShow details | Building baseline, day 1 of 14 | 1385 | baseline | 47% | 1 |
03 · ReportWhat every run took
Not a leaderboard. Models run at different effort settings and through different tools, and the difficulty mix was tuned to Claude Opus.
- Tasks graded
- 27032 timeouts · 0 not scored
- Time spent
- 14h 39mover 2 days
- Tokens used
- 5.39M1.78M in · 3.61M out
- Estimated cost
- $37.26at API list prices · 4 of 5 priced
Estimated cost is tokens at API list prices, not money spent: the Claude and Codex runs use plan usage. A dash means no list price is set for that model. Numbers cover every daily run on the current task mix.
| Model | Pass rate | Timeouts | Mean time | Median time | Output tokens / task | Thinking share | Output tokens / s | Est. cost / task | Est. cost / day |
|---|---|---|---|---|---|---|---|---|---|
| Claude Opus 5.5Claude Code · effort medium | 82%49/60 | 8 | 2m 21s | 1m 07s | 9.6k1.1k in | 93% | 101mean 103 | $0.197 | $5.12 |
| GPT-6 AstraCodex · effort medium | 68%41/60 | 5 | 3m 40s | 3m 21s | 5.0k15k in | 93% | 24mean 26 | $0.338 | $9.28 |
| GPT-6.1 SolCodex · effort medium | 75%45/60 | 7 | 3m 39s | 2m 59s | 5.3k16k in | 91% | 30mean 29 | – | – |
| Claude Sonnet 5.5Claude Code · effort high | 62%37/60 | 11 | 2m 47s | 1m 16s | 15k1.1k in | 96% | 129mean 138 | $0.150 | $3.68 |
| DeepSeek V4.1 FlashAPI · effort default | 47%14/30 | 1 | 4m 23s | 4m 20s | 63k289 in | 100% | 242mean 248 | $0.038 | $1.10 |
Time per task
How long each task took. A task with no answer in 8 minutes is a fail, and counts as 8 minutes of time spent.
By task family
Pass rate, mean time and mean output tokens for each kind of problem.
| Model | assign | orders | paths | queens | tsp |
|---|---|---|---|---|---|
| Claude Opus 5.5 | 67%4m 53s24k | 92%1m 27s9.0k | 83%1m 50s11k | 67%3m 06s4.3k | 100%29s2.7k |
| GPT-6 Astra | 50%3m 39s5.7k | 42%4m 06s5.1k | 100%3m 12s4.9k | 50%4m 15s4.6k | 100%3m 10s4.6k |
| GPT-6.1 Sol | 50%3m 37s5.7k | 100%2m 44s4.9k | 75%4m 41s6.8k | 58%4m 29s5.2k | 92%2m 47s4.1k |
| Claude Sonnet 5.5 | 42%6m 10s48k | 92%1m 55s15k | 83%1m 56s10k | 25%3m 31s13k | 67%26s3.2k |
| DeepSeek V4.1 Flash | 0%2m 07s32k | 67%5m 37s74k | 50%5m 25s72k | 33%4m 52s84k | 83%3m 51s55k |
Day by day
Claude Opus 5.5
2 days · time spent per day| Day | Passed | Timeouts | Time | Est. cost |
|---|---|---|---|---|
| 4 Oct | 25/30 | 4 | 1h 15m | $5.82 |
| 3 Oct | 24/30 | 4 | 1h 05m | $4.43 |
GPT-6 Astra
2 days · time spent per day| Day | Passed | Timeouts | Time | Est. cost |
|---|---|---|---|---|
| 4 Oct | 22/30 | 3 | 1h 51m | $8.91 |
| 3 Oct | 19/30 | 2 | 1h 49m | $9.65 |
GPT-6.1 Sol
2 days · time spent per day| Day | Passed | Timeouts | Time | Est. cost |
|---|---|---|---|---|
| 4 Oct | 21/30 | 3 | 1h 47m | – |
| 3 Oct | 24/30 | 4 | 1h 51m | – |
Claude Sonnet 5.5
2 days · time spent per day| Day | Passed | Timeouts | Time | Est. cost |
|---|---|---|---|---|
| 4 Oct | 17/30 | 5 | 1h 12m | $3.09 |
| 3 Oct | 20/30 | 6 | 1h 34m | $4.27 |
DeepSeek V4.1 Flash
1 day · time spent per day| Day | Passed | Timeouts | Time | Est. cost |
|---|---|---|---|---|
| 4 Oct | 14/30 | 1 | 2h 11m | $1.10 |
04 · MethodHow it works
- Fresh tasks every day.Each model gets 30 new exact-answer reasoning problems every day, generated at the difficulty it solves about half the time. No task is ever sent twice, so there is nothing to memorise.
- Graded by a program.An answer is right or wrong. No model judges another model.
- Rated like a chess player.Solving a problem counts as a win against that problem's difficulty. The rating comes with an error band. Movement inside the band means nothing.
- Compared with itself.The last 7 days are tested against the model's first 14 days, at the same difficulty.
- Two signals.Degraded means the pass rate fell. Changed means the model spends a different number of tokens on problems of the same difficulty while the pass rate held. Watch is an early warning: the pass rate dipped, but not enough to call.
How each number is computed
- Rating
- Fitted from passes and fails against each problem's difficulty, like a chess rating. The ± is one standard error.
- Pass rate
- Passes divided by passes plus fails. Errors and refusals are left out.
- Timeout
- No answer within 8 minutes. It is graded as a fail and counted as 8 minutes of time spent.
- Time per task
- Seconds from sending the problem to getting the answer. The mean and the median cover every task that answered or timed out.
- Tokens
- Input, output and thinking tokens as the provider reports them, for answered tasks. Thinking share is thinking tokens over output tokens.
- Output tokens per second
- A task's output tokens divided by its time, for answered tasks only. The table shows the median and the mean.
- Estimated cost
- Tokens multiplied by the API list price per million. It is not money spent: the Claude and Codex runs use plan usage. A model with no list price shows a dash.
- Task families
- Each day's 30 problems come from five kinds of search and counting puzzle, such as shortest round trips, grid paths and queen placements. The mix is the same every day.