StormEval Blog

Updated 2026-10-04 05:45 UTC

Same name. Fresh tasks. Every day.

Is your AI model getting worse?

StormEval gives each model fresh tasks of the same difficulty every day and compares it with its own past. A model that gets worse under the same name shows up here, with a date.

Baselines are still building. No verdicts yet.

Models watched
5
With a verdict
0 of 5
Degraded now
–
Answers graded
270

02 · Station reportsEach model against its own past

  • Stable
  • Watch
  • Changed
  • Degraded
ModelStatusTrendRatingChangePass rateDays
Claude Opus 5.5Claude Code · effort mediumShow details Building baseline, day 2 of 14 1773 +55 vs baseline 83% 2
GPT-6 AstraCodex · effort mediumShow details Building baseline, day 2 of 14 1628 +103 vs baseline 73% 2
GPT-6.1 SolCodex · effort mediumShow details Building baseline, day 2 of 14 1591 −127 vs baseline 70% 2
Claude Sonnet 5.5Claude Code · effort highShow details Building baseline, day 2 of 14 1466 −91 vs baseline 57% 2
DeepSeek V4.1 FlashAPI · effort defaultShow details Building baseline, day 1 of 14 1385 baseline 47% 1

03 · ReportWhat every run took

Not a leaderboard. Models run at different effort settings and through different tools, and the difficulty mix was tuned to Claude Opus.

Tasks graded
27032 timeouts · 0 not scored
Time spent
14h 39mover 2 days
Tokens used
5.39M1.78M in · 3.61M out
Estimated cost
$37.26at API list prices · 4 of 5 priced

Estimated cost is tokens at API list prices, not money spent: the Claude and Codex runs use plan usage. A dash means no list price is set for that model. Numbers cover every daily run on the current task mix.

ModelPass rateTimeoutsMean timeMedian timeOutput tokens / taskThinking shareOutput tokens / sEst. cost / taskEst. cost / day
Claude Opus 5.5Claude Code · effort medium 82%49/60 8 2m 21s 1m 07s 9.6k1.1k in 93% 101mean 103 $0.197 $5.12
GPT-6 AstraCodex · effort medium 68%41/60 5 3m 40s 3m 21s 5.0k15k in 93% 24mean 26 $0.338 $9.28
GPT-6.1 SolCodex · effort medium 75%45/60 7 3m 39s 2m 59s 5.3k16k in 91% 30mean 29 – –
Claude Sonnet 5.5Claude Code · effort high 62%37/60 11 2m 47s 1m 16s 15k1.1k in 96% 129mean 138 $0.150 $3.68
DeepSeek V4.1 FlashAPI · effort default 47%14/30 1 4m 23s 4m 20s 63k289 in 100% 242mean 248 $0.038 $1.10

Time per task

How long each task took. A task with no answer in 8 minutes is a fail, and counts as 8 minutes of time spent.

  • Claude Opus 5.5median 1m 07sslowest timed out · assign L10
  • GPT-6 Astramedian 3m 21sslowest timed out · queens L9
  • GPT-6.1 Solmedian 2m 59sslowest timed out · assign L12
  • Claude Sonnet 5.5median 1m 16sslowest timed out · assign L10
  • DeepSeek V4.1 Flashmedian 4m 20sslowest timed out · orders L14

By task family

Pass rate, mean time and mean output tokens for each kind of problem.

Modelassignorderspathsqueenstsp
Claude Opus 5.567%4m 53s24k92%1m 27s9.0k83%1m 50s11k67%3m 06s4.3k100%29s2.7k
GPT-6 Astra50%3m 39s5.7k42%4m 06s5.1k100%3m 12s4.9k50%4m 15s4.6k100%3m 10s4.6k
GPT-6.1 Sol50%3m 37s5.7k100%2m 44s4.9k75%4m 41s6.8k58%4m 29s5.2k92%2m 47s4.1k
Claude Sonnet 5.542%6m 10s48k92%1m 55s15k83%1m 56s10k25%3m 31s13k67%26s3.2k
DeepSeek V4.1 Flash0%2m 07s32k67%5m 37s74k50%5m 25s72k33%4m 52s84k83%3m 51s55k

Day by day

Claude Opus 5.5

2 days · time spent per day
DayPassedTimeoutsTimeEst. cost
4 Oct25/3041h 15m$5.82
3 Oct24/3041h 05m$4.43

GPT-6 Astra

2 days · time spent per day
DayPassedTimeoutsTimeEst. cost
4 Oct22/3031h 51m$8.91
3 Oct19/3021h 49m$9.65

GPT-6.1 Sol

2 days · time spent per day
DayPassedTimeoutsTimeEst. cost
4 Oct21/3031h 47m–
3 Oct24/3041h 51m–

Claude Sonnet 5.5

2 days · time spent per day
DayPassedTimeoutsTimeEst. cost
4 Oct17/3051h 12m$3.09
3 Oct20/3061h 34m$4.27

DeepSeek V4.1 Flash

1 day · time spent per day
DayPassedTimeoutsTimeEst. cost
4 Oct14/3012h 11m$1.10

04 · MethodHow it works

  1. Fresh tasks every day.Each model gets 30 new exact-answer reasoning problems every day, generated at the difficulty it solves about half the time. No task is ever sent twice, so there is nothing to memorise.
  2. Graded by a program.An answer is right or wrong. No model judges another model.
  3. Rated like a chess player.Solving a problem counts as a win against that problem's difficulty. The rating comes with an error band. Movement inside the band means nothing.
  4. Compared with itself.The last 7 days are tested against the model's first 14 days, at the same difficulty.
  5. Two signals.Degraded means the pass rate fell. Changed means the model spends a different number of tokens on problems of the same difficulty while the pass rate held. Watch is an early warning: the pass rate dipped, but not enough to call.

How each number is computed

Rating
Fitted from passes and fails against each problem's difficulty, like a chess rating. The ± is one standard error.
Pass rate
Passes divided by passes plus fails. Errors and refusals are left out.
Timeout
No answer within 8 minutes. It is graded as a fail and counted as 8 minutes of time spent.
Time per task
Seconds from sending the problem to getting the answer. The mean and the median cover every task that answered or timed out.
Tokens
Input, output and thinking tokens as the provider reports them, for answered tasks. Thinking share is thinking tokens over output tokens.
Output tokens per second
A task's output tokens divided by its time, for answered tasks only. The table shows the median and the mean.
Estimated cost
Tokens multiplied by the API list price per million. It is not money spent: the Claude and Codex runs use plan usage. A model with no list price shows a dash.
Task families
Each day's 30 problems come from five kinds of search and counting puzzle, such as shortest round trips, grid paths and queen placements. The mix is the same every day.