StormEval
  1. StormEval
  2. Models
  3. Claude Opus 5.5

Daily test results · Claude Code · effort medium

Is Claude Opus 5.5 getting worse?

Building baseline, day 2 of 14

No verdict yet: day 2 of a 14-day baseline.

Runs through
Claude Code
Effort
medium
Provider
claude-cli
Model id
claude-opus-5-5
Tasks
30 new ones every day
Updated
2026-10-04 05:45 UTC

Rating over time

Higher is better. The band is one standard error either side. The dashed line is the first day. Movement inside the band means nothing.

Was Claude Opus 5.5 nerfed?

No verdict yet

This is day 2 of the 14-day baseline for Claude Opus 5.5. The baseline began on 3 October 2026. If the model runs every day, the baseline is complete on 16 October 2026 and the first verdict is expected on 17 October 2026.

It is too early to say either way. Until the baseline is complete there is nothing to compare a new day with, and a single day moves by chance.

Rating
1773±102 · one standard error
Change since day 1
+55from 1718 on 3 October 2026
Pass rate
83%latest day · 30 graded tasks
Days of data
2since 3 October 2026

What the runs took

Median time per task
1m 07smean 2m 21s · 8 timeouts
Output tokens per task
9.6k93% thinking · 1.1k in
Output tokens per second
101median · mean 103
Estimated cost per task
$0.197at API list prices · $5.12 a day

Estimated cost is tokens at API list prices, not money spent. A dash means no list price is set. Numbers cover 60 graded tasks over 2 days.

Time per task

How long each task took. A task with no answer in 8 minutes is a fail, and counts as 8 minutes of time spent.

  • Claude Opus 5.5median 1m 07sslowest timed out · assign L10

By task family

Pass rate, mean time and mean output tokens for each kind of problem.

Modelassignorderspathsqueenstsp
Claude Opus 5.567%4m 53s24k92%1m 27s9.0k83%1m 50s11k67%3m 06s4.3k100%29s2.7k

Day by day

DayRatingPassedTimeoutsTime spentMedian timeOutput tokensOutput tokens / sEst. cost
4 Oct1773 ±10225/3041h 15m1m 16s285k102$5.82
3 Oct1718 ±9424/3041h 05m1m 07s216k101$4.43

How this is measured

Claude Opus 5.5 gets 30 new exact-answer reasoning problems every day, at a fixed difficulty mix. No problem is ever sent twice, so there is nothing to memorise. A program grades each answer as right or wrong.

The model is compared with itself: the last 7 days against its first 14 days, at the same difficulty. A verdict is only given when the difference is too large to be chance. A task with no answer in 8 minutes counts as a fail.

Read the full method

Not a leaderboard. Models run at different effort settings and through different tools, and the difficulty mix was tuned to Claude Opus.

Other models