StormEval
  1. StormEval
  2. Models
  3. Grok 4.7

Daily test results · Cursor · effort medium

Is Grok 4.7 getting worse?

Building baseline, day 1 of 14

No verdict yet: day 1 of a 14-day baseline.

Runs through
Cursor
Effort
medium
Provider
cursor-cli
Model id
grok-4.7-medium
Tasks
New issues in real code
Updated
2026-10-04 20:40 UTC

Rating over time

Higher is better. The band is one standard error either side. The dashed line is the first day. Movement inside the band means nothing.

Was Grok 4.7 nerfed?

No verdict yet

This is day 1 of the 14-day baseline for Grok 4.7. The baseline began on 4 October 2026. If the model runs every day, the baseline is complete on 17 October 2026 and the first verdict is expected on 18 October 2026.

It is too early to say either way. Until the baseline is complete there is nothing to compare a new day with, and a single day moves by chance.

Rating
1645±200 · one standard error
Change since day 1
–needs two days
Pass rate
100%latest day · 3 graded tasks
Days of data
1since 4 October 2026

What the runs took

Median time per task
8m 32smean 7m 21s · 0 timeouts
Output tokens per task
28k– thinking · 3.36M in
Output tokens per second
63median · mean 63
Estimated cost per task
–price not set

Estimated cost is tokens at API list prices, not money spent. A dash means no list price is set. Numbers cover 3 graded tasks over 1 day.

Time per task

How long each task took. A task not finished in 15 minutes is a fail, and counts as 15 minutes of time spent.

  • Grok 4.7median 8m 32sslowest 9m 47s · pull-request

By task family

Pass rate, mean time and mean output tokens for each kind of problem.

Modelpull-request
Grok 4.7100%7m 21s28k

Day by day

DayRatingPassedTimeoutsTime spentMedian timeOutput tokensOutput tokens / sEst. cost
4 Oct1645 ±2003/3022m 04s8m 32s83k63–

How this is measured

Grok 4.7 gets the same new tasks as every other model. Each task is an issue in a large, working codebase, and the model writes the change that resolves it with its own coding tool, as it would for a pull request. A program runs the tests to check the change. No task is ever sent twice to the same model, so there is nothing to memorise.

The model is compared with itself: the last 7 days against its first 14 days. A verdict is only given when the difference is too large to be chance. A task not finished in 15 minutes counts as a fail.

Read the full method

Not a leaderboard. Models run at different effort settings and through different tools.

Other models