StormEval
  1. StormEval
  2. Models
  3. Kimi K3

Daily test results · Kimi Code · effort default

Is Kimi K3 getting worse?

Waiting for first run

No verdict yet: it has not had its first run.

Runs through
Kimi Code
Effort
default
Provider
kimi-cli
Model id
kimi-code/k3
Tasks
New issues in real code
Updated
2026-10-04 20:40 UTC

Rating over time

Higher is better. The band is one standard error either side. The dashed line is the first day. Movement inside the band means nothing.

Waiting for first run

Was Kimi K3 nerfed?

No verdict yet

Kimi K3 has not had its first run. There is no data, so there is nothing to say either way.

A verdict needs 14 days of this model's own results first, and then at least one newer day to compare with them.

Rating
–no runs yet
Change since day 1
–needs two days
Pass rate
–no runs yet
Days of data
0waiting for first run

What the runs took

Median time per task
–mean – · 0 timeouts
Output tokens per task
–– thinking · – in
Output tokens per second
–median · mean –
Estimated cost per task
–price not set

Estimated cost is tokens at API list prices, not money spent. A dash means no list price is set. Numbers cover 0 graded tasks over 1 day.

Time per task

How long each task took. A task not finished in 15 minutes is a fail, and counts as 15 minutes of time spent.

  • Kimi K3No runs yet

By task family

Pass rate, mean time and mean output tokens for each kind of problem.

Modelpull-request
Kimi K3–––

Day by day

DayRatingPassedTimeoutsTime spentMedian timeOutput tokensOutput tokens / sEst. cost
4 OctNo graded tasks0/000.0s–0––

How this is measured

Kimi K3 gets the same new tasks as every other model. Each task is an issue in a large, working codebase, and the model writes the change that resolves it with its own coding tool, as it would for a pull request. A program runs the tests to check the change. No task is ever sent twice to the same model, so there is nothing to memorise.

The model is compared with itself: the last 7 days against its first 14 days. A verdict is only given when the difference is too large to be chance. A task not finished in 15 minutes counts as a fail.

Read the full method

Not a leaderboard. Models run at different effort settings and through different tools.

Other models