Daily test results · Codex · effort medium
Is GPT-6 Astra getting worse?
No verdict yet: day 2 of a 14-day baseline.
- Runs through
- Codex
- Effort
- medium
- Provider
- codex-cli
- Model id
gpt-6-astra- Tasks
- 30 new ones every day
- Updated
- 2026-10-04 05:45 UTC
Rating over time
Higher is better. The band is one standard error either side. The dashed line is the first day. Movement inside the band means nothing.
Was GPT-6 Astra nerfed?
No verdict yet
This is day 2 of the 14-day baseline for GPT-6 Astra. The baseline began on 3 October 2026. If the model runs every day, the baseline is complete on 16 October 2026 and the first verdict is expected on 17 October 2026.
It is too early to say either way. Until the baseline is complete there is nothing to compare a new day with, and a single day moves by chance.
- Rating
- 1628±83 · one standard error
- Change since day 1
- +103from 1525 on 3 October 2026
- Pass rate
- 73%latest day · 30 graded tasks
- Days of data
- 2since 3 October 2026
What the runs took
- Median time per task
- 3m 21smean 3m 40s · 5 timeouts
- Output tokens per task
- 5.0k93% thinking · 15k in
- Output tokens per second
- 24median · mean 26
- Estimated cost per task
- $0.338at API list prices · $9.28 a day
Estimated cost is tokens at API list prices, not money spent. A dash means no list price is set. Numbers cover 60 graded tasks over 2 days.
Time per task
How long each task took. A task with no answer in 8 minutes is a fail, and counts as 8 minutes of time spent.
By task family
Pass rate, mean time and mean output tokens for each kind of problem.
| Model | assign | orders | paths | queens | tsp |
|---|---|---|---|---|---|
| GPT-6 Astra | 50%3m 39s5.7k | 42%4m 06s5.1k | 100%3m 12s4.9k | 50%4m 15s4.6k | 100%3m 10s4.6k |
Day by day
| Day | Rating | Passed | Timeouts | Time spent | Median time | Output tokens | Output tokens / s | Est. cost |
|---|---|---|---|---|---|---|---|---|
| 4 Oct | 1628 ±83 | 22/30 | 3 | 1h 51m | 3m 22s | 132k | 24 | $8.91 |
| 3 Oct | 1525 ±73 | 19/30 | 2 | 1h 49m | 3m 16s | 143k | 24 | $9.65 |
How this is measured
GPT-6 Astra gets 30 new exact-answer reasoning problems every day, at a fixed difficulty mix. No problem is ever sent twice, so there is nothing to memorise. A program grades each answer as right or wrong.
The model is compared with itself: the last 7 days against its first 14 days, at the same difficulty. A verdict is only given when the difference is too large to be chance. A task with no answer in 8 minutes counts as a fail.
Not a leaderboard. Models run at different effort settings and through different tools, and the difficulty mix was tuned to Claude Opus.