<?xml version="1.0" encoding="UTF-8"?>
<feed xmlns="http://www.w3.org/2005/Atom">
  <id>https://stormeval.com/blog</id>
  <title>StormEval blog</title>
  <subtitle>Notes from StormEval on how the daily tests work and what they can and cannot tell you about an AI model.</subtitle>
  <link rel="self" type="application/atom+xml" href="https://stormeval.com/feed.xml"/>
  <link rel="alternate" type="text/html" href="https://stormeval.com/blog"/>
  <updated>2026-10-04T00:00:00Z</updated>
  <author><name>StormEval</name></author>
  <entry>
    <id>https://stormeval.com/blog/how-stormeval-tells-whether-an-ai-model-got-worse</id>
    <title>How StormEval tells whether an AI model got worse</title>
    <link rel="alternate" type="text/html" href="https://stormeval.com/blog/how-stormeval-tells-whether-an-ai-model-got-worse"/>
    <published>2026-10-04T00:00:00Z</published>
    <updated>2026-10-04T00:00:00Z</updated>
    <summary>Fresh tasks every day, a fixed difficulty mix, a program as the grader, and each model compared with its own past. What the test does, and what it cannot tell you.</summary>
  </entry>
</feed>
