<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0">
<channel>
  <title>assert(llm) wire</title>
  <link>https://assertllm.com/wire/</link>
  <description>Short notes on AI and LLM testing</description>
  <language>en</language>
  <lastBuildDate>Sun, 11 Oct 2026 00:00:00 GMT</lastBuildDate>
  <item>
    <title>assert(llm) now reports Wilson 95% confidence intervals</title>
    <link>https://assertllm.com/wire/2026-10-11-wilson-confidence-intervals/</link>
    <guid isPermaLink="true">https://assertllm.com/wire/2026-10-11-wilson-confidence-intervals/</guid>
    <pubDate>Sun, 11 Oct 2026 00:00:00 GMT</pubDate>
    <description>&lt;p&gt;Every result card, the live grid and the markdown export now show a Wilson 95% interval next to the pass rate. With 15 questions, 14 correct is not “93.3%”: it is somewhere between roughly 70% and 99%. When a model’s interval overlaps with the leader’s, the card says &lt;code&gt;≈ statistical tie&lt;/code&gt; instead of crowning a winner. Wilson is used over the textbook normal approximation because it stays sensible at small n and near 0% or 100%, which is exactly where small eval sets live.&lt;/p&gt;</description>
  </item>
  <item>
    <title>OpenAI Evals goes read-only on 31 October</title>
    <link>https://assertllm.com/wire/2026-10-09-openai-evals-read-only/</link>
    <guid isPermaLink="true">https://assertllm.com/wire/2026-10-09-openai-evals-read-only/</guid>
    <pubDate>Fri, 09 Oct 2026 00:00:00 GMT</pubDate>
    <description>&lt;p&gt;OpenAI’s hosted Evals turns read-only on 31 October 2026 and shuts down on 30 November 2026. If your regression suite lives there, export the eval definitions and the last good run before the read-only date, while you can still re-run them for a baseline. Keep golden items in plain JSON or CSV with an &lt;code&gt;id&lt;/code&gt;, a &lt;code&gt;prompt&lt;/code&gt; and an &lt;code&gt;expected&lt;/code&gt; answer: that shape moves between tools without a converter. The built-in suites in assert(llm) use exactly that shape; importing your own file is next on the roadmap.&lt;/p&gt;</description>
  </item>
  <item>
    <title>Two retired Claude Haiku ids now return 404</title>
    <link>https://assertllm.com/wire/2026-10-06-retired-haiku-ids/</link>
    <guid isPermaLink="true">https://assertllm.com/wire/2026-10-06-retired-haiku-ids/</guid>
    <pubDate>Tue, 06 Oct 2026 00:00:00 GMT</pubDate>
    <description>&lt;p&gt;&lt;code&gt;claude-3-5-haiku-20241022&lt;/code&gt; was retired on 19 February 2026, and &lt;code&gt;claude-3-haiku-20240307&lt;/code&gt; followed on 20 April. A request to either id now fails with &lt;code&gt;404 not_found_error&lt;/code&gt;, so an eval that still lists them scores every row as an error, not a wrong answer. A full row of errors for one model means: check the id before you blame the model.&lt;/p&gt;</description>
  </item>
</channel>
</rss>
