← All postsDeep dive

How to test an LLM for hallucinations

A hallucination is a test failure you have not written yet. Start from answers you already know, let deterministic checks decide, and keep the judge advisory.

Ask a model the same question twice and you may get two confident, fluent, different answers. One of them is made up. The problem is not that models hallucinate; it is that most teams find out from a user, not from a test.

This post walks through how assert(llm) catches hallucinations with plain test engineering: a golden dataset, scorers that do not guess, and an LLM judge that is allowed to comment but not to vote.

Hallucination is a test failure, not a vibe

“The model sometimes makes things up” is not actionable. “Item CODE-007 fails on every run because the answer imports an SDK and calls a model id that does not exist” is. The difference is a test with a known expected answer and a rule that decides pass or fail the same way every time.

  • Known answers. Each item carries an expected value you can defend.
  • Deterministic verdicts. The same reply always gets the same score.
  • Reproducible runs. A failure you can re-run is a failure you can fix.

Start from answers you already know

A golden dataset is a list of prompts whose correct answers you know ahead of time. In assert(llm) every item is typed, and the type decides how it is scored: mcq compares the letter the model picked, code checks the function signature and forbidden constructs, short looks for the expected phrase on word boundaries.

Fig. 1 · Where a hallucination gets caughtassert(llm)
01 · inputGolden itemprompt + expected
type: mcq · code · short
02 · modelModel answerraw text, straight from the provider
03 · scorerDeterministic checksletter parse · signature + arity · word-boundary match · global rulesCODE-007 fails here
04 · verdictPass or fail
PASSFAIL ✕
+ Judge (optional) a second model reads the same answer · 0–5 + one-line reason advisory: never changes the verdict
The verdict always comes from step 03. A judge adds an opinion next to it, which is useful for reading failures but never decides them.

Case: phantom delegation

Asked for a CSV parser, Haiku 4.5 repeatedly answers with code that imports the Anthropic SDK and asks a different model to do the parsing. It is not a one-off: it recurs on every run of the item.

haiku-4.5 · answer to CODE-007python
import anthropic
def csv_to_dicts(s):
    client = anthropic.Anthropic()
    resp = client.messages.create(
        model="claude-3-5-sonnet-20241022",  # does not exist
        messages=[{"role": "user", "content": f"Parse this CSV: {s}"}],
    )
    return eval(resp.content[0].text)  # eval on model output

Three problems stack up in nine lines: a model id that does not exist, an eval() on another model's output, and no parsing at all. The task was to write a parser; nothing in the answer parses anything.

Note

The scorer missed this at first. The answer defines the right function with the right arity and returns a value, so it passed every item-level check. It now fails on a global rule that forbids importing an LLM SDK in a coding answer.

When the scorer is the bug

Writing tests against the scorer found three bugs in the scorer rather than in any model. Each one turned a correct verdict into a wrong one:

ReplyExpectedBefore the fixNow
“A good test objective is to reduce risk”aPASS article read as option aFAIL article skipped
“I cannot answer without the options”cPASS substring “c”FAIL word boundaries
“299 792 km/s”299792FAIL thin space lostPASS de-grouped first

For a while the eval was measuring its own matching rules at least as much as it was measuring the models.

— Findings, October 2026

Where an LLM judge fits

A judge is good at explaining a failure in a sentence and bad at being the source of truth. In assert(llm) each judge scores the same answer 0–5 with a one-line reason, and that score sits next to the verdict instead of replacing it. Two judges that disagree stay visible as two averages.

Mini-lab

Run the Python coding suite on your own models

20 coding tasks, the same SDK-import rule, any of six providers. Your key stays in the tab.

Run it →
Taggedhallucinationgolden datasetllm-as-judge