Which LLM actually gets it right?
Test Claude, GPT, Gemini, Llama and more on a ready suite or on your own prompts. Straight from your browser, with your own key.
01
What to test
a ready suite, a copy of one, or your own
Your prompts
0
start from
Prompts live in this tab, like your keys. Export to keep them.
#PromptExpected answerScored as
02
Models & keys
a key unlocks its provider's models
Test
on
Judge
A model can't judge itself.
A second model scores every answer 0–5 with a one-line reason. Advisory: it never changes pass/fail. Each judge adds one call per answer.
With judge enabled, each question costs 2× API calls.
Ready
0 prompts × 0 models
Where your key goes
Your browser
this tab holds the key and your prompts
Provider API
Anthropic, OpenAI, Google, Groq, Mistral, OpenRouter
Report, here
export JSON or Markdown
No proxy and no account. The page is static files; every call goes from your browser directly to the provider you picked.
Read before you run
Finding · coding
Asked for code, the model called an API that doesn't exist
repeatable across runs
Read the finding → Learn · glossaryGolden dataset, LLM-as-judge, hallucination: the terms, short
Open the glossary →
Learn · soon
How to test an LLM for hallucinations
ends in a mini-lab you run with your own key
0/0
answers in
Elapsed
0:00
Left, about
—
Tokens in → out
0 → 0
Judged
0/0
Live grid
passfail
errorasking now
queued
Live lognewest first
Keep this tab open. Closing it stops the run.