Which LLM actually gets it right?

Test Claude, GPT, Gemini, Llama and more on a ready suite or on your own prompts. Straight from your browser, with your own key.

01

What to test

a ready suite, a copy of one, or your own
Your prompts 0
start from
Prompts live in this tab, like your keys. Export to keep them.
#PromptExpected answerScored as
02

Models & keys

a key unlocks its provider's models
Test on
Judge
A model can't judge itself.

A second model scores every answer 0–5 with a one-line reason. Advisory: it never changes pass/fail. Each judge adds one call per answer.

Ready

0 prompts × 0 models

Where your key goes

Your browser

this tab holds the key and your prompts

→

Provider API

Anthropic, OpenAI, Google, Groq, Mistral, OpenRouter

→

Report, here

export JSON or Markdown

No proxy and no account. The page is static files; every call goes from your browser directly to the provider you picked.

Read before you run

Pass rate
Every answer, at a glance
passfailerror
Show run log ↓