LLM-as-Judge: When It Breaks and How to Fix It
The Kitebase support bot from article 04 passes its citation check: every id it cites was in the prompt. Then a customer asks how long the password reset link lasts, and the bot says 24 hours, citing [reset-password#1]. The excerpt says 30 minutes. No regex catches that, and you can’t read 200 answers by hand every time you change the prompt.
So you hand the reading to a second model. That’s LLM-as-judge, and article 08 showed where it sits in an eval pipeline. The catch: the judge is a model too, it’s wrong in its own ways, and the only way to know how often is to compare it with a person.
What you’ll build: a judge that grades 12 Kitebase answers pass or fail against a short rubric, returning a validated Python object, and a script that measures how often it agrees with a human who labelled the same 12. It runs offline with canned verdicts; an API key swaps in claude-sonnet-5.
What a judge is
A judge is a model call whose only job is to grade another model’s output. It reads what the bot read, what the bot wrote, and a rubric (the grading rules, written out), and returns a verdict. Think of it as a test assertion written in English: assert answer_matches_excerpts(answer), where the function body is a prompt.
Use rules first and a judge for what rules can’t check. The unknown_citations check from article 04 is a rule: free, instant and never wrong about whether an id was sent. Whether “24 hours” matches “30 minutes” in meaning is a judge’s job.
Step 1: Write a rubric the judge can follow
The obvious first judge prompt is “Rate this answer from 1 to 5.” It gives you a number you can’t act on. What’s the difference between a 3 and a 4? Ask twice and you might get both.
A rubric that works is a short list of yes/no questions, each about one thing, that two people would answer the same way. Hamel Husain’s guide to LLM judges recommends pass/fail over numeric scales for exactly this reason: a fail tells you what to fix. For the Kitebase bot two questions cover most failures:
You grade answers from Kitebase's support bot. The bot must answer only from the
help-center excerpts it was given.
You get the customer's question, the excerpts and the bot's answer. Judge two
things separately:
grounded: true if every fact in the answer is stated in the excerpts. A fact that
isn't in the excerpts makes it false, even if it sounds reasonable.
answers_question: true if the answer tells the customer what they need for the
question they asked.
Write the critique first: one or two sentences quoting the part of the answer your
verdict depends on.
The user message carries one case, tagged the same way the bot’s prompt was:
<question>How long does the password reset link last?</question>
<excerpts>
<excerpt id="reset-password#1">
Reset your password: Send yourself a reset link
On the sign-in page, click **Forgot password?** ... The link works once and expires after 30 minutes.
</excerpt>
</excerpts>
<answer>
The reset link works once and expires after 24 hours [reset-password#1].
</answer>
The judge gets the same excerpts the bot got. Without them it can only judge whether the answer sounds right, and “24 hours” sounds fine.
Keep the criteria separate and combine them in your code. When grounded fails, the bot made something up. When answers_question fails with grounded true, the bot stuck to the excerpts but they were the wrong ones, which points at retrieval. Start with two to four criteria; each extra one is another thing to calibrate.
Step 2: Get a verdict you can parse
Ask for “JSON with these fields” in plain text and most replies are fine. Then one comes back wrapped in a code fence, or with "grounded": "no", and your eval run crashes at case 143.
Structured outputs fix this. You describe the shape as a Pydantic model (a Python class that validates data), and client.messages.parse makes the API return exactly that shape, which the SDK checks and hands back as an object:
class Verdict(BaseModel):
critique: str # first, so the judge explains before it decides
grounded: bool
answers_question: bool
@property
def passed(self) -> bool:
return self.grounded and self.answers_question
def judge(client, case: Case) -> Verdict | None:
response = client.messages.parse(
model="claude-sonnet-5",
max_tokens=1024,
system=JUDGE_SYSTEM,
messages=[{"role": "user", "content": build_judge_prompt(case)}],
output_format=Verdict,
)
if response.stop_reason in ("refusal", "max_tokens"):
return None # no trustworthy verdict; count it, don't guess
return response.parsed_output
For kb-02 you get back a real Verdict, not a string:
>>> verdict = judge(client, kb02)
>>> verdict.grounded, verdict.answers_question, verdict.passed
(False, True, False)
>>> verdict.critique
'The answer says the link expires after 24 hours; reset-password#1 says 30 minutes.'
passed lives in your code, not in the prompt, so the rule “both criteria must hold” can’t drift between runs. And a refusal or a cut-off reply returns None, which the calibration script counts separately. Treating a missing verdict as a pass hides exactly the cases you most need to see.
Why put the critique before the true/false fields?
The model writes its reply in order. With critique first, it has quoted the evidence (“24 hours” against “30 minutes”) before it writes grounded, like asking someone to show their work.
The bigger payoff is yours: when the judge and a human disagree, the critique is what you read to find out why. That’s how you fix the rubric in Step 4.
Step 3: Pick a cheaper model for the judge
The bot answers with claude-opus-5. The judge doesn’t have to. It runs on every case in every eval run, so its cost multiplies, and its job is narrower than the bot’s: compare one answer against a few paragraphs it’s been handed.
A judge call for a real bot answer, with three excerpts, is roughly 700 input tokens and 100 output tokens. Over a 200-case eval set, at current list prices:
| Judge model | Per case | Per 200-case run |
|---|---|---|
claude-opus-5 ($5 in / $25 out per million) | $0.0060 | $1.20 |
claude-sonnet-5 ($2 / $10) | $0.0024 | $0.48 |
claude-haiku-4-5 ($1 / $5) | $0.0012 | $0.24 |
The example defaults to claude-sonnet-5: checking every fact against the excerpts needs careful reading, and it’s still 2.5 times cheaper than Opus. Try claude-haiku-4-5 next. Whether a cheaper judge is good enough isn’t something to guess; the next step measures it. (Sonnet and Opus are the same family, which matters for self-preference, below.)
Step 4: Calibrate the judge against human labels
Run the judge and it passes 9 of 12 answers. Is the bot good, or is the judge lenient? From the judge’s output alone you can’t tell.
So you test the tester. Calibration means a person labels a sample of answers pass or fail, the judge grades the same sample, and you compare. In the example, data/cases.json holds 12 labelled cases: 9 pass, 3 fail. python main.py grades them and prints:
kb-01 human pass judge pass grounded=True answers=True
kb-02 human fail judge fail grounded=False answers=True
...
kb-05 human pass judge fail grounded=True answers=False <-
kb-07 human fail judge pass grounded=True answers=True <-
...
Judge agreement 83% kappa 0.56 fails caught 2 of 3 passes kept 8 of 9
Always pass agreement 75% kappa 0.00 fails caught 0 of 3 passes kept 9 of 9
The “Always pass” row is a fake judge that passes everything, and it’s there to make a point: it agrees with the human 75% of the time and catches nothing. When most answers are fine, percent agreement (how often the two labels match) looks good even for a useless judge. Two more numbers fix that:
- Cohen’s kappa is agreement after subtracting what two raters with these pass rates would hit by luck. 0 means no better than luck, 1 means perfect. Always-pass scores exactly 0.
- Fails caught is how many of the human’s fails the judge also failed. If the judge gates merges, this is the number that matters: every missed fail is a regression that ships.
def agreement(human: list[bool], judged: list[bool]) -> Agreement:
n = len(human)
observed = sum(h == j for h, j in zip(human, judged)) / n
h_pass, j_pass = sum(human) / n, sum(judged) / n
chance = h_pass * j_pass + (1 - h_pass) * (1 - j_pass)
kappa = (observed - chance) / (1 - chance) if chance < 1 else 0.0
fails = [j for h, j in zip(human, judged) if not h]
passes = [j for h, j in zip(human, judged) if h]
return Agreement(observed, kappa, (fails.count(False), len(fails)),
(passes.count(True), len(passes)))
For the judge: observed agreement is 10/12 = 0.83. Both raters pass 9 of 12, so luck alone gives 0.75 × 0.75 + 0.25 × 0.25 = 0.625. Kappa is (0.833 − 0.625) / (1 − 0.625) = 0.56. On the scale Landis and Koch proposed in 1977, 0.41 to 0.60 is “moderate” and 0.61 to 0.80 “substantial”. Treat those bands as rough labels, not pass marks.
Read the disagreements, then fix the rubric
The numbers tell you how much to trust the judge. The disagreements tell you what to change. The script prints each one with both notes:
kb-05 Can I pay by bank transfer?
human: The excerpts don't cover it, so saying so is the right answer.
judge: The customer asked about bank transfers and the answer doesn't say whether they're possible.
kb-07 Where do I download my invoices?
human: Emailed invoices and the ZIP export aren't in the excerpt. Made up.
judge: A thorough answer: gives the path from invoices#1, explains who can see billing, and adds helpful tips.
kb-05 is a rubric bug: the bot is told to say “I don’t know” when the excerpts don’t cover a question, and the rubric never said that counts. kb-07 is worse, a real failure waved through: the right path, then two features Kitebase doesn’t have, wrapped in “Great question!”. Two lines go into the rubric:
If the excerpts don't cover the question, saying so and pointing to support counts as answering it.
Length and friendliness earn nothing. Check each extra detail against the excerpts.
Then run the judge again on the same labelled cases. If kb-05 flips, agreement goes to 11 of 12 and kappa to 0.75. Recalibrate whenever you change the rubric or the judge model: a new rubric is a new judge.
Twelve cases is enough to show the method and far too few to trust: one case moves agreement by 8 points. My default is 50 to 100 labelled answers with at least 15 or so real failures in them, taken from real traffic. Before a judge gates merges, I want it catching at least 8 of every 10 failures a human finds.
This isn’t hopeless. In Zheng et al.’s MT-Bench study, GPT-4 as a judge agreed with human experts 85% of the time on pairwise votes that weren’t ties, higher than the 81% the experts managed with each other. A calibrated judge can be as consistent as a person. You just have to check yours.
The biases, and what’s actually measured
Calibration catches problems on your data, but it helps to know what to look for. The numbers below come from 2023 and 2024 models; today’s judges may do better or worse, which is one more reason to measure your own.
Position bias
When a judge compares two answers, position bias means it favours one slot, usually the first, regardless of content. Zheng et al. showed judges the same pair twice with the order swapped. GPT-4 gave the same verdict both times in 65% of cases; Claude-v1 in only 23.8%, and in 75% of cases it favoured whichever answer came first.
The fix is the one the paper proposes: ask twice with the order swapped, and only count a win that survives both.
def compare(client, case: Case, first: str, second: str) -> str:
one = prefer(client, case, first, second) # first shown as A
two = prefer(client, case, second, first) # first shown as B
if (one, two) == ("A", "B"):
return "first"
if (one, two) == ("B", "A"):
return "second"
return "tie" # it changed its mind when the order changed, or said tie
It doubles the cost of each comparison, and it’s worth it. This only applies to pairwise judging (which of two answers is better), the natural fit for choosing between prompt v3 and v4: compare each pair with the swap and count how often v4 wins. For a gate on every change, the pass/fail rubric is simpler and has no slots to be biased about.
Verbosity bias
Verbosity bias is rating longer answers higher because they’re longer. Zheng et al. tested it by padding answers with repeated content: Claude-v1 and GPT-3.5 preferred the padded version 91.3% of the time, GPT-4 8.7%. It varies a lot by judge. kb-07 is the Kitebase version. The defence is the “length earns nothing” line in the rubric, plus padded answers in your labelled set so calibration catches a relapse.
Self-preference bias
Self-preference bias is a judge rating answers from its own model higher. The evidence is weaker than for the other two. Zheng et al. saw GPT-4 give its own answers about a 10% higher win rate and Claude-v1 about 25%, but wrote that their data couldn’t settle whether the bias was real. Panickssery et al. (2024) found that GPT-4 and Llama 2 can often recognise their own summaries, and that the better a model got at recognising its own writing, the more it preferred it.
So Sonnet judging Opus might be more lenient than a human. You don’t need to guess: an over-lenient judge shows up in calibration as missed failures. If you’re worried, also grade the labelled set with another provider’s model and compare both against the human.
Try it yourself
The companion example grades the 12 Kitebase answers, prints the agreement numbers and the disagreements, and has a pairwise demo with a fake judge that always picks A.
Download the runnable example (zip)
cd 09-llm-as-judge
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
python main.py --pairwise
Then try these:
- Add the two rubric lines from Step 4 to
JUDGE_SYSTEM. Offline, stand in for the rerun by settingkb-05to"answers_question": trueindata/canned_verdicts.json: agreement goes to 92% and kappa to 0.75. WithANTHROPIC_API_KEYset, rerun the real judge instead. - Make failures rarer. Change
kb-02andkb-10to"pass"indata/cases.json. Always-pass now agrees 92% of the time, and its kappa is still 0.00. - Set
JUDGE_MODEL = "claude-haiku-4-5"and, with a key, run it again. If fails caught holds up, keep the cheaper judge.
pip install pytest && pytest -q runs the offline tests. They need no keys and no network.
Common beginner mistakes
- One 1-to-5 score for everything. “Correct, on topic and well written” squashed into a 3 tells you nothing about what to fix. Use separate pass/fail criteria.
- Trusting the judge without labels. An unchecked judge is a second model you haven’t tested. Label 50 answers before its scores decide anything.
- Reporting only percent agreement. When 9 of 10 answers pass, a judge that passes everything scores 90%. Report kappa and fails caught.
- Counting a missing verdict as a pass. A refusal or cut-off reply isn’t a verdict. Count it separately.
Questions you will face in production
“Who should label the calibration set?” One person who knows what a correct answer is, like the support lead for Kitebase, with a short note per label. One consistent labeller beats a committee that disagrees with itself.
“The judge passes 98% of answers. Is the bot that good?” Maybe, but check the judge first. Pull 30 of its passes and have a person label them. If they find failures, tighten the rubric wording and add those answers to the labelled set.
“Can I run the judge on live traffic?” Yes, on a sample. Judge a few hundred real answers a day and track the fail rate over time; a jump after a deploy or a docs change is worth a look. Article 11 covers tracking numbers like this.
Check your understanding
Your judge agrees with the human on 45 of 50 answers. Only 4 of the 50 are human fails, and the judge caught 1 of them. Is it ready to gate merges?
No. 90% agreement sounds strong, but a judge that passed everything would score 92% here. It caught 1 of 4 real failures, so three of every four regressions would merge. Read the three it missed, fix the rubric, and label more failing answers so the number means something.
The judge fails an answer on answers_question but passes it on grounded. What would you look at first?
Retrieval. The bot stuck to its excerpts but didn’t answer the question, so the excerpts were probably the wrong ones. Check which chunks search returned for that question, as in article 04.
You compare prompt v3 and v4 pairwise without swapping, and v4 wins 70% of the time. v4's answers were always shown first. What's wrong?
Some of that 70% may be position bias, not v4. Run each comparison twice with the order swapped and only count wins that survive both orders. If the win rate drops a lot, the first result was mostly about the slot.
What to remember
- A judge is a model call that grades another model’s output against a written rubric. Use rules for what rules can check, and a judge for the rest.
- Write the rubric as a few pass/fail criteria, each about one thing, and give the judge everything the bot saw.
- Use
messages.parsewith a Pydantic model so every verdict is a validated object, critique first. Don’t count a missing verdict as a pass. - A cheaper model like
claude-sonnet-5is usually enough for the judge, since it runs on every case. Calibration decides. - Calibrate against human labels. Report agreement, kappa and fails caught, then read the disagreements and fix the rubric.
- Position, verbosity and self-preference bias vary by judge. Swap the order for pairwise judging, and let your labelled set show what yours does.
What to study next
You can now measure whether the Kitebase bot is getting better or worse, and trust the measurement. Next is making it cheaper and faster without changing its answers: article 10: Caching for LLM Apps covers prompt caching, semantic caching and embedding caches.
Further reading
- Zheng et al. (2023): Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. The source of the position, verbosity and self-preference numbers above, and the swap-the-order fix.
- Panickssery, Bowman and Feng (2024): LLM Evaluators Recognize and Favor Their Own Generations. The self-recognition and self-preference result.
- Hamel Husain: Using LLM-as-a-Judge for Evaluation. A practical guide to building judges with one domain expert, pass/fail labels and written critiques.
- Landis and Koch (1977): The Measurement of Observer Agreement for Categorical Data. Where the kappa bands come from.
- LMSYS Chatbot Arena. Pairwise human judging at large scale; the methodology is worth reading.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the broad mechanics and specific numbers come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.