How to Know If Your AI Is Actually Working
The Kitebase support bot from article 04 sends three help-center chunks to Claude with every question. A teammate wants to cut that to one, since that’s two thirds fewer excerpt tokens per question. They try three questions by hand, the answers look fine, and they ship it.
A week later a customer writes: “I lost my phone and can’t get a two-factor code. How do I sign in?” The one chunk search now returns is “Turn on two-factor”: go to Settings and scan the QR code. The chunk about backup codes, the actual answer, came second, and second no longer makes it into the prompt. At best, Claude says it doesn’t know. At worst, it tells someone who just lost their phone to scan a QR code with it.
Nothing crashed and no test failed, because there were no tests for this. What catches this before a customer does is an eval (short for evaluation): a fixed set of questions you run through the system after every change, scored automatically, so you can see what got better and what got worse.
What you’ll build: an eval harness for the Kitebase bot. It runs 12 labelled questions, scores retrieval and answers, prints a scorecard, compares it with the last accepted version, and fails the pull request’s checks when something gets worse. It runs offline, and one flag switches it to real answers from Claude.
Why AI needs a different kind of test
A unit test compares exact values: add(2, 3) returns 5, or the test fails. You can’t do that with an LLM. It can phrase a correct answer a hundred ways, and it phrases it differently from one run to the next.
So an eval doesn’t check the exact text. It checks things that must be true of any good answer:
- It contains the facts. An answer about reset links must say “30 minutes”.
- It cites the right source. The citation must point at the reset-password article.
- It knows its limits. If the docs don’t cover the question, it says so.
And for a RAG bot, you check two stages separately. If search didn’t find the right chunk, no prompt can save the answer. If search found it and the answer is still wrong, the problem is the prompt or the model. Separate numbers tell you where to look.
The golden set
A golden set is the list of questions you test with, each labelled with what a correct result looks like. Here’s one case from the companion example’s golden.json:
{"id": "no-recovery-email",
"question": "How do I reset my password if I never set a recovery email?",
"sources": ["account-recovery"],
"facts": ["I can't access my email"]}
sources is the help article that answers it. facts are strings a correct answer must contain; here it’s the name of the button the customer has to click. The example has 12 cases like this, and they aren’t all easy on purpose:
- Ten ordinary questions, one or two per help article.
secure-accountneeds two articles (recovery email and two-factor), so it has two facts.paraphraseis the “I’m locked out and have no backup address” from article 04: same need as the first case, almost no shared words.bank-transferasks something no article covers. Itssourcesandfactsare empty, and the only right answer is “I don’t know”.
In a real product, the best questions come from users: sample them from your logs and label each one. Every bug report becomes a case you never delete. Add questions your docs don’t cover, and a few phrased the way customers actually type.
How many questions do you need?
Enough that one question moving doesn’t swing the score. With 12 questions, each one is 8 percentage points, so every change in the score is a specific question you can go and read. That’s fine for learning and for a first version.
With 200 questions, each one is half a point, and you can split them by topic (billing, sign-in) and still have enough in each group to trust. Start with 20 to 50 you’ve labelled carefully, and grow from bug reports and logs. A few hundred careful cases beat thousands of sloppy ones: a wrong label looks exactly like a bot failure.
Here’s what one eval run does with each case, and with all 12 at the end:
Scoring retrieval: hit rate and recall at k
First question: did search bring back what the answer needs? k is how many chunks search returns, the TOP_K setting in the bot. Two numbers answer it:
- Hit rate@k: the share of questions where at least one expected fact came back in the top k chunks.
- Recall@k: the share of all expected facts that came back, averaged over questions. It only differs from hit rate when a question needs more than one thing.
A fact counts as retrieved when a chunk from the right article contains it:
hits = search(case["question"], index, top_k)
retrieved = [c.text.lower() for _, c in hits if c.source in case["sources"]]
found = [f for f in case["facts"] if any(f.lower() in text for text in retrieved)]
hit = bool(found)
recall = len(found) / len(case["facts"])
For three of the cases, with the bot as it is (one chunk per section, top 3):
question retrieval
no-recovery-email hit
secure-account 50%
paraphrase MISS
no-recovery-email finds its fact in the top chunk. secure-account gets the recovery-email text but not the two-factor text, so it’s a hit with 50% recall. And paraphrase is a clean miss: the offline stand-in embedder matches words, not meaning, and the eval now puts that weakness in the numbers. Across the 11 questions the docs cover, hit rate@3 is 91% and recall@3 is 86%. bank-transfer has nothing to retrieve, so it’s left out of both.
Why label facts instead of chunk ids like account-recovery#1? Because chunk ids change every time you change how you chunk, and changing the chunking is exactly the kind of thing you want to evaluate. A fact is the same text however the article gets split. Article 07 uses the same idea, recall@10 on your own questions, to pick an embedding model.
Scoring answers: facts, citations, “I don’t know”
Second question: is the answer right? check_answer returns a list of problems, and an empty list is a pass:
def check_answer(case: dict, reply: str, hits) -> list[str]:
text = reply.lower().replace("’", "'") # models often write a curly ’ in "don’t"
if not case["sources"]: # the docs don't cover it
return [] if "don't know" in text else ["should say it doesn't know"]
problems = [f'missing "{fact}"' for fact in case["facts"] if fact.lower() not in text]
cited = re.findall(r"\[([\w-]+#\d+)\]", reply)
if any(cid not in {c.id for _, c in hits} for cid in cited):
problems.append("cites a chunk it was never given")
elif not any(cid.split("#")[0] in case["sources"] for cid in cited):
problems.append("wrong article cited")
return problems
The citation check is unknown_citations from article 04, plus one more rule: the citation has to point at the right article.
The answers need to come from somewhere. By default the example uses a stand-in answerer that pastes the top retrieved chunk as the answer, with its id as the citation. It’s a dumb bot on purpose: free and deterministic, so the whole harness runs offline. python main.py --claude runs the same checks on real answers from claude-opus-5. Here’s the full scorecard with the stand-in:
Kitebase bot: chunks by section, top_k=3, 13 chunks, 12 golden questions
Answers from: offline stand-in (pastes the top chunk)
question retrieval answer
no-recovery-email hit pass
link-expiry hit pass
...
lost-phone hit FAIL missing "backup codes"
change-sign-in-email hit FAIL missing "Settings > Profile"
secure-account 50% FAIL missing "authenticator app"; wrong article cited
paraphrase MISS FAIL missing "I can't access my email"; wrong article cited
bank-transfer - pass
hit rate@3 91% recall@3 86%
answers pass 67% prompt ~249 tokens per question
Read lost-phone: retrieval is a hit, the answer fails. The backup-codes chunk came back second, and the stand-in only reads the first. That’s what the two separate numbers buy you: a failure with a retrieval hit is an answering problem. bank-transfer passes because every chunk scores below the bot’s minimum score of 0.2, so the bot says “I don’t know” instead of pasting a sign-in chunk.
67% is fine. The point isn’t 100% on day one; it’s that the score never quietly drops.
String checks are cheap, fast and deterministic, and they’re brittle. A real answer that says “half an hour” fails the “30 minutes” check. Pick facts that a correct answer can’t avoid, like a button label, a menu path or a number, and read the failures before you trust them. For qualities a string can’t check, like “is this answer polite and complete?”, you ask another model to grade the answer against a short rubric. That’s LLM-as-judge, and it has its own traps, covered in article 09.
Comparing two versions
A scorecard on its own tells you where you are. The useful part is comparing it with a baseline: the saved scorecard of the version you last accepted. python main.py --save-baseline writes it to baseline.json, and you commit that file. Every later run compares against it and lists each regression, a question that used to pass and now fails. Here’s the teammate’s change from the opening:
$ python main.py --top-k 1
...
vs baseline (chunks by section, top_k=3):
hit rate 91% -> 64%
recall 86% -> 64%
answers pass 67% -> 67%
prompt tokens 249 -> 160
WORSE: lost-phone: retrieval 100% -> 0%
WORSE: change-sign-in-email: retrieval 100% -> 0%
WORSE: secure-account: retrieval 50% -> 0%
Prompts are about 90 tokens smaller, and three questions lost the chunk that answers them, including the lost phone. The answer score didn’t move only because the stand-in reads nothing past the top chunk anyway. Claude reads every chunk you send, so with real answers those three would suffer.
Now a different change: fixed 50-word chunks instead of one per section, the naive fixed-size chunking from article 05. The answer score is 67% again, same as the baseline. But the per-question diff says something else:
$ python main.py --chunk-words 50
...
fixed: lost-phone: answer now passes
fixed: change-sign-in-email: answer now passes
fixed: secure-account: retrieval 50% -> 100%
WORSE: sso-reset: answer now fails, missing "IT admin"
WORSE: downgrade-timing: retrieval 100% -> 0%
WORSE: downgrade-timing: answer now fails, missing "end of the billing period"
Two answers fixed, two broken, and the total didn’t move. Compare totals and you’d call it neutral and ship it. So the harness diffs every question, and fails on “anything that passed before and fails now”, not “the total dropped”.
Here’s what broke downgrade-timing:
The chunker cut “Downgrades apply at the” from “end of the billing period.” The second half is a 5-word chunk that shares no words with the question, so it never comes back. You’d never spot that reading the chunking code.
Running it in CI
CI (continuous integration) is the service that runs your checks on every pull request, like GitHub Actions. It treats a non-zero exit code as a failed check, so the harness ends with:
if worse:
sys.exit(1) # a non-zero exit is what fails the CI job
The workflow is the same as for any Python test. From evals.yml in the example:
on: pull_request
jobs:
offline:
runs-on: ubuntu-latest
steps:
- uses: actions/checkout@v4
- uses: actions/setup-python@v5
with:
python-version: "3.12"
- run: pip install -r requirements.txt
- run: python main.py # exits 1 on any regression, which fails the check
Now the --top-k 1 pull request gets a red check listing the three questions it broke. When a change is an improvement, or a trade-off you accept, run --save-baseline and commit the new baseline.json in the same pull request, so the reviewer sees what moved.
Split what you run by cost and by how repeatable it is:
- Every pull request: retrieval and the offline checks. With a fixed embedder, search gives the same results every run, and it’s free. A failure always means the code changed something.
- Nightly, and on prompt or model changes: real answers.
--claudemakes 11 calls (the bank-transfer question never reaches the model), about 250 input tokens and 100 to 200 output tokens each. Atclaude-opus-5’s $5 per million input and $25 per million output tokens, that’s about 5 cents a run. At 500 questions it’s a couple of dollars, which adds up per pull request.
Real answers vary between runs, so one flipped check isn’t proof of a regression: re-run the failing question a few times and read the answer. For the same reason, the harness doesn’t compare --claude runs with the stand-in baseline.
After you ship
The golden set only covers what you thought to ask. Keep feeding it from production: sample real questions, read the answers users marked unhelpful, and add every failure you find to golden.json. Article 11 covers logging what you need for that.
Try it yourself
The companion example is the whole harness: the Kitebase bot, 12 golden questions, a saved baseline and a CI workflow. It runs offline with no keys.
Download the runnable example (zip)
cd 08-llm-evaluation-pipeline
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
Then try these:
python main.py --top-k 1, thenecho $?. The exit code is 1: that’s the red check in CI. Try--top-k 5and see which numbers move and which don’t.python main.py --chunk-words 25. Hit rate falls to 64% and answers to 42%. Find the question whose fact got cut in half.- Add a case to
golden.json:"Do backup codes work more than once?"with sourcetwo-factorand fact"Each one works once". Runpytest -qfirst: one test checks every labelled fact really is in its article, so a typo fails a test instead of looking like a bot bug. Then save a new baseline.
pip install pytest && pytest -q runs the offline tests. They need no keys and no network.
Common beginner mistakes
- Only checking by hand. Three questions you already know the answers to aren’t an eval. Write them down, label them, run them every time.
- Trusting one total. 67% to 67% hid two fixes and two breaks. Diff per question.
- Only scoring the final answer. Without a retrieval number you can’t tell a search problem from a prompt problem.
- Labelling chunk ids. They change the moment you re-chunk, and re-chunking is what you wanted to test. Label facts and source articles.
- Never checking the labels. A typo in a fact looks exactly like a bot failure. Test that every fact appears in its source.
Questions you will face in production
“Can we just use a framework?” Yes, once you know what you’re measuring. RAGAS, Inspect AI and OpenAI’s evals give you runners, reports and ready-made metrics. The golden set is still yours to write, and it’s most of the value. Starting with a small script like this one teaches you what those frameworks do.
“Our answers are free text. What do we check?” Start with what a correct answer can’t avoid: a number, a menu path, a product name, a citation. Add an LLM judge with a short rubric for the rest, and check the judge against your own ratings on a sample first. Article 09 shows how.
“What score should we aim for?” Don’t pick a number up front. Your first scorecard is the baseline, and the rule is that nothing that passes today may fail tomorrow. Raise the bar as fixes land.
Check your understanding
Hit rate@3 is 100% but answers pass only 50%. Where do you look first?
The answering step: the prompt, the model, or how the answer uses the chunks. Search is bringing back the facts. Read a few failing answers next to the chunks they were given.
A pull request changes the chunk size. Hit rate goes from 91% to 91%. Is it safe to merge?
Not yet. Check the per-question diff: one question can be fixed and another broken with the same total. Merge if nothing that passed before now fails, or if you’ve looked at what broke and decided the trade-off is worth it, and saved a new baseline.
Your golden set labels each question with a chunk id like invoices#2. You switch to 50-word chunks. What goes wrong?
invoices#2 now names a different piece of text, a 5-word fragment, so the eval scores against the wrong target. Label the source article and a fact string instead; those don’t change when the chunking does.
A nightly run with real answers fails one question that passed last night. Nothing changed in the code. What do you do?
Re-run that question a few times and read the answers. Model output varies, so a single flip can be noise, or a string check that’s too strict (“half an hour” vs “30 minutes”). If it keeps failing with no code change, the model or provider changed, and the eval caught it.
What to remember
- An eval is a fixed set of labelled questions, run after every change, scored automatically.
- Label each question with its source and the facts a correct answer must contain, not chunk ids.
- Score retrieval (hit rate and recall at k) and answers (facts, citations, “I don’t know”) separately, so a failure tells you where to look.
- Compare every run with a saved baseline, question by question. Fail on anything that used to pass.
- Run the cheap, deterministic checks on every pull request and the real-model checks nightly.
- Grow the golden set from real questions and every bug you find.
What to study next
String checks cover facts and citations. For everything else you need a model to grade answers, and judges are biased in ways that look like bugs in your bot. Article 09: LLM-as-Judge covers how to write a rubric, check a judge against people, and avoid the known biases.
Further reading
- Hamel Husain: Your AI Product Needs Evals. The best practical write-up on building evals into a product. Read this if you read nothing else.
- Eugene Yan: Task-Specific LLM Evals that Do & Don’t Work. Which metrics hold up for classification, summarization and Q&A, and which don’t.
- RAGAS. A RAG evaluation framework with metrics for retrieval and for whether answers stick to the context.
- OpenAI evals. Open-source eval framework with many built-in evals to learn from.
- Inspect AI. An eval framework from the UK AI Security Institute, good for graded evaluations.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the broad mechanics and specific numbers come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.