How to Know If Your AI Is Actually Working

The Kitebase support bot from article 04 sends three help-center chunks to Claude with every question. A teammate wants to cut that to one, since that’s two thirds fewer excerpt tokens per question. They try three questions by hand, the answers look fine, and they ship it.

A week later a customer writes: “I lost my phone and can’t get a two-factor code. How do I sign in?” The one chunk search now returns is “Turn on two-factor”: go to Settings and scan the QR code. The chunk about backup codes, the actual answer, came second, and second no longer makes it into the prompt. At best, Claude says it doesn’t know. At worst, it tells someone who just lost their phone to scan a QR code with it.

Nothing crashed and no test failed, because there were no tests for this. What catches this before a customer does is an eval (short for evaluation): a fixed set of questions you run through the system after every change, scored automatically, so you can see what got better and what got worse.

What you’ll build: an eval harness for the Kitebase bot. It runs 12 labelled questions, scores retrieval and answers, prints a scorecard, compares it with the last accepted version, and fails the pull request’s checks when something gets worse. It runs offline, and one flag switches it to real answers from Claude.

Why AI needs a different kind of test

A unit test compares exact values: add(2, 3) returns 5, or the test fails. You can’t do that with an LLM. It can phrase a correct answer a hundred ways, and it phrases it differently from one run to the next.

So an eval doesn’t check the exact text. It checks things that must be true of any good answer:

  • It contains the facts. An answer about reset links must say “30 minutes”.
  • It cites the right source. The citation must point at the reset-password article.
  • It knows its limits. If the docs don’t cover the question, it says so.

And for a RAG bot, you check two stages separately. If search didn’t find the right chunk, no prompt can save the answer. If search found it and the answer is still wrong, the problem is the prompt or the model. Separate numbers tell you where to look.

The golden set

A golden set is the list of questions you test with, each labelled with what a correct result looks like. Here’s one case from the companion example’s golden.json:

{"id": "no-recovery-email",
 "question": "How do I reset my password if I never set a recovery email?",
 "sources": ["account-recovery"],
 "facts": ["I can't access my email"]}

sources is the help article that answers it. facts are strings a correct answer must contain; here it’s the name of the button the customer has to click. The example has 12 cases like this, and they aren’t all easy on purpose:

  • Ten ordinary questions, one or two per help article.
  • secure-account needs two articles (recovery email and two-factor), so it has two facts.
  • paraphrase is the “I’m locked out and have no backup address” from article 04: same need as the first case, almost no shared words.
  • bank-transfer asks something no article covers. Its sources and facts are empty, and the only right answer is “I don’t know”.

In a real product, the best questions come from users: sample them from your logs and label each one. Every bug report becomes a case you never delete. Add questions your docs don’t cover, and a few phrased the way customers actually type.

How many questions do you need?

Enough that one question moving doesn’t swing the score. With 12 questions, each one is 8 percentage points, so every change in the score is a specific question you can go and read. That’s fine for learning and for a first version.

With 200 questions, each one is half a point, and you can split them by topic (billing, sign-in) and still have enough in each group to trust. Start with 20 to 50 you’ve labelled carefully, and grow from bug reports and logs. A few hundred careful cases beat thousands of sloppy ones: a wrong label looks exactly like a bot failure.

Here’s what one eval run does with each case, and with all 12 at the end:

For each of the 12 golden questions. This one is no-recovery-email. 1. GOLDEN CASE question How do I reset my password if I never set a recovery email? source account-recovery fact the answer needs "I can't access my email" 2. RUN THE BOT search, top 3 0.61 account-recovery#1 0.51 recovery-email#1 0.50 reset-password#1 answer ... click Forgot password? and then I can't access my email ... [account-recovery#1] 3a. RETRIEVAL CHECK Is the fact in a retrieved chunk from account-recovery? Yes: a hit. 3b. ANSWER CHECK Does the answer contain the fact, and cite a chunk it was given? Yes and yes: a pass. 12 results 4. SCORECARD hit rate@3 91% recall@3 86% answers pass 67% prompt ~249 tokens not 100%, and that's fine 5. COMPARE WITH baseline.json the scorecard of the version you last accepted NOTHING GOT WORSE exit 0: the PR can merge SOMETHING GOT WORSE lost-phone: retrieval 100% -> 0% exit 1: CI blocks the PR
One eval run. The numbers are what the companion code prints.

Scoring retrieval: hit rate and recall at k

First question: did search bring back what the answer needs? k is how many chunks search returns, the TOP_K setting in the bot. Two numbers answer it:

  • Hit rate@k: the share of questions where at least one expected fact came back in the top k chunks.
  • Recall@k: the share of all expected facts that came back, averaged over questions. It only differs from hit rate when a question needs more than one thing.

A fact counts as retrieved when a chunk from the right article contains it:

hits = search(case["question"], index, top_k)
retrieved = [c.text.lower() for _, c in hits if c.source in case["sources"]]
found = [f for f in case["facts"] if any(f.lower() in text for text in retrieved)]

hit = bool(found)
recall = len(found) / len(case["facts"])

For three of the cases, with the bot as it is (one chunk per section, top 3):

  question              retrieval
  no-recovery-email     hit
  secure-account        50%
  paraphrase            MISS

no-recovery-email finds its fact in the top chunk. secure-account gets the recovery-email text but not the two-factor text, so it’s a hit with 50% recall. And paraphrase is a clean miss: the offline stand-in embedder matches words, not meaning, and the eval now puts that weakness in the numbers. Across the 11 questions the docs cover, hit rate@3 is 91% and recall@3 is 86%. bank-transfer has nothing to retrieve, so it’s left out of both.

Why label facts instead of chunk ids like account-recovery#1? Because chunk ids change every time you change how you chunk, and changing the chunking is exactly the kind of thing you want to evaluate. A fact is the same text however the article gets split. Article 07 uses the same idea, recall@10 on your own questions, to pick an embedding model.

Scoring answers: facts, citations, “I don’t know”

Second question: is the answer right? check_answer returns a list of problems, and an empty list is a pass:

def check_answer(case: dict, reply: str, hits) -> list[str]:
    text = reply.lower().replace("’", "'")  # models often write a curly ’ in "don’t"
    if not case["sources"]:  # the docs don't cover it
        return [] if "don't know" in text else ["should say it doesn't know"]
    problems = [f'missing "{fact}"' for fact in case["facts"] if fact.lower() not in text]
    cited = re.findall(r"\[([\w-]+#\d+)\]", reply)
    if any(cid not in {c.id for _, c in hits} for cid in cited):
        problems.append("cites a chunk it was never given")
    elif not any(cid.split("#")[0] in case["sources"] for cid in cited):
        problems.append("wrong article cited")
    return problems

The citation check is unknown_citations from article 04, plus one more rule: the citation has to point at the right article.

The answers need to come from somewhere. By default the example uses a stand-in answerer that pastes the top retrieved chunk as the answer, with its id as the citation. It’s a dumb bot on purpose: free and deterministic, so the whole harness runs offline. python main.py --claude runs the same checks on real answers from claude-opus-5. Here’s the full scorecard with the stand-in:

Kitebase bot: chunks by section, top_k=3, 13 chunks, 12 golden questions
Answers from: offline stand-in (pastes the top chunk)

  question              retrieval  answer
  no-recovery-email     hit        pass
  link-expiry           hit        pass
  ...
  lost-phone            hit        FAIL  missing "backup codes"
  change-sign-in-email  hit        FAIL  missing "Settings > Profile"
  secure-account        50%        FAIL  missing "authenticator app"; wrong article cited
  paraphrase            MISS       FAIL  missing "I can't access my email"; wrong article cited
  bank-transfer         -          pass

  hit rate@3  91%      recall@3  86%
  answers pass 67%      prompt   ~249 tokens per question

Read lost-phone: retrieval is a hit, the answer fails. The backup-codes chunk came back second, and the stand-in only reads the first. That’s what the two separate numbers buy you: a failure with a retrieval hit is an answering problem. bank-transfer passes because every chunk scores below the bot’s minimum score of 0.2, so the bot says “I don’t know” instead of pasting a sign-in chunk.

67% is fine. The point isn’t 100% on day one; it’s that the score never quietly drops.

String checks are cheap, fast and deterministic, and they’re brittle. A real answer that says “half an hour” fails the “30 minutes” check. Pick facts that a correct answer can’t avoid, like a button label, a menu path or a number, and read the failures before you trust them. For qualities a string can’t check, like “is this answer polite and complete?”, you ask another model to grade the answer against a short rubric. That’s LLM-as-judge, and it has its own traps, covered in article 09.

Comparing two versions

A scorecard on its own tells you where you are. The useful part is comparing it with a baseline: the saved scorecard of the version you last accepted. python main.py --save-baseline writes it to baseline.json, and you commit that file. Every later run compares against it and lists each regression, a question that used to pass and now fails. Here’s the teammate’s change from the opening:

$ python main.py --top-k 1
...
vs baseline (chunks by section, top_k=3):
  hit rate       91% -> 64%
  recall         86% -> 64%
  answers pass   67% -> 67%
  prompt tokens  249 -> 160
  WORSE:  lost-phone: retrieval 100% -> 0%
  WORSE:  change-sign-in-email: retrieval 100% -> 0%
  WORSE:  secure-account: retrieval 50% -> 0%

Prompts are about 90 tokens smaller, and three questions lost the chunk that answers them, including the lost phone. The answer score didn’t move only because the stand-in reads nothing past the top chunk anyway. Claude reads every chunk you send, so with real answers those three would suffer.

Now a different change: fixed 50-word chunks instead of one per section, the naive fixed-size chunking from article 05. The answer score is 67% again, same as the baseline. But the per-question diff says something else:

$ python main.py --chunk-words 50
...
  fixed:  lost-phone: answer now passes
  fixed:  change-sign-in-email: answer now passes
  fixed:  secure-account: retrieval 50% -> 100%
  WORSE:  sso-reset: answer now fails, missing "IT admin"
  WORSE:  downgrade-timing: retrieval 100% -> 0%
  WORSE:  downgrade-timing: answer now fails, missing "end of the billing period"
BASELINE TOP_K = 1 50-WORD CHUNKS by section, top_k = 3 what changed what changed retrieval answer retrieval answer retrieval answer no-recovery-email hit pass link-expiry hit pass password-length hit pass failed-attempts hit pass sso-reset hit pass FAIL lost-phone hit FAIL MISS pass download-invoice hit pass downgrade-timing hit pass MISS FAIL change-sign-in-email hit FAIL MISS pass secure-account 50% FAIL MISS hit paraphrase MISS FAIL bank-transfer n/a pass score 91% 67% 64% 67% 82% 67% same as baseline got worse got better The answer score is 67% in all three columns. With 50-word chunks that hides two answers fixed and two broken. Diff per question, not just the totals.
The answer score is 67% in all three versions. The per-question diff is where the changes show.

Two answers fixed, two broken, and the total didn’t move. Compare totals and you’d call it neutral and ship it. So the harness diffs every question, and fails on “anything that passed before and fails now”, not “the total dropped”.

Here’s what broke downgrade-timing:

GOLDEN CASE: DOWNGRADE-TIMING "When does a plan downgrade take effect?" source invoices fact "end of the billing period" BY SECTION (THE BASELINE) invoices#2 score 0.35, rank 1 Invoices and billing: Change your plan Upgrade or downgrade under Admin > Billing > Plan. Upgrades apply straight away. Downgrades apply at the end of the billing period. The whole section is one chunk. retrieval: hit, answer: pass 50-WORD CHUNKS invoices#1 score 0.23, rank 1 ... Change your plan Upgrade or downgrade under Admin > Billing > Plan. Upgrades apply straight away. Downgrades apply at the 50 WORDS, CUT invoices#2 score 0.00, not in the top 3 end of the billing period. retrieval: MISS, answer: FAIL The answer is invoices#1, ending at "apply at the". The golden case names the fact, not a chunk id. After re-chunking, invoices#2 is a different chunk, but "end of the billing period" is the same text, so the same golden set scores both versions.
A fixed-size chunker cut the answer in half. The fact-based label still works after re-chunking.

The chunker cut “Downgrades apply at the” from “end of the billing period.” The second half is a 5-word chunk that shares no words with the question, so it never comes back. You’d never spot that reading the chunking code.

Running it in CI

CI (continuous integration) is the service that runs your checks on every pull request, like GitHub Actions. It treats a non-zero exit code as a failed check, so the harness ends with:

if worse:
    sys.exit(1)  # a non-zero exit is what fails the CI job

The workflow is the same as for any Python test. From evals.yml in the example:

on: pull_request
jobs:
  offline:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-python@v5
        with:
          python-version: "3.12"
      - run: pip install -r requirements.txt
      - run: python main.py  # exits 1 on any regression, which fails the check

Now the --top-k 1 pull request gets a red check listing the three questions it broke. When a change is an improvement, or a trade-off you accept, run --save-baseline and commit the new baseline.json in the same pull request, so the reviewer sees what moved.

Split what you run by cost and by how repeatable it is:

  • Every pull request: retrieval and the offline checks. With a fixed embedder, search gives the same results every run, and it’s free. A failure always means the code changed something.
  • Nightly, and on prompt or model changes: real answers. --claude makes 11 calls (the bank-transfer question never reaches the model), about 250 input tokens and 100 to 200 output tokens each. At claude-opus-5’s $5 per million input and $25 per million output tokens, that’s about 5 cents a run. At 500 questions it’s a couple of dollars, which adds up per pull request.

Real answers vary between runs, so one flipped check isn’t proof of a regression: re-run the failing question a few times and read the answer. For the same reason, the harness doesn’t compare --claude runs with the stand-in baseline.

After you ship

The golden set only covers what you thought to ask. Keep feeding it from production: sample real questions, read the answers users marked unhelpful, and add every failure you find to golden.json. Article 11 covers logging what you need for that.

Try it yourself

The companion example is the whole harness: the Kitebase bot, 12 golden questions, a saved baseline and a CI workflow. It runs offline with no keys.

Download the runnable example (zip)

cd 08-llm-evaluation-pipeline
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

Then try these:

  1. python main.py --top-k 1, then echo $?. The exit code is 1: that’s the red check in CI. Try --top-k 5 and see which numbers move and which don’t.
  2. python main.py --chunk-words 25. Hit rate falls to 64% and answers to 42%. Find the question whose fact got cut in half.
  3. Add a case to golden.json: "Do backup codes work more than once?" with source two-factor and fact "Each one works once". Run pytest -q first: one test checks every labelled fact really is in its article, so a typo fails a test instead of looking like a bot bug. Then save a new baseline.

pip install pytest && pytest -q runs the offline tests. They need no keys and no network.

Common beginner mistakes

  • Only checking by hand. Three questions you already know the answers to aren’t an eval. Write them down, label them, run them every time.
  • Trusting one total. 67% to 67% hid two fixes and two breaks. Diff per question.
  • Only scoring the final answer. Without a retrieval number you can’t tell a search problem from a prompt problem.
  • Labelling chunk ids. They change the moment you re-chunk, and re-chunking is what you wanted to test. Label facts and source articles.
  • Never checking the labels. A typo in a fact looks exactly like a bot failure. Test that every fact appears in its source.

Questions you will face in production

“Can we just use a framework?” Yes, once you know what you’re measuring. RAGAS, Inspect AI and OpenAI’s evals give you runners, reports and ready-made metrics. The golden set is still yours to write, and it’s most of the value. Starting with a small script like this one teaches you what those frameworks do.

“Our answers are free text. What do we check?” Start with what a correct answer can’t avoid: a number, a menu path, a product name, a citation. Add an LLM judge with a short rubric for the rest, and check the judge against your own ratings on a sample first. Article 09 shows how.

“What score should we aim for?” Don’t pick a number up front. Your first scorecard is the baseline, and the rule is that nothing that passes today may fail tomorrow. Raise the bar as fixes land.

Check your understanding

Hit rate@3 is 100% but answers pass only 50%. Where do you look first?

The answering step: the prompt, the model, or how the answer uses the chunks. Search is bringing back the facts. Read a few failing answers next to the chunks they were given.

A pull request changes the chunk size. Hit rate goes from 91% to 91%. Is it safe to merge?

Not yet. Check the per-question diff: one question can be fixed and another broken with the same total. Merge if nothing that passed before now fails, or if you’ve looked at what broke and decided the trade-off is worth it, and saved a new baseline.

Your golden set labels each question with a chunk id like invoices#2. You switch to 50-word chunks. What goes wrong?

invoices#2 now names a different piece of text, a 5-word fragment, so the eval scores against the wrong target. Label the source article and a fact string instead; those don’t change when the chunking does.

A nightly run with real answers fails one question that passed last night. Nothing changed in the code. What do you do?

Re-run that question a few times and read the answers. Model output varies, so a single flip can be noise, or a string check that’s too strict (“half an hour” vs “30 minutes”). If it keeps failing with no code change, the model or provider changed, and the eval caught it.

What to remember

  • An eval is a fixed set of labelled questions, run after every change, scored automatically.
  • Label each question with its source and the facts a correct answer must contain, not chunk ids.
  • Score retrieval (hit rate and recall at k) and answers (facts, citations, “I don’t know”) separately, so a failure tells you where to look.
  • Compare every run with a saved baseline, question by question. Fail on anything that used to pass.
  • Run the cheap, deterministic checks on every pull request and the real-model checks nightly.
  • Grow the golden set from real questions and every bug you find.

What to study next

String checks cover facts and citations. For everything else you need a model to grade answers, and judges are biased in ways that look like bugs in your bot. Article 09: LLM-as-Judge covers how to write a rubric, check a judge against people, and avoid the known biases.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the broad mechanics and specific numbers come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.