Build It: A Small RAG Service That Ships

By now the Kitebase support bot exists as a dozen separate scripts. The one from article 04 re-embeds every help article each time it starts. Nothing is cached, so the tenth person to ask how to reset a password costs as much as the first. Nothing is logged, so when a customer pastes a wrong answer into a support ticket you can’t say which prompt produced it, which chunks it saw or what it cost. And the eval harness from article 08 tests a copy of the bot, not the bot you actually run.

Shipping it means putting those pieces into one process with one request path, so what you test, log and cache is what users get.

What you’ll build: “Ask the Docs”, a small HTTP service that answers Kitebase questions from six help articles, cites the sections it used, says “I don’t know” without calling the model when nothing matches, logs one JSON line per request with its cost, caches repeat questions, and has a 10-question eval set that fails CI when quality drops. It runs offline with stand-ins; one API key switches on Claude.

Each part links back to the article that teaches it. What’s new here is the seams between them.

The service on one page

The service has the same two halves as every RAG system. Setup turns the docs into an index and runs when the docs change. Query runs for every question. The worked example follows one question through both: “How do I reset my password if I never set a recovery email?”

Setup: python main.py ingest, whenever data/ changes DATA/ 6 help articles reset-password.md + 5 more INGEST one chunk per ## section, then embed only chunks whose hash is new 13 embedded, 0 reused 04, 05, 07 INDEX.JSON 13 chunks + vectors version 241a3244 Query: POST /ask, for every question. Small grey numbers are the articles that teach each part. POST /ASK "How do I reset my password if I never set a recovery email?" 1. CACHE key = question + prompt fingerprint + model + index version miss 10 2. SEARCH, MIN SCORE 0.2 0.61 account-recovery#1 0.51 recovery-email#1 0.50 reset-password#1 04, 06 3. PROMPT answer@v1 fingerprint f4a8b85c 3 excerpts + question 03 4. CLAUDE claude-opus-5 client passed in, a stand-in offline 02 5. CHECK every cited id was in the prompt [account-recovery#1] 04 ANSWER text + sources account-recovery#1 saved to the cache 6. LOG: ONE JSON LINE PER REQUEST "cache": "miss", "outcome": "answered", "prompt": "answer@v1", "fingerprint": "f4a8b85c", "input_tokens": 402, "output_tokens": 118, "cost_usd": 0.00496, "latency_ms": 1 11, 12
One request through the service. Every value is what the companion code prints.

The companion code has one small file per box:

FileJobTaught in
ingest.pysplit docs into chunks, embed only new ones, save index.json04, 05, 07
retrieve.pyembed the question, search, drop weak hits04, 06
prompts.py + prompts/versioned prompt template with a fingerprint03
llm.pythe Claude call, token counts, cost02, 12
service.pyone request: cache, search, prompt, Claude, citations, logthis article
cache.pyresponse cache10
server.pyPOST /ask, GET /healththis article
evals.py + golden.json10 labelled questions and a pass bar08, 09

The rule that holds it together: AskService.ask in service.py is the only way to get an answer. The HTTP server calls it, the command line calls it, and the evals call it. If the evals used their own copy of the pipeline, they would pass while the real service drifted.

Step 1: Ingest once, not on every request

The bot in article 04 embedded all 13 chunks on every start. That’s free with the stand-in embedder, and slow and paid with a real one.

So ingest saves the index to disk with a hash of each chunk’s text next to its vector. A hash is a short fingerprint of some text: the same text always gives the same hash, and any edit gives a different one. On the next run, a chunk whose hash is already in the index keeps its old vector:

old = json.loads(index_file.read_text()) if index_file.exists() else None
reusable = {}
if old and old["embedder"] == EMBEDDER:
    reusable = {row["hash"]: row["vector"] for row in old["chunks"]}

chunks = load_chunks(data_dir)
hashes = [text_hash(c.text) for c in chunks]
todo = [c for c, h in zip(chunks, hashes) if h not in reusable]
fresh = dict(zip((text_hash(c.text) for c in todo), embed([c.text for c in todo])))

load_chunks is the one-chunk-per-section splitter from article 04. Run ingest twice, then change “30 minutes” to “1 hour” in reset-password.md and run it again:

$ python main.py ingest
{"chunks": 13, "embedded": 13, "reused": 0, "removed": 0, "version": "241a3244"}
$ python main.py ingest
{"chunks": 13, "embedded": 0, "reused": 13, "removed": 0, "version": "241a3244"}
$ python main.py ingest     # after the edit
{"chunks": 13, "embedded": 1, "reused": 12, "removed": 0, "version": "51ab8f9e"}

The index version is a hash of the embedder name plus every chunk hash, so it changes whenever any doc does. The log and the cache key use it to know the docs moved.

The gotcha is the old["embedder"] == EMBEDDER check. Vectors from two embedding models can’t be compared, and mixing them doesn’t crash: the scores just become nonsense (article 07). So the index records its embedder, and the service refuses to start on a mismatch:

The index was built with stand-in-4096, but the code now embeds questions with
text-embedding-3-small. Run: python main.py ingest

Step 2: Retrieve, and let it come back empty

Search is the cosine-similarity search from article 04, with its gotcha fixed: search always finds the top k, even when none of them is any good. The last line keeps only hits above a minimum score:

scored.sort(key=lambda hit: hit.score, reverse=True)
return [hit for hit in scored[:k] if hit.score >= min_score]

For the worked question all three hits clear MIN_SCORE = 0.2: account-recovery#1 at 0.61, recovery-email#1 at 0.51 and reset-password#1 at 0.50. For “Can I pay by bank transfer?” every chunk scores 0.00, search returns nothing, and the service answers “I don’t know” without calling the model. That’s cheaper and safer than hoping Claude notices that three sign-in excerpts don’t answer a billing question.

The 0.2 belongs to the stand-in. A real embedding model scores on a different scale, so when you swap one in, re-pick the minimum from the scores your eval questions get. Adding keyword search (article 06) only changes search; the service doesn’t care how the hits were found.

Step 3: A prompt with a version and a fingerprint

The prompt is the one from article 04, moved out of the code into two files, as article 03 recommends:

prompts/answer/v1/system.txt   rules: answer only from the excerpts, cite [ids], say "I don't know"
prompts/answer/v1/user.txt     $excerpts, a blank line, then Question: $question

load_prompt() reads whichever version ACTIVE in prompts.py points at and gives it two names. The id, answer@v1, is what a person picks. The fingerprint is a hash of the exact text:

@property
def fingerprint(self) -> str:
    return hashlib.sha256((self.system + "\0" + self.user).encode()).hexdigest()[:8]

For v1 it’s f4a8b85c. Two names, because someone will edit v1/system.txt in place “just to fix a typo”. The id stays answer@v1 but the fingerprint changes, so the logs show which exact text produced each answer and the cache stops serving answers written by the old text. Rendering uses Template.substitute, which raises an error on a missing variable instead of sending a literal $question to the model.

Step 4: Claude behind a client you pass in

The model call is the Messages API call from article 02. What matters here is where it lives: in a function that takes the client as an argument.

def ask_claude(client, system: str, user: str, model: str = MODEL) -> Reply:
    response = client.messages.create(
        model=model,
        max_tokens=MAX_TOKENS,
        system=system,
        messages=[{"role": "user", "content": user}],
    )
    text = "".join(block.text for block in response.content if block.type == "text")
    return Reply(text, response.stop_reason,
                 response.usage.input_tokens, response.usage.output_tokens)

The service never creates its own client. main.py passes anthropic.Anthropic(timeout=30.0) when ANTHROPIC_API_KEY is set and StandInClaude() otherwise, and the tests pass fakes with whatever reply a test needs. That’s dependency injection: the caller hands a function the thing it depends on, which is why the whole service runs and tests offline.

StandInClaude pastes the first excerpt it was sent, with its citation, and estimates tokens at 4 characters each. It’s dumb on purpose: free, instant and the same every run. With a key, the text comes from claude-opus-5 and reads like the answer in article 04.

The 30-second timeout replaces the SDK’s 10-minute default, because somebody is waiting. The SDK already retries rate limits and server errors twice with backoff (a longer wait before each retry), so the service adds no retry loop of its own.

Step 5: Check the reply before you show it

A reply from the model is not yet an answer. check_reply in service.py turns it into one of five outcomes, a one-word label for how the request ended:

OutcomeWhenWhat the user getsCached?
answeredevery cited id was in the promptthe answer and its sourcesyes
no_matchnothing scored above the minimum”I don’t know”, no model callyes
bad_citationthe reply cites a chunk it was never sentthe titles of the retrieved articlesno
truncatedstop_reason is "max_tokens"the titles of the retrieved articlesno
refusedstop_reason is "refusal"a short apology and the support emailno

The citation check is the one from article 04:

cited = list(dict.fromkeys(re.findall(r"\[([\w-]+#\d+)\]", reply.text)))
if any(cid not in sent for cid in cited):
    # The model cited a chunk it was never given, so it made at least one thing up.
    return {"answer": fallback, "sources": [], "outcome": "bad_citation"}

If Claude replies “Call us on 555-0100 [phone-support#1]”, there’s no such chunk, so the customer sees “I couldn’t put together a reliable answer. These help articles look relevant: Locked out of your account, …” instead of a made-up phone number.

For the worked question the outcome is answered, and each source carries its title so a UI can link it:

{"answer": "Without a recovery email, Kitebase can't send you a reset link. Instead, click ... [account-recovery#1]",
 "sources": [{"id": "account-recovery#1", "title": "Locked out of your account"}],
 "outcome": "answered", "cached": false}

Step 6: One log line per request, with its cost

When a customer reports a wrong answer, you need to see what the service did for that request. Every call to ask writes one structured log line: a JSON object on a single line, so you can search and total it with ordinary tools instead of reading prose. For the worked question:

{"ts": "2026-09-27T18:52:23+00:00", "question": "How do I reset my password if I never set a recovery email?",
 "prompt": "answer@v1", "fingerprint": "f4a8b85c", "model": "claude-opus-5", "index": "241a3244",
 "chunks": {"account-recovery#1": 0.61, "recovery-email#1": 0.51, "reset-password#1": 0.5},
 "cache": "miss", "outcome": "answered", "input_tokens": 402, "output_tokens": 118,
 "cost_usd": 0.00496, "latency_ms": 2}

(One line in the real log, wrapped here.) chunks with their scores tells you whether retrieval or the prompt is to blame for a bad answer. fingerprint and index say which prompt text and which docs were live. outcome lets you count “I don’t know” answers and bad citations per day.

The cost comes from the token counts the API returns and the model’s price, claude-opus-5 at $5 per million input tokens and $25 per million output tokens:

def cost_usd(model: str, input_tokens: int, output_tokens: int) -> float:
    price_in, price_out = PRICES[model]
    return (input_tokens * price_in + output_tokens * price_out) / 1_000_000

402 input tokens are $0.00201 and 118 output tokens are $0.00295: $0.00496 for this answer. Output costs five times as much per token, which is why short answers matter (article 12).

Two gotchas. Log cache hits and no_match answers too, at a cost of 0, or your cost per question looks worse than it is and you can’t see what the cache saves. And question is customer text that will contain email addresses and worse, so decide how long you keep it and who can read it (article 11 covers PII, traces and dashboards).

Step 7: Cache repeat answers

Support questions repeat. A response cache stores the final answer under a key built from the question, so a repeat skips search and the model call. The hard part is the key. If it’s just the question, the cache keeps serving an answer after whatever produced it has changed:

WHAT GOES INTO THE CACHE KEY question how long does the password reset link last prompt f4a8b85c fingerprint of answer@v1 model claude-opus-5 index 241a3244 hash of every chunk SHA-256 → KEY 88d5bcdcdbe5bbcc answer stored for an hour Then someone changes "30 minutes" to "1 hour" in reset-password.md and runs ingest. KEY INCLUDES THE INDEX VERSION index 241a3244 → 51ab8f9e key db8897aca03b218e: miss Claude answers from the new doc: 1 hour KEY IS THE QUESTION ALONE key unchanged: hit Serves "30 minutes" for up to an hour, with a citation to a doc that now says 1 hour
Anything that can change the answer goes into the key.
def normalize(question: str) -> str:
    return " ".join(question.lower().split()).rstrip("?!. ")

def cache_key(question: str, prompt_fingerprint: str, model: str, index_version: str) -> str:
    raw = json.dumps([normalize(question), prompt_fingerprint, model, index_version])
    return hashlib.sha256(raw.encode()).hexdigest()[:16]

normalize lowercases, collapses spaces and drops trailing punctuation, so trivial typing differences still hit. Entries also expire after an hour, a TTL (time to live), as a backstop for anything the key misses. Here’s what three requests cost, plus the bank-transfer question without a minimum score:

REQUEST 1. CACHE 2. SEARCH 3. CLAUDE COST How do I reset my password if I never set a recovery email? first time anyone asks miss 3 hits top 0.61 402 in 118 out $0.00496 answered how do i reset my password if i never set a recovery email same question, other casing HIT skipped skipped $0 cached Can I pay by bank transfer? no help article covers it miss 0 hits best 0.00 < 0.2 skipped $0 I don't know Can I pay by bank transfer? with MIN_SCORE = 0 miss 3 hits all 0.00 358 in 118 out $0.00474 wrong topic The key lowercases and trims the question, so row 2 finds row 1's answer. Row 4 pays for an answer about password resets. Row 3 pays nothing.
Two ways a request avoids paying for a model call: the cache, and the minimum score.

Only answered and no_match results are cached. A cached truncated reply or bad citation would turn one failure into an hour of them; uncached, the next request tries again.

cache.py is a dict saved to cache.json, which is right for one process. Two servers behind a load balancer need a shared store like Redis. Article 10 covers that, semantic caching, and the provider’s prompt caching, which needs a longer prompt than this one’s 400 tokens.

Step 8: The HTTP API

server.py puts AskService behind two routes using only Python’s standard library:

$ python main.py serve
Serving on http://127.0.0.1:8000 with offline stand-in. Ctrl+C to stop.

$ curl -s -X POST localhost:8000/ask -d '{"question": "Where do I download an invoice?"}'
{"answer": "Go to **Admin > Billing > Invoices** and click an invoice date to download it as a PDF.
 Only workspace admins can see billing. [invoices#1]", "sources": [{"id": "invoices#1",
 "title": "Invoices and billing"}], "outcome": "answered", "cached": false}

$ curl -s localhost:8000/health
{"ok": true, "index": "241a3244", "prompt": "answer@v1 (f4a8b85c)"}

The answer JSON is one line, wrapped here. The handler returns a 400 for a body that isn’t JSON or has no question, and for questions over 500 characters, since a pasted 20-page document is a cost you didn’t plan for. If ask raises after the SDK’s retries, it returns a 503 with a short message, not a stack trace. /health reports the index version and prompt fingerprint, so after a deploy you can check which ones are live.

Python’s docs say http.server isn’t meant for production. It’s fine for this example and for an internal tool behind a VPN. For anything public, move the same handler to FastAPI: the route is about ten lines, because all the work is in AskService.

Step 9: Evals that run the same code path

The eval set is golden.json: 10 questions from the article 08 golden set, each labelled with the article that answers it and a fact a correct answer must contain. One of them, bank transfers, has no answer in the docs, so the only right reply is “I don’t know.”

run_evals builds an AskService with no cache and a list for its log, then asks every question:

records = []  # the service's log lines, so the eval sees retrieval and cost too
service = AskService(index, client, prompt, cache=None, log=records.append)

No cache, because an eval that hits the cache measures the cache. Reading the service’s own log lines lets the eval score retrieval and total the cost without a second code path:

$ python main.py eval
Evals: 10 questions, prompt answer@v1 (f4a8b85c)

  no-recovery-email      hit   pass
  ...
  lost-phone             hit   FAIL  missing "backup codes"
  change-sign-in-email   hit   FAIL  missing "Settings > Profile"
  paraphrase             MISS  FAIL  missing "I can't access my email"; doesn't cite the right article
  bank-transfer          hit   pass

  retrieval 90%   answers 70%   bar 70%   OK
  9 model calls, $0.0238

Two failures are the stand-in reading only the first excerpt when the fact was in another; Claude reads all three. The paraphrase is the real retrieval miss from article 04, which a real embedding model fixes. Bank transfers made no model call, hence 9 calls for 10 questions.

PASS_BAR is today’s score, 70%. Below it the command exits with code 1, which fails a CI job. Raise the bar as the bot improves; never lower it to make a build green. Article 08 compares each question with a saved baseline instead, and article 09 grades what string checks can’t.

Now repeat the doc edit from Step 1 and run the evals. link-expiry fails and the build goes red, though the bot is right: the golden set is out of date. That’s a useful failure. The eval set is part of the docs, and whoever changes a doc updates the fact that tests it.

Try it yourself

The companion example is the whole service. It runs offline, and ANTHROPIC_API_KEY switches the stand-in for claude-opus-5 everywhere, evals included (9 calls, a few cents).

Download the runnable example (zip)

cd 13-capstone-build-rag-service
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py            # ingest, ask twice (the second is cached), ask off-topic
python main.py eval
python main.py serve      # then curl it from another terminal

Then try these:

  1. Change a doc. Run python main.py ask "How long does the password reset link last?" twice; the second is a cache hit. Now edit “30 minutes” to “1 hour” in data/reset-password.md and run python main.py ingest: 1 embedded, 12 reused, a new index version. Ask again: a miss, and the answer says 1 hour. python main.py eval fails on link-expiry until you update its fact.
  2. Change the prompt. Add a line to prompts/answer/v1/system.txt and ask the same question again. The fingerprint changed, so it’s a miss even though the question and the docs didn’t.
  3. Move the minimum score. Set MIN_SCORE = 0.45 in retrieve.py and run the evals: link-expiry, whose best chunk scores 0.44, now gets “I don’t know”. Set it to 0 and ask about bank transfers: the log shows a $0.00474 model call for an answer about password resets.

pip install pytest && pytest -q runs 15 offline tests, including one that starts the HTTP server on a free port and calls it. They need no keys and no network.

What this build skips on purpose

These are the gaps between this and a service for customers, roughly in the order I’d close them:

  • A real embedding model and hybrid search. Swap embed() for text-embedding-3-small, re-ingest, re-pick MIN_SCORE, then add keyword search (article 06). The paraphrase eval tells you when it worked.
  • Auth and rate limits. Before this leaves your network, require an API key and cap requests per key, or anyone can spend your Claude budget.
  • A real database. index.json is searched row by row in memory, fine for thousands of chunks. Past that, use Postgres with pgvector (article 07).
  • A shared cache and real traces. Redis instead of cache.json, and the log lines shipped to a tracing tool (article 10, article 11).
  • Streaming. The endpoint returns the whole answer at once. A chat UI would stream it (article 02).
  • A bigger eval set. Ten questions catch obvious breakage. Add every wrong answer a customer reports, and aim for 50 or more.

Common beginner mistakes

  • Evals that test a copy of the bot. A separate eval script that rebuilds the pipeline passes while the real service drifts. Run evals through the same ask the server calls.
  • Cache keys that ignore the prompt and the docs. The cache keeps serving answers from the old prompt or the old doc until the TTL runs out, with citations that no longer match.
  • Caching failures. A truncated reply or a bad citation, cached for an hour, turns one bad request into hundreds.
  • Logging only model calls. Cache hits and “I don’t know” answers vanish from the numbers, and cost per question looks worse than it is.
  • Re-embedding on every start. It’s invisible with a stand-in and a real bill with a real model. Hash the chunks and embed only what changed.

Questions you will face in production

“A customer says the bot gave a wrong answer. Where do I look first?” Find the log line for that request. If chunks doesn’t contain the section with the right answer, it’s retrieval: chunking (article 05), the embedder or the minimum score. If the right chunk was there and the answer was still wrong, it’s the prompt or the model, and fingerprint tells you which prompt text to look at. Either way, add the question to golden.json so it can’t come back unnoticed.

“Cost is creeping up. What’s the cheapest thing to check?” The cache hit rate in the logs. If many misses differ only in wording, better normalization or a semantic cache helps (article 10). If output_tokens is high, ask for shorter answers. The rest is in article 12.

“How do I know a model upgrade is worth it?” Change MODEL in llm.py, run python main.py eval, and compare the pass rate and the cost line with the current model’s. Changing MODEL also changes the cache key, so nobody gets an answer from the old model after the switch.

Check your understanding

You fix a typo in prompts/answer/v1/system.txt and deploy. Why don't users keep getting cached answers from the old text?

The cache key includes the prompt fingerprint, a hash of the exact template text. The typo fix changes the fingerprint, so every key changes and the first request for each question is a miss. The id is still answer@v1, which is why the fingerprint exists.

The eval pass rate drops from 70% to 60% right after someone edits a help article. What happened, and what's the right fix?

Probably a golden fact no longer matches the doc, like link-expiry expecting “30 minutes” after the doc changed to “1 hour”. Check the failing question’s answer against the doc. If the bot is right, update golden.json. Don’t lower PASS_BAR.

The logs show a request with outcome no_match, cost 0, and an empty chunks map. The docs do cover the question. What do you check?

Retrieval: every chunk scored under MIN_SCORE. Run the question through search with min_score=0 and look at the scores. If the right chunk is just under the bar, the bar is too high for this embedder. If it’s far down the list, the question is phrased differently from the doc, which a real embedding model or keyword search fixes.

Why does the eval run with cache=None?

With a cache, the second eval run would replay stored answers instead of running search and the model, so a prompt or retrieval change could break the bot and the evals would still pass. Evals measure the pipeline; the cache is tested on its own.

What to remember

  • One request path: the API, the CLI and the evals all call the same ask, so what you test is what you serve.
  • Ingest when docs change, not on every start, and re-embed only chunks whose hash changed. Record which embedder built the index.
  • Put everything that can change an answer into the cache key: question, prompt fingerprint, model, index version. Never cache failures.
  • Check the reply before showing it: missing matches, refusals, truncation and invented citations each get a safe fallback.
  • Log one JSON line per request, cache hits included, with the chunks, the prompt fingerprint, the tokens and the cost.
  • Every bad answer a customer finds becomes a new eval question. Over time the eval set is worth more than the code.

What to study next

That’s the AI Engineering topic. The next step is to point this at something real and grow golden.json from real users’ questions:

  • Your team’s internal docs. Swap data/ for a Confluence or Notion export and adapt load_chunks to its format. Everything after it stays the same.
  • A codebase. Split by function instead of by heading and ask “where do we handle Stripe webhooks?”
  • Old support tickets. Index resolved tickets and surface similar ones, with a draft reply, for each new ticket.

If you want Claude to use tools and live data instead of only your docs, the MCP Development topic starts with What Is MCP?.

Further reading

  • Python docs: http.server. The module server.py uses, including the warning about production use.
  • FastAPI documentation. Where to move the handler when the service goes beyond an internal tool.
  • pgvector. Vector search in Postgres, the usual next step after an in-memory index.
  • OpenTelemetry for Python. For turning the per-request log line into traces once you have a tracing backend.

Where this article comes from. This is the working pattern most small-team RAG services converge on after six months of production. It is not a citation of any single paper, it is a synthesis. The code samples are real Python you can copy. If a snippet fails or a library version moves, the article gets fixed within a day; send me a note.


Auto-marks when you reach the end. Click to toggle.