RAG Explained: How to Build an AI That Answers Questions From Your Documents
You’re adding a support bot to Kitebase, a small project-tracking app. A customer writes: “How do I reset my password if I never set a recovery email?” You send that straight to a model. It has never seen Kitebase’s help center, so it answers the way most apps work: click “Forgot password” and follow the link in your email. That’s the one path that can’t work for this customer, because the link goes to the recovery email they never set.
The model isn’t broken. It fills the gap with the most likely-sounding text, the hallucination problem from article 01. The fix is to find the right paragraph in your own docs and hand it to the model with the question. That pattern is called RAG, and it’s the first non-trivial thing most teams ship with LLMs.
What you’ll build: a support bot that answers from six short Kitebase help articles, shows which passages it picked and how they scored, and cites the one it used. It runs offline; two API keys upgrade it to a real embedding model and a real answer from Claude.
What RAG is
RAG stands for Retrieval Augmented Generation, and the name is the recipe:
- Retrieve: search your docs for the passages most related to the question.
- Augment: paste those passages into the prompt.
- Generate: the model writes the answer from them, not from memory.
It’s a search engine whose results go into a prompt instead of onto a page.
Why not paste every doc into every prompt? For Kitebase’s six articles you could: they come to under 1,000 tokens (the unit models read and bill in, roughly 4 characters of English). A real help center with 500 articles of about 400 tokens each is 200,000 tokens per question. At claude-opus-5’s $5 per million input tokens, that’s $1.00 per question before the model writes a word. The prompt RAG builds for the Kitebase question is about 400 tokens, or $0.002.
Why not train the model on your docs instead?
That’s fine-tuning: training a model further on your own examples. It’s the wrong tool for facts that change.
- When a doc changes, RAG re-indexes one file. Fine-tuning means retraining.
- RAG can show which passage an answer came from. Knowledge baked into a model can’t be cited or checked.
- RAG can filter what each user is allowed to retrieve. A model can’t un-learn the HR doc some users shouldn’t see.
Two pipelines: setup and query
A RAG system is two pieces of code that run at different times. Setup prepares your docs, once and again whenever they change. Query runs for every question, so it has to be fast; all the slow work on your docs has already happened.
Step 1: Split the docs into chunks
Kitebase’s “Locked out of your account” article covers three situations: no recovery email, single sign-on, and too many failed attempts. The customer needs the first. Retrieve the whole article and the model gets the other two as noise.
So you split each document into chunks: passages small enough to be about one thing. Chunks, not documents, are what you search and what goes into the prompt. Kitebase’s help articles have a ## heading per situation, so the example makes one chunk per section:
def load_chunks(data_dir: Path = DATA_DIR) -> list[Chunk]:
chunks = []
for path in sorted(data_dir.glob("*.md")):
title, *sections = re.split(r"^## ", path.read_text(), flags=re.MULTILINE)
title = title.strip().removeprefix("# ")
for n, section in enumerate(sections, start=1):
heading, _, body = section.partition("\n")
text = f"{title}: {heading.strip()}\n{body.strip()}"
chunks.append(Chunk(id=f"{path.stem}#{n}", text=text))
return chunks
Six files become 13 chunks of roughly 40 to 130 tokens. The one that matters, trimmed:
id: account-recovery#1
Locked out of your account: If you never set a recovery email
Without a recovery email, Kitebase can't send you a reset link. Instead, click
**Forgot password?** and then **I can't access my email**. Fill in the form ...
Two details do real work. The id (file name plus section number) is what the model will cite and what your code will check. And the article title is glued onto every chunk, because the heading “If you never set a recovery email” on its own never says it’s about being locked out.
Real documents are messier: strip HTML menus with trafilatura, pull PDF text with pdfplumber, and start long documents at about 500-token chunks. Chunking is the decision that matters most: if the answer is cut in half or buried in a huge chunk, nothing later can fix it. Article 05 shows how to choose.
Step 2: Turn each chunk into a vector
Now you need to rank chunks by how well they match a question, fast, even across 500,000 of them. Comparing strings misses that “locked out” and “can’t sign in” mean the same thing.
An embedding is a fixed-length list of numbers (a vector) that stands for a piece of text. It’s like a hash, except similar inputs get similar outputs. An embedding model does the conversion: OpenAI’s text-embedding-3-small turns any text into 1,536 numbers, arranged so texts with similar meaning get similar vectors.
By default the companion example uses a stand-in instead, so it runs with no account and you can see every number. It isn’t a real embedding model. It drops filler words (“how”, “do”, “my”), hashes each remaining word to one of 4,096 slots, and counts:
def stand_in_embed(texts: list[str]) -> list[list[float]]:
vectors = []
for text in texts:
vec = [0.0] * DIMENSIONS # 4,096
for word in words(text):
slot = int(hashlib.md5(word.encode()).hexdigest(), 16) % DIMENSIONS
vec[slot] += 1.0
norm = math.sqrt(sum(x * x for x in vec)) or 1.0
vectors.append([x / norm for x in vec])
return vectors
The last two lines scale each vector to length 1, so a long chunk doesn’t win just by having more words. Here’s the question and three of the chunks:
The red box is why real embedding models exist. “No backup address” and “never set a recovery email” are the same problem in different words, so word counting ranks the right section 7th. A real model is trained to put those phrases close together. Set OPENAI_API_KEY and the example switches to text-embedding-3-small, at $0.02 per million tokens: indexing all 13 chunks costs about $0.000015.
One rule holds for any embedder: embed chunks and questions with the same model. Vectors from different models can’t be compared, and the scores still look plausible. Change models, re-embed everything.
How does a real model know that "backup address" means "recovery email"?
It learned it. An embedding model is trained on huge numbers of text pairs that belong together (a question and its answer, two phrasings of one idea) and pairs that don’t, and rewarded for putting related pairs close together.
After training, every text maps to a point in a space with 1,536 dimensions. You can’t picture it, but the math is the same as distance on a map: nearby points mean similar things. Article 07 goes further.
Step 3: Search for the closest chunks
Search means: embed the question with the same embedder, score it against every stored vector, and keep the top few. The score is cosine similarity, how closely two vectors point the same way: 1.0 is the same direction, 0 is nothing in common.
def cosine(a: list[float], b: list[float]) -> float:
dot = sum(x * y for x, y in zip(a, b))
norms = math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b))
return dot / norms if norms else 0.0
def search(question: str, index: Index, embed, k: int = TOP_K) -> list[tuple[float, Chunk]]:
[q_vec] = embed([question])
scored = [(cosine(q_vec, vec), chunk) for chunk, vec in index]
scored.sort(key=lambda pair: pair[0], reverse=True)
return scored[:k]
For the Kitebase question it prints:
Top 3 chunks:
0.61 account-recovery#1 Locked out of your account: If you never set a recovery email
0.51 recovery-email#1 Add or change your recovery email: Why you need one
0.50 reset-password#1 Reset your password: Send yourself a reset link
The right section comes first, with related background behind it. The index here is a Python list, and scoring 13 vectors is instant. At 500,000 chunks you’d use a vector database, storage that finds the nearest vectors without scoring every row. Postgres with the pgvector extension is the usual first choice.
The gotcha is in the last line of search: it always returns k results, even when none of them is any good.
No article covers bank transfers, so Claude gets three sign-in excerpts for a billing question. The prompt in the next step tells it to say it doesn’t know, but it’s cheaper and safer to catch this first: drop hits below a minimum score, and skip the model call when nothing is left. The cut-off depends on the embedder. 0.2 suits the stand-in and means nothing for text-embedding-3-small, so pick it from the scores real questions get.
Real embedding models have the opposite weakness to the stand-in: good at meaning, weak on exact strings like error codes and SKUs. Production systems run keyword and vector search together and merge the results. That’s hybrid search, in article 06.
Step 4: Build the prompt
The prompt has three parts: a system prompt with the rules (instructions the model follows for the whole conversation), the retrieved chunks, and the question.
You are the support assistant for Kitebase, a project-tracking app.
Answer using only the help-center excerpts in the user's message.
After each fact, cite the excerpt it came from in square brackets, like [reset-password#1].
If the excerpts don't contain the answer, say you don't know and suggest emailing support@kitebase.example.
The user message from build_prompt, trimmed:
<excerpt id="account-recovery#1">
Locked out of your account: If you never set a recovery email
Without a recovery email, Kitebase can't send you a reset link. Instead, click ...
</excerpt>
<excerpt id="recovery-email#1"> ... </excerpt>
<excerpt id="reset-password#1"> ... </excerpt>
Question: How do I reset my password if I never set a recovery email?
“Only the excerpts” stops the model mixing in how other apps work. The <excerpt id="..."> tags mark where each passage ends and give it an id to cite. “Say you don’t know” offers a way out that isn’t guessing. The whole prompt is about 400 tokens. Article 03 covers keeping prompts like this in versioned templates.
Step 5: Ask Claude, then check the citations
The call is an ordinary Messages API call, as in article 02:
def ask_claude(client, question: str, hits: list[tuple[float, Chunk]]) -> str:
response = client.messages.create(
model="claude-opus-5",
max_tokens=1024,
system=SYSTEM_PROMPT,
messages=[{"role": "user", "content": build_prompt(question, hits)}],
)
if response.stop_reason == "refusal":
return "Sorry, I can't help with that one. Please email support@kitebase.example."
return "".join(block.text for block in response.content if block.type == "text")
Check stop_reason before trusting the text: "refusal" means the model declined, and "max_tokens" means the answer was cut off, which the full version in main.py flags. With ANTHROPIC_API_KEY set you get something like this (the wording changes between runs; the facts shouldn’t):
Without a recovery email, Kitebase can't send you a reset link [account-recovery#1].
Instead, click **Forgot password?** on the sign-in page, then **I can't access my email**,
and fill in your workspace URL and the date of your last invoice. Support checks the
details and emails you a one-time sign-in link, usually within one business day
[account-recovery#1]. Once you're in, set a new password under **Settings > Security**
and add a recovery email so this doesn't happen again [account-recovery#1].
Compare that with “follow the link in your email” from the opening. Every step comes from the help center, and each points to its source.
Grounding cuts hallucination a lot but doesn’t remove it. The cheapest check is that every cited id was actually in the prompt:
def unknown_citations(answer: str, hits: list[tuple[float, Chunk]]) -> list[str]:
sent = {chunk.id for _, chunk in hits}
return [cid for cid in re.findall(r"\[([\w-]+#\d+)\]", answer) if cid not in sent]
An answer citing [phone-support#1] cites a chunk that doesn’t exist, so don’t show it as is. A real id can still sit next to a misquoted fact, so this catches the obvious failures, not all of them. Measuring answers properly is article 08.
The next upgrade: re-ranking
Search compares the question’s vector with each chunk’s vector separately. It never reads the two together, so the best chunk sometimes lands 7th instead of 1st. Even here, recovery-email#1 edges out reset-password#1 by 0.01, though the reset instructions are more useful.
A re-ranker is a second model that reads the question and one chunk together and scores how well that chunk answers it. It’s too slow to run on every chunk, so search returns the top 20 to 50, the re-ranker re-scores those, and the best 3 to 5 go into the prompt. In Anthropic’s contextual retrieval experiments, adding a re-ranker cut retrieval failures (the right chunk missing from the top 20) from 2.9% to 1.9%. It costs an extra model call per question. The companion example skips it: with 13 chunks, search already puts the right one first.
Try it yourself
The companion example is the whole bot. It runs offline and prints the prompt it would send; the two API keys switch on the real parts.
Download the runnable example (zip)
cd 04-rag-system-design
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
Then try these:
python main.py "Can I pay by bank transfer?". Every score is 0.00 and you still get three chunks. Add a minimum score of 0.2 insearchand skip the Claude call when nothing is left.python main.py "I'm locked out and have no backup address". The stand-in ranks the right section 7th. SetOPENAI_API_KEYand ask again.- Set
TOP_K = 13so every chunk goes into the prompt. It grows from about 401 tokens to 997: fine for six articles, $1.00 a question at 500.
pip install pytest && pytest -q runs the offline tests. They need no keys and no network.
Common beginner mistakes
- Chunks that are too big or too small. A 5,000-token chunk buries the answer. A 20-token chunk like “See section 4.2” matches lots of questions and answers none.
- Only using vector search. It misses product names, error codes and SKUs. Add keyword search.
- Trusting the parser. Retrieving the right chunk does nothing if a PDF table turned it into word soup. Read a sample of chunks before you embed.
- No way to say “I don’t know.” Search always returns something. Add a minimum score and a prompt that allows “I don’t know.”
- No evaluation. Write 30 to 50 real questions with the chunk that should answer each, and check how often it comes back in the top 3. Article 08 shows how.
Questions you will face in production
“What if you have 10 million documents?”
Start with Postgres and pgvector: one less system to run, and it goes further than most teams need. Move to a dedicated vector database (Qdrant, Pinecone, Turbopuffer) when your own load tests say Postgres is the bottleneck.
“How do you handle permissions?” Tag every chunk with the groups allowed to see it and filter in the search query. Never retrieve everything and ask the model to hide what the user can’t see: anything in the prompt can end up in the answer.
“How do you keep the index fresh?” Re-run setup when docs change, re-embedding only chunks whose text changed. Store a hash of each chunk’s text next to its vector (article 10 calls this an embedding cache), and delete chunks for removed pages or the bot keeps quoting them.
Check your understanding
Your bot answers "I don't know" to a question that's clearly covered in the docs. What do you look at first?
The retrieved chunks and scores for that question, before touching the prompt. If the right chunk isn’t in the top k, it’s a retrieval problem (chunking, the embedder, or wording). If it’s there and the model still says “I don’t know”, it’s a prompt problem.
You switch from the stand-in to text-embedding-3-small and keep the 0.2 minimum score. What could go wrong?
Different embedders score on different scales, so 0.2 might keep everything or drop good matches. Re-pick it from the scores real questions get. And re-embed every chunk, because the old vectors came from a different model.
A customer asks "What does error KB-4012 mean?" and vector search returns sign-in chunks, though a troubleshooting page lists KB-4012. Why, and what's the fix?
“KB-4012” has no meaning for an embedding model to match on. Keyword search would find the literal string, so run both and merge the results: hybrid search, from article 06.
An answer ends with [reset-password#3]. What happened, and what should your code do?
reset-password.md only has two sections, so the model invented that citation. unknown_citations flags it. Don’t show the answer as is: retry once or fall back to showing the retrieved articles, and log it.
What to remember
- RAG is search plus a prompt: retrieve the relevant chunks, put them in the prompt, have the model answer only from them.
- Setup (split, embed, store) runs when docs change. Query (embed, search, prompt, answer) runs per question.
- Embeddings make similar meaning give similar vectors. Use the same model for chunks and questions.
- Search always returns k results. Add a minimum score and give the model a way to say “I don’t know.”
- Give every chunk an id, ask for citations, and check that every cited id was sent.
What to study next
Chunking is where most RAG quality is won or lost, so it comes next: article 05: Choosing Chunking Strategies covers the main ways to split documents and how to pick a size. After that, article 06 adds keyword search and re-ranking, and article 07 covers choosing an embedding model.
Further reading
- Lewis et al. (2020): Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. The original RAG paper. Worth reading for the framing.
- Anthropic: Introducing Contextual Retrieval. Adds a short note about the surrounding document to each chunk before embedding. The source of the re-ranking numbers above.
- Pinecone Learning Center. Practical guides on vector search, RAG and retrieval evaluation.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the broad mechanics and specific numbers come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.