Choosing Chunking Strategies: A Practical Framework

In article 04 the Kitebase support bot split its help articles at every ## heading, and the right passage came back first. That split was a choice, and it did more work than it looked.

Split the same six articles every 300 characters instead and the best match for “How do I reset my password if I never set a recovery email?” ends in the middle of a sentence: “Support checks the”. The bot tells the customer to fill in a form and never says what happens next. Split them with the method many tutorials recommend and the best match is a heading with no instructions under it. Same docs, same embedder, same question. Only the chunker changed.

What you’ll build: a script that splits Kitebase’s six help articles four ways, runs the same question against each, and checks whether the top 3 chunks hold the four facts a correct answer needs. It runs offline with no API keys.

What a chunker decides

A chunk is a passage of a document that gets embedded, searched and pasted into the prompt on its own. A chunker is the code that decides where each chunk starts and stops (its boundaries) and how big it can get (the chunk size).

Those two choices decide what the model can see. Search returns whole chunks, so if the answer is split across two and only one comes back, the model gets half an answer, and nothing later in the pipeline can put it back. That’s why article 04 called chunking the decision that matters most.

The worked example is the same one: the question above and the article that answers it, account-recovery.md. It’s 879 characters, about 220 tokens (the unit models read and bill in, roughly 4 characters of English). A correct answer needs four facts from its first section. The companion code checks for them in whatever search returns:

ANSWER_FACTS = ["I can't access my email", "date of your last invoice",
                "one-time sign-in link", "one business day"]

def facts_found(hits: list[tuple[float, Chunk]], facts: list[str] = ANSWER_FACTS) -> list[str]:
    return [f for f in facts if any(f in chunk.text for _, chunk in hits)]

It’s a tiny retrieval test: before any model is involved, did the right text come back? Kitebase’s articles are short, so the example scales sizes down: a 300-character chunk is about 75 tokens. On real documents you’d start around 500 tokens, and the failures are the same.

Here’s where each chunker cuts that file, and what the top 3 held:

Where each chunker cuts account-recovery.md (879 characters) the file §1 If you never set a recovery email §2 SSO §3 lockout title 4 answer facts TOP 3 HOLDS whole file no chunking #1 4 of 4 FACTS but 445 tokens fixed-size 300 chars #1 #2 #3 2 of 4 FACTS 225 tokens fixed + overlap 300 + 60 chars #1 #3 #2 #4 2 of 4 FACTS 225 tokens recursive 300 chars #3 #4 #5 #6 0 of 4 FACTS 73 tokens by section ## headings #1 #2 #3 4 of 4 FACTS 267 tokens cut mid-sentence: "Support checks the / details" each chunk repeats the last 60 characters of the one before #2 is the heading alone, and the top hit at 0.82 each chunk starts "Locked out of your account: ..."
Same file, same question, five chunkers. The right-hand column is what the top 3 chunks held.

Fixed-size chunking

The simplest chunker cuts every N characters, wherever that lands:

def fixed_size(text: str, size: int = CHUNK_SIZE, overlap: int = 0) -> list[str]:
    step = size - overlap
    return [text[start:start + size] for start in range(0, max(len(text) - overlap, 1), step)]

Ignore overlap for now. At 300 characters, account-recovery.md becomes three chunks, and the first cut lands in the middle of the answer:

account-recovery#1 ends:   ...the date of your last invoice. Support checks the
account-recovery#2 starts: details and emails a one-time sign-in link to the address ...

Search still ranks account-recovery#1 first, at 0.60. But the second half of the answer, the part that says support emails you a one-time link within one business day, sits in #2, which ranks 4th. The top 3 holds 2 of the 4 facts:

fixed-size: 13 chunks of 31 to 300 characters, 7 of 34 sentences cut
  0.60  account-recovery#1   # Locked out of your a ... ce. Support checks the
  0.53  recovery-email#1     # Add or change your r ...  > Security > Recovery
  0.53  reset-password#1     # Reset your password  ...  recovery email, see "
  top 3: about 225 tokens, 2 of 4 answer facts (missing: one-time sign-in link, one business day)

“7 of 34 sentences cut” counts sentences, across all six articles, that no chunk holds whole.

Real fixed-size chunkers usually count tokens rather than characters. That changes the unit, not the problem: they still cut through sentences, steps and tables. It’s fine for a first prototype, and rarely what you ship.

Overlap: repeat a little at every cut

The cut sentence is the most fixable part. Overlap means each chunk starts a few characters before the previous one ended, so the text around every boundary appears in two chunks. That’s the overlap argument: with overlap=60, chunks start every 240 characters instead of every 300, and each one repeats the last 60 characters of the one before.

FIXED-SIZE, 300 CHARACTERS, NO OVERLAP account-recovery#1 ends …with your workspace URL and the date of your last invoice. Support checks the account-recovery#2 starts details and emails a one-time sign-in link to the address you give, usually within one business day… "Support checks the details and emails a one-time sign-in link" is whole in neither chunk. SAME, WITH 60 CHARACTERS OF OVERLAP account-recovery#1 ends …Fill in the form with your workspac e URL and the date of your last invoice. Support checks the account-recovery#2 starts e URL and the date of your last invoice. Support checks the details and emails a one-time sign-in… #2 starts 60 characters earlier, mid-word, so the cut sentence is whole in #2. Across all six articles: 7 of 34 sentences cut becomes 0, for 17% more text to embed.
Overlap repeats the text around each cut, so a sentence that straddles one survives whole in the next chunk.

Cut sentences drop from 7 to 0, for one extra chunk and 17% more text to embed and store. The top 3 doesn’t change, though: still 2 of 4 facts. The repaired #2 now holds the whole second half of the answer, and it still ranks 4th. Overlap fixes sentences cut at a boundary. It doesn’t change which chunks win. (Notice #2 now starts at “e URL”, mid-word: overlap is as blind to the text as the cut it repairs.)

The default is 10 to 20% of the chunk size: 50 tokens on a 500-token chunk. Too little and a long sentence still falls through: set OVERLAP = 30 in the example and 3 sentences are cut again. Much more and you’re mostly storing and paying for duplicates.

Recursive chunking

Fixed-size ignores the text. A recursive chunker splits at the biggest natural boundary that works: paragraphs first, then lines, then sentences, then words. It then merges neighbouring pieces back together while they fit the size limit. This is what LangChain’s RecursiveCharacterTextSplitter does, and it’s the default in many tutorials. The core of it:

def recursive(text, size=CHUNK_SIZE, separators=("\n\n", "\n", ". ", " ")):
    if len(text) <= size:
        return [text.strip()]
    sep, *rest = separators
    chunks, current = [], ""
    for piece in split_keeping(text, sep):  # "a\n\nb" -> ["a\n\n", "b"]
        if len(current) + len(piece) <= size:
            current += piece  # still fits: keep merging
            continue
        if current.strip():
            chunks.append(current.strip())
        if len(piece) <= size:
            current = piece
        else:  # too big on its own: split it at the next, smaller boundary
            chunks.extend(recursive(piece, size, tuple(rest)))
            current = ""
    if current.strip():
        chunks.append(current.strip())
    return chunks

The version in main.py also falls back to fixed-size if it runs out of separators. It never cuts a sentence in these articles: 0 of 34. And it does worse than anything else:

recursive: 18 chunks of 21 to 275 characters, 0 of 34 sentences cut
  0.82  account-recovery#2   ## If you never set a recovery email
  0.58  reset-password#1     # Reset your password
  0.51  recovery-email#1     # Add or change your r ... you leave the company.
  top 3: about 73 tokens, 0 of 4 answer facts

The first section, heading plus body, is about 520 characters, too big for one chunk. So the chunker drops down to line breaks, and the heading line becomes a chunk on its own. Every word in it that the embedder counts (never, set, recovery, email) is in the question, so it scores 0.82, the highest score anywhere in the run. It holds none of the answer. The body sits in #3 and #4, which don’t make the top 3. The same thing happens to the title of reset-password.md in second place. The stand-in embedder exaggerates this, since it only counts words, but real embedding models also tend to rank a short heading that restates the question above the longer passage that answers it.

That’s the too-small failure: a chunk so short its few words all match, with nothing behind them. Recursive chunking isn’t wrong; it just doesn’t know a heading belongs to the text under it. Fixes: a bigger size (at 600 characters it gets 4 of 4 here), a minimum size that merges tiny pieces into the next, or a chunker that understands headings.

Splitting by section

Most docs you’ll index have structure: help centers, READMEs, wikis, API references. Headings are the author telling you where one topic ends and the next begins. Section chunking (also called structural chunking) uses them as the boundaries. It’s what article 04 did, plus one safety net:

def by_section(text: str, max_size: int = MAX_SECTION) -> list[str]:
    title, *sections = re.split(r"^## ", text, flags=re.MULTILINE)
    title = title.strip().removeprefix("# ")
    chunks = []
    for section in sections:
        heading, _, body = section.partition("\n")
        prefix = f"{title}: {heading.strip()}\n"
        for piece in recursive(body.strip(), max_size - len(prefix)):
            chunks.append(prefix + piece)
    return chunks

Two things make it work. Every chunk starts with the article title and its heading, so "If you never set a recovery email" also says it’s about being locked out. And a section longer than max_size (here twice the normal chunk size) is split with recursive, with the prefix repeated on every piece, so no chunk loses what it’s about.

by section: 13 chunks of 158 to 515 characters, 0 of 34 sentences cut
  0.61  account-recovery#1   Locked out of your acc ...  doesn't happen again.
  0.51  recovery-email#1     Add or change your rec ... you leave the company.
  0.50  reset-password#1     Reset your password: S ...  out of your account".
  top 3: about 267 tokens, 4 of 4 answer facts

All four facts, in one chunk, in first place. These are the same chunks and scores as in article 04.

The gotcha is that sections vary a lot in size: a real wiki has one-line sections next to 5,000-word ones. The max_size fallback handles the long ones, and short ones are fine with the title prefix. And it all depends on your parser keeping the headings. A PDF through a naive text extractor comes out as one flat wall of text, so check what your parser produces before you pick a chunker.

How big should a chunk be?

Chunk size pulls in two directions, and both ends fail.

The top chunk for the same question, at three sizes TOO BIG: WHOLE FILE account-recovery#1, ~220 tokens # Locked out of your account ## If you never set a… Without a recovery email… ## If your workspace uses… ## If your account is… 0.54 The answer, plus two situations this customer isn't in. RIGHT: ONE SECTION account-recovery#1, ~130 tokens Locked out of your account: If you never set a recovery… Without a recovery email, Kitebase can't send you a reset link. Instead, click… (all four steps follow) 0.61 One situation, title glued on. All 4 answer facts. TOO SMALL: A HEADING recursive chunk #2, ~9 tokens ## If you never set a recovery email 0.82 Nearly every word is in the question. None of the answer is. The highest score went to the chunk with nothing in it. Search measures how alike two texts are, not whether one answers the other.
The highest score went to the chunk with nothing in it.

Too big. The whole-file chunk holds the answer plus two situations this customer isn’t in. Those extra words pull its vector away from the question: 0.54 against the section’s 0.61. The top 3 whole files are also 445 tokens against 267 for sections. Harmless here, but with 3,000-token articles that’s 9,000 tokens of mostly irrelevant context per question.

Too small. The 9-token heading wins on score and carries nothing. Search measures how alike two texts are, not whether one answers the other.

The defaults:

  • Docs with headings: split by section with the title on every chunk, and cap sections at about 1,000 tokens, splitting longer ones recursively.
  • Docs without structure: recursive chunking at about 500 tokens with 50 tokens of overlap.
  • Code, tables and numbered steps: keep each one whole, even if it runs over the size. Half a table or steps 1 to 3 of 6 is worse than a slightly big chunk.

Change the numbers when your tests say so, not before. The ANSWER_FACTS check is the smallest version of that test. The real one is 30 to 50 real questions, each with the chunk that should answer it, counting how often it comes back in the top 3. Article 08 shows how to build it.

What about semantic chunking?

Semantic chunking embeds every sentence and starts a new chunk wherever the meaning of consecutive sentences jumps, so boundaries follow topic changes even with no headings.

It sounds like the best of everything, but it isn’t the default. It needs an embedding for every sentence at index time, a similarity threshold you have to tune, and it can still produce tiny or huge chunks. A 2024 study (Qu et al., linked below) found it wasn’t consistently better than simple fixed-size chunking across the retrieval tasks it tested. Try it when your docs have no structure and recursive chunking measurably fails.

Search small, return big

The recursive heading chunk was the best match in the run. The problem was only what it carried.

Small-to-big retrieval (LangChain calls it a parent document retriever) uses that. You index small chunks for precise matching, and each remembers the bigger chunk it came from, its parent, usually its section. Search runs on the small chunks; the parents go into the prompt. Here, the 0.82 heading match would return the full section and all four facts.

It costs a second store and a lookup step. Reach for it when sections are long: a 2,000-token section as one chunk averages away the paragraph that matters, and small-to-big lets you match that paragraph and still send its context.

Try it yourself

The companion example runs all five chunkers against the same question and prints where each cut, what came back, and how many answer facts made the top 3.

Download the runnable example (zip)

cd 05-chunking-strategies
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py

It uses only the standard library. Then try these:

  1. Set OVERLAP = 30. Cut sentences go from 0 back up to 3: 30 characters isn’t enough to cover a whole sentence at every boundary.
  2. Set TOP_K = 4. Fixed-size now gets all four facts, because the second half of the answer ranked 4th. Recursive gets only two. A bigger k can hide a chunking problem, and it costs tokens on every question.
  3. Set CHUNK_SIZE = 600. Every chunker gets 4 of 4, because each chunk is now most of an article. Look at the token counts: section chunks are still the cheapest context.

pip install pytest && pytest -q runs the offline tests. They check where each chunker cuts, that overlap repeats exactly the right characters, and what each one returns for the worked question.

Common beginner mistakes

  • Taking the library default without looking. A splitter’s default size and separators were picked for someone else’s documents. Print twenty chunks and read them before you embed anything.
  • Splitting headings from their text. A heading on its own matches everything about its topic and answers nothing. Keep each heading with the text under it, and put the document title on every chunk.
  • Cutting tables, code and numbered steps. Half a table or steps 1 to 3 of 6 reads as a complete answer and isn’t. Detect them when you parse and keep each one whole.
  • Tuning chunk size by feel. Changing 500 to 700 because it seems better is guessing. Change one number and rerun the same test questions.

Questions you will face in production

“How do I pick the chunk size?” Start with the defaults above, then build 30 to 50 test questions from real support tickets or search logs, each with the chunk that should answer it. Measure how often that chunk is in the top 3. Try one other size or strategy at a time and keep the change only if the number goes up.

“What happens when a document changes?” Re-chunk it and re-embed only chunks whose text changed; a hash of each chunk’s text next to its vector tells you which. Section chunking helps: editing one section changes one chunk. With fixed-size chunks, a sentence added near the top shifts every boundary after it, so the whole document gets re-embedded.

“Should chunks include context from outside themselves?” Yes, a little. The title prefix is the cheap version. Anthropic’s contextual retrieval is the thorough one: an LLM writes a sentence or two about where each chunk sits in its document, and that’s embedded with the chunk. In Anthropic’s tests it cut retrieval failures by 35%, and by 49% combined with keyword search. It costs one LLM call per chunk at index time. Try the title prefix first.

Check your understanding

Your bot's answers to "how do I export my data?" list steps 1 to 3 of a 6-step guide and stop. Retrieval found the right article. What do you check?

Where the chunk boundaries fall in that guide. A fixed-size or recursive split most likely cut the numbered list, and only the first half ranked in the top k. Keep numbered lists whole, or split by section so the whole guide is one chunk.

You add 20% overlap to a fixed-size chunker and your retrieval test doesn't improve at all. Is overlap useless?

Not useless, just not the problem. Overlap only repairs sentences cut at a boundary, like the “Support checks the / details” cut. If the missing facts sit in a chunk that ranks too low, as in the Kitebase run, overlap can’t help. Look at where the answer is split and what outranks it.

The top-scoring chunk for most questions is a one-line heading. What's going on, and what would you change?

Short chunks whose few words all match the question score very high and carry no answer. Fixes: keep headings attached to the text below them (section chunking), set a minimum chunk size that merges tiny pieces into their neighbour, or use small-to-big retrieval so a heading match returns its whole section.

You're indexing 2,000 scanned PDF contracts with no reliable headings. Which chunker do you start with?

Recursive chunking at about 500 tokens with about 50 tokens of overlap, after checking that the PDF text extraction is readable. Section chunking needs headings you don’t have. Keep clauses and tables whole where your parser can detect them, and test with real questions before changing anything.

What to remember

  • The chunker decides what the model can see. An answer split across chunks becomes half an answer.
  • Fixed-size cuts wherever the count lands. Overlap repairs sentences cut at a boundary but doesn’t change which chunks rank.
  • A chunk can be too small: a lone heading scores highest and answers nothing.
  • Docs with headings: split by section and put the title on every chunk. Docs without: recursive, about 500 tokens, about 50 overlap.
  • Keep tables, code and numbered steps whole, and tune sizes with test questions, not by feel.

What to study next

Even perfect chunks can lose to a bad search. Embedding search struggles with exact strings like error codes and product names, which article 06: Hybrid Search fixes by combining it with keyword search and re-ranking. After that, article 07 covers how to choose the embedding model that scores your chunks.

Further reading

Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the broad mechanics and specific numbers come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.


Auto-marks when you reach the end. Click to toggle.