Embeddings Deep Dive
The Kitebase support bot from article 04 handles “How do I reset my password if I never set a recovery email?” well: the right help section comes first. Then a customer types “I’m locked out and have no backup address”. Same problem, different words. The bot’s offline stand-in embedder ranks the right section 7th of 13 and puts “Change your sign-in email” first.
Article 04 said a real embedding model does better with paraphrases and left it there. This article is the why, plus the decisions a real one brings: which model, how many numbers per vector, what switching costs, and how to store a million vectors.
What you’ll build: a script that scores the locked-out paraphrase by hand on 3-number vectors, benchmarks the stand-in on 8 labelled Kitebase questions (it gets 6), and refuses to search when the question and the index came from different models.
What an embedding is
String matching can’t tell that “no backup address” and “never set a recovery email” are the same complaint. You need a form of text where meaning decides what’s close, not spelling.
An embedding is that: a fixed-length list of numbers (a vector) that stands for a piece of text. It’s like a hash, except similar inputs get similar outputs. An embedding model makes them. You send text, you get floats back:
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from the environment
response = client.embeddings.create(
model="text-embedding-3-small",
input=["I'm locked out and have no backup address"],
)
vector = response.data[0].embedding
print(len(vector))
print(round(sum(x * x for x in vector), 3))
1536
1.0
Three things to notice. The length is always 1,536, whether you send six words or six pages (up to 8,192 tokens, the pieces of text models read and bill in, about 4 characters each). The squares sum to 1, because OpenAI scales every vector to length 1, which matters in the next section. And no position has a name: position 412 doesn’t mean “email”. Meaning is spread across all of them.
The meaning comes from training. An embedding model sees huge numbers of text pairs that belong together (a question and its answer, two phrasings of one idea) and pairs that don’t, and is rewarded for putting the first kind close and the second far apart. That’s contrastive learning, and it’s how “backup address” can end up near “recovery email” without sharing a word.
Two consequences. A model is best at the kind of text it trained on (web text, code, one language). And every model builds its own space, so its vectors mean nothing to another model. That one causes the worst bug in this article.
Cosine similarity, worked by hand
You have a vector for the question and one per chunk (a passage of a document, the unit you search; see article 05). Now you need one number for “how close are these two?”
The usual answer is cosine similarity: picture each vector as an arrow from the origin, and measure the angle between two arrows. Same direction scores 1.0. At right angles, 0. The length of the arrows doesn’t count.
1,536 numbers are hard to check by hand, so here are 3, made up so each position means something: being locked out, an email address, billing. A real model has no named positions.
| Text | Vector |
|---|---|
| question: “I’m locked out and have no backup address” | [4, 3, 0] |
| account-recovery#1: “If you never set a recovery email” | [3, 4, 0] |
| change-email#1: “Update the address” | [0, 5, 0] |
| invoices#1: “Download an invoice” | [0, 0, 5] |
The formula is the dot product (multiply matching positions, add them up) divided by the two lengths multiplied together. A vector’s length is the square root of the sum of its squares, so the question’s is √(16 + 9 + 0) = 5, and so is every chunk’s here.
- account-recovery#1: (4·3 + 3·4 + 0·0) / (5 × 5) = 24 / 25 = 0.96
- change-email#1: (4·0 + 3·5 + 0·0) / (5 × 5) = 15 / 25 = 0.60
- invoices#1: (4·0 + 3·0 + 0·5) / (5 × 5) = 0 / 25 = 0.00
In code, from the companion example:
def dot(a: list[float], b: list[float]) -> float:
return sum(x * y for x, y in zip(a, b))
def length(v: list[float]) -> float:
return math.sqrt(dot(v, v))
def cosine(a: list[float], b: list[float]) -> float:
# zip() would quietly compare the first 1,536 numbers of a 4,096-number vector.
if len(a) != len(b):
raise ValueError(f"can't compare a {len(a)}-number vector with a {len(b)}-number one")
lengths = length(a) * length(b)
return dot(a, b) / lengths if lengths else 0.0
It prints the same arithmetic, plus two lines worth reading:
same direction, twice as long [6, 8, 0]: cosine 0.96, dot product 48
scaled to length 1, [0.8, 0.6, 0.0] . [0.6, 0.8, 0.0] = 0.96: the dot product is the cosine
The first line is why cosine ignores length: a chunk that says the same thing twice as loudly shouldn’t jump the queue. Its dot product doubled; its cosine didn’t move. The second is why OpenAI and Voyage scale their vectors to length 1: the division then does nothing, and the dot product alone is the cosine. pgvector has a separate inner-product operator (<#>) for exactly that case.
Why cosine and not plain distance between the points?
For vectors of length 1 it makes no difference: Voyage’s docs note that cosine and Euclidean (straight-line) distance rank their normalized vectors identically.
Cosine is the default because it still ignores length when vectors aren’t normalized, like the toy ones above. Set your vector database to the metric your model’s docs recommend.
The gotcha: a score only means something next to scores from the same model. The stand-in gave account-recovery#1 0.14, the toy vectors 0.96, and text-embedding-3-small will give a third number. Pick any minimum-score cut-off per model, from the scores real questions get.
Why the stand-in misses the paraphrase
Here’s what the stand-in does with the locked-out question (step 2 of main.py):
2. stand-in (4,096 numbers per vector): "I'm locked out and have no backup address"
#1 0.23 change-email#1
#2 0.21 account-recovery#3
#3 0.19 account-recovery#2
... account-recovery#1 is #7 of 13 at 0.14
The stand-in gives each word its own slot, so texts only score when they share words. The question shares “address” with “Update the address”, so that wins. “Backup” and “recovery” land in different slots and never meet.
A trained model has no slot per word. It learned from pairs of text that “backup address” and “recovery email” do the same job, so it can point them the same way, as the toy vectors do by design. That’s the whole difference.
Set OPENAI_API_KEY and the script runs the same question through text-embedding-3-small. Look at where account-recovery#1 lands. That check on your data beats any general claim about the model.
Dimensions
Each number in the vector is a dimension. Models range from 384 (all-MiniLM-L6-v2, a small open model) through 1,024 (Voyage’s default) and 1,536 (text-embedding-3-small) to 3,072 (text-embedding-3-large).
More dimensions give the model more room to separate close meanings. They also cost you: every vector is stored, held in memory and compared at that size. Doubling dimensions doubles storage.
You don’t always have to take the full size. The text-embedding-3 models accept a dimensions parameter that returns a shorter vector:
client.embeddings.create(model="text-embedding-3-small", input=texts, dimensions=512)
That’s 3x less to store. OpenAI’s docs report that text-embedding-3-large shortened to 256 dimensions still beats their previous model, text-embedding-ada-002, at its full 1,536, on the MTEB benchmark. Whether 512 is enough for your questions is what the benchmark in the next section answers.
How can you cut numbers off a vector and keep it working?
These models are trained so the first dimensions carry the most important information and later ones add detail. It’s called Matryoshka representation learning, after the nested dolls: each shorter prefix is a usable, smaller embedding.
Only models trained this way survive it; cut an ordinary model’s vector in half and you just lose half. If you shorten a vector yourself instead of via the API, scale it back to length 1, as OpenAI’s docs point out, or the dot-product shortcut breaks.
The gotcha is in your database. Postgres’s pgvector extension can store vectors up to 16,000 dimensions, but its fast HNSW index (the structure that finds near neighbours without scanning every row) takes at most 2,000 for its standard vector type. text-embedding-3-large at 3,072 won’t index. Shorten it with dimensions, or use pgvector’s half-precision halfvec type, which indexes up to 4,000.
Choosing a model
There are dozens of embedding models and a public leaderboard, MTEB (Massive Text Embedding Benchmark), that ranks them across many tasks. It’s good for a shortlist. It isn’t your data.
Start here:
| Model | Dimensions | Price per 1M tokens | Pick it when |
|---|---|---|---|
text-embedding-3-small | 1,536, shortenable | $0.02 | The default. Cheap, and a good first model to benchmark. |
text-embedding-3-large | 3,072, shortenable | $0.13 | Your benchmark shows small missing questions that large gets. |
Voyage voyage-4 family | 1,024 (256 to 2,048) | See Voyage’s pricing page | You want the provider Anthropic’s docs point to (Anthropic has no embeddings API), or their code, legal or finance models. |
all-MiniLM-L6-v2 | 384 | Free, runs on your machine | Text can’t leave your servers. Keep chunks short: it cuts input off after 256 word pieces. |
For scale: OpenAI reports an MTEB average of 62.3% for small and 64.6% for large, 2.3 points for 6.5 times the price. On your questions the gap might be bigger, smaller or zero.
Benchmark on your own questions. Write down real questions, each with the chunk that answers it, and count how often that chunk comes back in the top k. That count, as a fraction, is recall@k (recall at k). The companion has 8 Kitebase questions:
LABELLED = [
("How do I reset my password if I never set a recovery email?", "account-recovery#1"),
("I'm locked out and have no backup address", "account-recovery#1"),
("Where can I get a PDF of last month's bill?", "invoices#1"),
# ...5 more
]
def benchmark(index: Index, embed, labelled=LABELLED) -> list[tuple[str, int]]:
"""For each labelled question, the rank its right chunk came in at."""
vectors = embed([q for q, _ in labelled])
return [(q, rank_of(search(index, v, index.model), want))
for (q, want), v in zip(labelled, vectors)]
3. Benchmark: right chunk in the top 3? (8 labelled questions)
stand-in: 6/8
ok #1 How do I reset my password if I never set a recovery email?
MISS #7 I'm locked out and have no backup address
ok #1 Where can I get a PDF of last month's bill?
...
MISS #6 I typed my password wrong too many times
Recall@3 is 6/8 for the stand-in, and both misses are paraphrases. With OPENAI_API_KEY set, text-embedding-3-small gets scored on the same 8. Run every candidate through the same list and pick the cheapest one that clears your bar.
Eight questions show the idea but are too few to decide on. Use 50 or more from real tickets or search logs, not ones you wrote to match the docs. Article 08 turns this into an eval you run on every change.
Switching models means re-embedding everything
Your benchmark says a new model wins. The tempting move is to change the model name in the query code and ship. Every stored vector is still from the old model, though, and the two spaces don’t line up.
The companion shows it with a second stand-in that hashes words with a different salt: same 4,096 dimensions, different “model”.
5. Same dimensions, different model (stand-in with another hash salt)
guard: index was built with stand-in, question embedded with stand-in-v2
without the guard, top 3: invoices#2 0.14, account-recovery#1 0.00, account-recovery#2 0.00
A password question gets “Change your plan” as its top hit, and nothing errors. With real models the wrong scores look like ordinary scores, which is worse. If the dimensions differ, zip compares the first 1,536 of 4,096 numbers and returns a number anyway, which is why cosine checks lengths.
The fix is to store the model name with the index and check it on every search:
def search(index: Index, query_vector: list[float], query_model: str) -> list[tuple[float, Chunk]]:
if query_model != index.model:
raise ValueError(f"index was built with {index.model}, question embedded with {query_model}")
scored = [(cosine(query_vector, v), c) for c, v in zip(index.chunks, index.vectors)]
return sorted(scored, key=lambda pair: pair[0], reverse=True)
The money is rarely the problem: a million chunks of 500 tokens is 500 million tokens, $10 on text-embedding-3-small. The real costs are the hours the backfill takes against the provider’s rate limits, and holding two copies of every vector meanwhile.
Storing a million vectors
One vector from text-embedding-3-small is 1,536 numbers. Stored as standard 4-byte floats (float32), that’s 6,144 bytes, 6 KB. Kitebase’s 13 chunks are 80 KB. A million chunks is 6.1 GB before any search index:
4. Storage for 1,000,000 vectors (vectors only, no index)
dimensions float32 float16 int8
384 1.5 GB 0.8 GB 0.4 GB
1,536 6.1 GB 3.1 GB 1.5 GB
3,072 12.3 GB 6.1 GB 3.1 GB
The other columns are smaller number formats: float16 is 2 bytes a number (pgvector’s halfvec), int8 is 1 byte. Voyage can also return binary, one bit per dimension, 32x smaller than float32. Each step down loses precision; your benchmark tells you whether recall notices.
For most teams the right store is the Postgres you already run, with pgvector:
CREATE EXTENSION vector;
CREATE TABLE chunks (
id text PRIMARY KEY, -- 'account-recovery#1'
body text NOT NULL,
embedding_model text NOT NULL, -- 'text-embedding-3-small'
embedding vector(1536) NOT NULL
);
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);
SELECT id, 1 - (embedding <=> $1) AS cosine
FROM chunks
ORDER BY embedding <=> $1
LIMIT 3;
<=> is pgvector’s cosine distance, which is 1 minus the similarity, so smaller is closer. The embedding_model column is the guard from the last section, in SQL. The vector(1536) column also refuses a vector of the wrong length, which catches the zip bug for free.
Try it yourself
The companion runs every output in this article offline: the hand-worked cosines, the paraphrase ranking, the 8-question benchmark, the storage table and the model guard. Set OPENAI_API_KEY and text-embedding-3-small joins the ranking and the benchmark.
Download the runnable example (zip)
cd 07-embeddings-deep-dive
python -m venv .venv && source .venv/bin/activate
pip install -r requirements.txt
python main.py
Then try these:
- Set
TOY["question"]to[3, 4, 0], the same direction as account-recovery#1: that score becomes 1.00. Then try[0, 0, 1]: only invoices#1 scores. - Add your own paraphrases to
LABELLED, like("I forgot which email I signed up with", "change-email#1"). Each one the stand-in misses is a question a real model should get. With the key set, check that it does. - With the key set, embed with
dimensions=256inmain(). Vectors shrink 6x. Run the benchmark again and see whether recall moves.
pip install pytest && pytest -q runs the offline tests. They need no keys and no network.
Common beginner mistakes
- Different models for chunks and questions. No error, wrong results. Store the model name with every vector and check it.
- One minimum score for every model. Scores from different models sit on different scales. Re-pick cut-offs when you switch.
- Choosing from the leaderboard alone. MTEB is a shortlist. Your 50 labelled questions decide.
- Not checking the input limit.
all-MiniLM-L6-v2cuts input off after 256 word pieces, so the end of a long chunk is never embedded. Know your model’s limit and chunk below it. - Cutting dimensions without re-normalising. A truncated vector isn’t length 1 any more, so dot-product scores drift. Use the API’s
dimensionsparameter or scale it back yourself.
Questions you will face in production
“Should you fine-tune your own embedding model?” Almost never first. Try hybrid search (article 06), a re-ranker and better chunking before training anything. A cheaper trick is contextual retrieval: add a sentence about the source document to each chunk before embedding it. Anthropic reports it cut retrieval failures in the top 20 results by 35%, from 5.7% to 3.7%.
“Are embeddings sensitive data?” Treat them like the text they came from. In one study, researchers rebuilt 92% of 32-token inputs exactly from their embeddings alone (Morris et al., 2023). Apply the same access controls and encryption as the source.
“When do you re-embed?” When a chunk’s text changes, and when you change models. Store a hash of each chunk’s text beside its vector so a re-index only embeds what changed; article 10 calls this an embedding cache.
Check your understanding
Work out the cosine similarity of [1, 2, 2] and [2, 1, 2] by hand.
The dot product is 1·2 + 2·1 + 2·2 = 8. Both lengths are √(1 + 4 + 4) = 3. So the cosine is 8 / 9 = 0.89: close, not identical.
A teammate upgrades the question embedder to text-embedding-3-large and leaves the index alone. What happens, and what would have caught it?
The question now has 3,072 numbers and the stored chunks have 1,536. A length check or a vector(1536) column errors straight away. Without one, a zip-based cosine compares the first 1,536 numbers and returns plausible junk. The model-name guard catches it even when lengths match.
Model A's top score on a question is 0.82 and model B's is 0.45. Which model is better?
You can’t tell. Scores from different models sit on different scales. Run both through the same labelled questions and compare recall@k, which only cares about rank.
You have 2 million chunks and text-embedding-3-small. Roughly how much space do the vectors take, and how could you halve it?
2,000,000 × 1,536 × 4 bytes is about 12.3 GB before the index. Store them as pgvector’s halfvec (2 bytes a number) or ask for dimensions=768, then run your benchmark to check recall held up.
What to remember
- An embedding is a fixed-length list of numbers from a trained model. Similar meaning, nearby direction.
- Cosine similarity is the dot product divided by both lengths. For length-1 vectors, it’s just the dot product.
- Scores only compare within one model. So do minimum-score cut-offs.
- Start with
text-embedding-3-small. Change models when your own labelled questions say so, not the leaderboard. - Switching models means re-embedding everything. Store the model name with every vector and refuse mismatches.
- Storage is dimensions × bytes per number × vectors. Shorter vectors and smaller number formats are the levers.
What to study next
You can now build retrieval and choose its parts. The next question is how you know any of it works: article 08: LLM Evaluation Pipelines shows how to turn a benchmark like the 8 questions here into evals for the whole bot, and article 09: LLM-as-Judge covers grading answers that don’t have one right string.
Further reading
- OpenAI: Vector embeddings guide. Dimensions, the
dimensionsparameter, normalization, and the MTEB scores quoted above. - Anthropic: Embeddings. Why Anthropic points to Voyage, their model list, and notes on quantization and query vs document inputs.
- Muennighoff et al. (2022): MTEB, Massive Text Embedding Benchmark. The benchmark most models report. Good for a shortlist; benchmark on your own data.
- Kusupati et al. (2022): Matryoshka Representation Learning. The paper behind shortenable embeddings.
- pgvector. The README covers types, dimension limits, distance operators and HNSW indexes.
- Morris et al. (2023): Text Embeddings Reveal (Almost) As Much As Text. Why embeddings deserve the same protection as the text.
- Sentence Transformers documentation. The library behind many open embedding models, with clear explainers on the math.
Where this article comes from. This is a synthesis of common practice in AI engineering as of 2026, not a citation of any single paper. The sources above are where the broad mechanics and specific numbers come from. If you find an error or have a better source for a claim, the article gets fixed within a day, send me a note.