Learn / RAG in 7 lessons / Embeddings and vector search
Embeddings and vector search
What an embedding actually represents, how similarity search finds relevant chunks, and a runnable TF-IDF + cosine retriever that mirrors the same idea without a model.
What an embedding is
An embedding model maps a piece of text to a fixed-length list of numbers (say, 768 or 1536 of them) such that texts with similar meaning end up close together in that numeric space, and unrelated texts end up far apart. The model learns this mapping from massive amounts of text during training - nobody hand-designs the dimensions.
“Close together” is almost always measured with cosine similarity: the cosine of the angle between two vectors, ranging from -1 (opposite) to 1 (identical direction). It ignores vector length and only cares about direction, which matters because embedding magnitude isn’t a meaningful signal on its own.
Vector search, step by step
- Index time: embed every chunk once, store the vector alongside the chunk’s text and metadata (source, url, heading - see lesson 2).
- Query time: embed the incoming question with the same embedding model.
- Search: find the stored vectors closest to the question’s vector by cosine similarity.
- Return: the top-k chunks (typically k=3 to k=8) go into the prompt.
At small scale (thousands of chunks) brute-force comparison is fast enough - this is exactly
what this site’s own chatbot does with BM25 term-matching (src/lib/chat/bm25.ts), no
embeddings API at all. At real scale (millions of vectors) you need an approximate
nearest-neighbour (ANN) index - HNSW graphs or IVF clustering, which is what a vector
database (Qdrant, pgvector, Azure AI Search, Pinecone) actually implements under the hood, in
exchange for occasionally missing the true best match.
Building intuition without a model
You don’t need a neural network to see how vector search works. TF-IDF (term frequency - inverse document frequency) turns text into a vector using pure word-counting statistics: common words across all documents get down-weighted, words that are frequent in one document but rare elsewhere get up-weighted. It’s a much older, much simpler technique than neural embeddings, but it produces vectors you can compare with cosine similarity exactly the same way. Run this to see it end to end, standard library only:
:::pyrun
import math
from collections import Counter
documents = {
"d1": "Kubernetes pods run on nodes and are managed by deployments.",
"d2": "A Lambda function has a memory limit and a timeout, up to 900 seconds.",
"d3": "Chunking splits a document into smaller overlapping pieces before indexing.",
"d4": "Retrieval augmented generation retrieves relevant chunks before the model answers.",
}
def tokenize(text):
return [t.lower() for t in text.replace(".", "").replace(",", "").split()]
tokenized = {doc_id: tokenize(text) for doc_id, text in documents.items()}
# Inverse document frequency across the whole corpus.
n_docs = len(tokenized)
df = Counter()
for tokens in tokenized.values():
for term in set(tokens):
df[term] += 1
idf = {term: math.log(n_docs / count) + 1 for term, count in df.items()}
def tfidf_vector(tokens):
tf = Counter(tokens)
return {term: freq * idf.get(term, 0.0) for term, freq in tf.items()}
def cosine(vec_a, vec_b):
shared = set(vec_a) & set(vec_b)
dot = sum(vec_a[t] * vec_b[t] for t in shared)
norm_a = math.sqrt(sum(v * v for v in vec_a.values()))
norm_b = math.sqrt(sum(v * v for v in vec_b.values()))
if norm_a == 0 or norm_b == 0:
return 0.0
return dot / (norm_a * norm_b)
doc_vectors = {doc_id: tfidf_vector(tokens) for doc_id, tokens in tokenized.items()}
def search(query, top_k=2):
query_vector = tfidf_vector(tokenize(query))
scored = [(doc_id, cosine(query_vector, vec)) for doc_id, vec in doc_vectors.items()]
return sorted(scored, key=lambda pair: pair[1], reverse=True)[:top_k]
query = "how does chunking help retrieval"
for doc_id, score in search(query):
print(f"{doc_id} score={score:.3f} {documents[doc_id]}")
:::
Notice d3 and d4 - both about chunking/retrieval - win over the Kubernetes and Lambda
documents even though the query shares no exact words with d4. That’s TF-IDF’s weighting
doing real work; a neural embedding model does the same job with far more semantic nuance
(it also matches “auto” to “car”), but the retrieval mechanism - vectorize, compare, rank -
is identical either way.
Key takeaways
- An embedding is a fixed-length vector where geometric closeness approximates semantic closeness, learned by a model trained for that purpose.
- Vector search finds the stored chunks whose embeddings are closest to the question's embedding, almost always via cosine similarity.
- TF-IDF is not a neural embedding, but it's the same shape of idea - text becomes a vector, and vectors are compared - which is why it's a good way to build intuition before adding a model.
- At real scale you need an approximate nearest-neighbour index (HNSW, IVF, etc. - what Qdrant, pgvector, or Azure AI Search implement) because brute-force comparison against millions of vectors doesn't scale.
Quick check
3 questions - see how much stuck.