30 - Embeddings and Vector Databases¶
Previous: 29 - Tool Use (Function Calling) | Index: All guides | Next: 31 - RAG (Retrieval-Augmented Generation)
Quick reference for turning text into vectors (embeddings), measuring similarity, and storing / searching vectors with NumPy, FAISS, Chroma, pgvector, Qdrant and Azure AI Search.
Last verified: 2026-09-27. For newer changes, check the Official docs links in the Introduction.
Introduction¶
Before you start¶
You should know: Python lists and NumPy arrays (17), and the idea of a model turning input into output (26). No advanced maths is needed: "vectors that are close together mean similar things" is enough to start.
The problem it solves: keyword search fails when the words differ: a user searches "cancel my subscription", the help page says "how to end your membership", and nothing is found. You need search by meaning, fast, over thousands or millions of texts.
Before embeddings: search engines matched words (with tricks like stemming and synonym lists that people maintained by hand). Word vectors (word2vec, 2013) showed that meaning can be captured as numbers; modern embedding models do this for whole sentences and paragraphs, and vector databases make searching millions of them fast.
Think of it like: a map where every text gets coordinates, and texts about similar things are placed near each other (the mental model below). Searching becomes "find the points closest to this one".
What are embeddings and vector databases?¶
An embedding is a list of numbers (a vector, e.g. 384 or 1024 numbers) that represents the meaning of a piece of text (or an image). An embedding model is trained so that texts with similar meaning get vectors that are close together, even when they share no words ("car" and "automobile"). A vector database stores millions of these vectors and quickly finds the ones closest to a query vector. Together they power semantic search, RAG, recommendations, clustering and duplicate detection.
Mental model: a map of meaning¶
^ dimension 2
|
"reset password" * * "forgot my login" <- close: similar meaning
| * "can't sign in"
|
----------------------------------------------------> dimension 1
|
| * "shipping takes 3 days"
| * "delivery time" <- another cluster
|
Real embeddings have hundreds of dimensions, but the idea is the same:
search = embed the question -> find the nearest points on the map.
INDEXING (once / when data changes) SEARCHING (every query)
documents -> chunks -> embedding model question -> embedding model -> query vector
-> vectors + text + metadata -> vector DB: nearest k vectors
-> vector database -> matching chunks (+ scores)
Keyword search (LIKE '%password%') finds exact words; semantic search finds meaning. Combining both (hybrid search) is often best.
Why use them?¶
- Search by meaning in documents, tickets, products, code.
- RAG: find the right context to give an LLM (31 - RAG).
- Cheap and fast: embedding models are much smaller and cheaper than chat LLMs.
- Other uses: clustering, deduplication, recommendations, classification features, anomaly detection.
Key terms¶
| Term | Meaning |
|---|---|
| Embedding / vector | List of floats representing meaning |
| Dimension | Length of the vector (384, 768, 1024, 1536, 3072 ...) |
| Embedding model | Model that converts text to vectors |
| Cosine similarity | Similarity of direction between two vectors (-1 to 1; higher = more similar) |
| Distance | Opposite of similarity (smaller = closer) |
| Nearest neighbours (kNN) | The k most similar vectors |
| ANN | Approximate nearest neighbour: fast, slightly inexact search |
| HNSW | Popular graph-based ANN index |
| Collection / index | A named set of vectors in a vector DB |
| Metadata / payload | Extra fields stored with each vector (source, date, user) for filtering |
| Hybrid search | Combining vector similarity with keyword (BM25) search |
| Normalisation | Scaling vectors to length 1 (then dot product = cosine) |
Where it fits: concept in 26 - LLM Fundamentals; models from 25 - Hugging Face or Ollama (36); used by 31 - RAG and agents' memory (32); pgvector builds on 20 - SQL; databases run in 43 - Docker.
Official docs¶
Where to read the latest, authoritative documentation:
| Resource | Link |
|---|---|
| Sentence Transformers | https://sbert.net/ |
| MTEB embedding leaderboard | https://huggingface.co/spaces/mteb/leaderboard |
| OpenAI embeddings guide | https://platform.openai.com/docs/guides/embeddings |
| Voyage AI docs | https://docs.voyageai.com/ |
| Chroma | https://docs.trychroma.com/ |
| pgvector | https://github.com/pgvector/pgvector |
| Qdrant | https://qdrant.tech/documentation/ |
| FAISS | https://github.com/facebookresearch/faiss/wiki |
| Azure AI Search | https://learn.microsoft.com/en-us/azure/search/ |
Contents¶
- Creating Embeddings
- Choosing an Embedding Model
- Similarity Metrics
- Brute-Force Search with NumPy
- How Vector Indexes Work (ANN, HNSW)
- FAISS (In-Memory Library)
- Chroma (Local Vector DB)
- pgvector (PostgreSQL)
- Qdrant (Vector DB Server)
- Azure AI Search and Other Managed Options
- Metadata Filtering
- Hybrid Search (Vector + Keyword)
- Other Uses: Clustering, Dedup, Classification
- Which Vector Store When
- Performance and Cost Tips
- Troubleshooting
- Try It
1. Creating Embeddings¶
Converting text into vectors with an embedding model. Local models (sentence-transformers, Ollama) or hosted APIs (OpenAI, Voyage, Azure OpenAI, Cohere); send a list of texts, get a list of vectors.
Use it for indexing documents and embedding each search query with the SAME model.
# Local, free: sentence-transformers
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
vectors = model.encode(["How do I reset my password?", "Shipping takes 3 days"],
normalize_embeddings=True) # numpy array, shape (2, 384)
# Local via Ollama (ollama pull bge-m3)
import ollama
resp = ollama.embed(model="bge-m3", input=["How do I reset my password?"])
vectors = resp["embeddings"] # list of lists (1024 dims)
# Hosted: OpenAI
from openai import OpenAI
client = OpenAI()
resp = client.embeddings.create(model="text-embedding-3-small", input=["How do I reset my password?"])
vectors = [d.embedding for d in resp.data] # 1536 dims
Anthropic does not offer its own embedding model; Voyage AI models are a common pairing with Claude (pip install voyageai), with an input_type of "document" for indexing and "query" for searches.
2. Choosing an Embedding Model¶
Picking the model that turns your text into vectors. Balance quality, language support, dimension (storage), speed, cost and privacy. Check the MTEB leaderboard on Hugging Face for rankings.
Use it before indexing; changing the model later means re-embedding everything.
| Model (examples) | Where | Dims | Notes |
|---|---|---|---|
all-MiniLM-L6-v2 |
Local (sentence-transformers) | 384 | Tiny, fast, English, great for learning |
bge-m3 |
Local (HF / Ollama) | 1024 | Multilingual, long inputs, strong |
multilingual-e5-large |
Local (HF) | 1024 | Multilingual (needs "query: " / "passage: " prefixes) |
nomic-embed-text |
Ollama | 768 | Good local default |
text-embedding-3-small / -large |
OpenAI / Azure OpenAI | 1536 / 3072 | Easy hosted option; can shorten dimensions |
| Voyage models | Voyage API | varies | Strong retrieval quality, domain variants (code, finance, law) |
Rules:
- Same model for documents and queries. Vectors from different models are not comparable.
- Store the model name with your index so you know when to re-embed.
- Multilingual data -> multilingual model.
3. Similarity Metrics¶
How "closeness" between two vectors is measured. Cosine similarity compares direction; dot product equals cosine for normalised vectors; Euclidean measures straight-line distance.
Use it for configuring a vector DB; use what the embedding model recommends (usually cosine).
| Metric | Formula idea | Range | Use |
|---|---|---|---|
| Cosine similarity | angle between vectors | -1 to 1 (higher = closer) | Default for text embeddings |
| Dot product | sum of a[i] * b[i] | any | Same as cosine if normalised; fastest |
| Euclidean (L2) distance | straight-line distance | 0 to inf (lower = closer) | Some models / image embeddings |
import numpy as np
def cosine(a: np.ndarray, b: np.ndarray) -> float:
return float(a @ b / (np.linalg.norm(a) * np.linalg.norm(b)))
Similarity scores are relative: 0.8 might be "very similar" for one model and "average" for another. Compare scores within the same model, and choose thresholds from your own data.
4. Brute-Force Search with NumPy¶
Searching by comparing the query to every stored vector. One matrix multiplication with normalised vectors gives all cosine scores; take the top k.
Use it for up to tens of thousands of chunks, prototypes, tests. Exact and dependency-free.
import numpy as np
docs = ["How do I reset my password?", "Shipping takes 3 days", "Forgot my login details"]
doc_vecs = model.encode(docs, normalize_embeddings=True) # (n_docs, dim)
def search(query: str, k: int = 2) -> list[tuple[str, float]]:
q = model.encode(query, normalize_embeddings=True) # (dim,)
scores = doc_vecs @ q # cosine for normalised vectors
top = np.argsort(-scores)[:k]
return [(docs[i], float(scores[i])) for i in top]
search("I can't log in") # [('Forgot my login details', 0.71), ('How do I reset my password?', 0.55)]
np.save("doc_vecs.npy", doc_vecs) # persist
5. How Vector Indexes Work (ANN, HNSW)¶
Data structures that find near vectors without checking every one. HNSW builds a multi-layer graph linking each vector to its neighbours; search walks the graph from coarse to fine. Results are approximate but very fast.
Use it for hundreds of thousands to billions of vectors, low latency needs.
Layer 2 (few nodes, long jumps) A ----------------------- F
Layer 1 A ------- C ------- F --- H
Layer 0 (all nodes, short links) A - B - C - D - E - F - G - H - I
Search: start at the top, jump toward the query, go down a layer, refine, repeat.
| Setting (HNSW) | Higher value means |
|---|---|
M |
More links per node: better recall, more memory |
ef_construction |
Better index quality, slower build |
ef / ef_search |
Better recall at query time, slower search |
Recall = share of the true nearest neighbours that the ANN search actually returns.
6. FAISS (In-Memory Library)¶
Meta's library for fast vector search in memory (CPU / GPU). Build an index object, add vectors, search; save / load the index as a file. Stores vectors only (keep text / metadata in your own list or DB).
Use it for fast local search inside one Python process; research; large static datasets.
import faiss # pip install faiss-cpu
import numpy as np
dim = doc_vecs.shape[1]
index = faiss.IndexFlatIP(dim) # exact inner product (= cosine with normalised vectors)
index.add(doc_vecs.astype(np.float32))
q = model.encode(["I can't log in"], normalize_embeddings=True).astype(np.float32)
scores, ids = index.search(q, 3) # top 3
[docs[i] for i in ids[0]]
faiss.write_index(index, "docs.faiss")
index = faiss.read_index("docs.faiss")
hnsw = faiss.IndexHNSWFlat(dim, 32) # approximate, faster for big data
7. Chroma (Local Vector DB)¶
An easy, open-source vector database that runs inside Python (or as a server). Collections store ids, documents, metadata and embeddings; Chroma can embed text for you with a default or custom embedding function.
Use it for prototypes, small / medium RAG apps, notebooks, local tools.
import chromadb # pip install chromadb
client = chromadb.PersistentClient(path="./chroma_db") # saved to disk
col = client.get_or_create_collection("support_docs", metadata={"hnsw:space": "cosine"})
col.add(
ids=["doc1", "doc2", "doc3"],
documents=["How do I reset my password?", "Shipping takes 3 days", "Forgot my login details"],
metadatas=[{"topic": "account"}, {"topic": "shipping"}, {"topic": "account"}],
) # uses Chroma's default embedding model
res = col.query(query_texts=["I can't log in"], n_results=2, where={"topic": "account"})
res["documents"][0], res["distances"][0], res["metadatas"][0]
col.upsert(ids=["doc2"], documents=["Shipping takes 2 to 4 days"], metadatas=[{"topic": "shipping"}])
col.delete(ids=["doc3"])
col.count()
To use your own model, pass embeddings= in add / query_embeddings= in query, or set an embedding_function on the collection.
8. pgvector (PostgreSQL)¶
A PostgreSQL extension that adds a
vectorcolumn type and similarity search. Store embeddings next to your normal data; query with distance operators in SQL; add an HNSW index for speed.Use this when you already use Postgres; you want vectors, metadata, permissions and transactions in one database.
CREATE EXTENSION IF NOT EXISTS vector;
CREATE TABLE chunks (
id BIGSERIAL PRIMARY KEY,
source TEXT,
content TEXT,
embedding VECTOR(384) -- must match your model's dimension
);
CREATE INDEX ON chunks USING hnsw (embedding vector_cosine_ops);
-- 5 most similar chunks (<=> cosine distance, <-> L2, <#> negative inner product)
SELECT source, content, 1 - (embedding <=> '[0.1, 0.2, ...]') AS similarity
FROM chunks
WHERE source LIKE 'handbook%'
ORDER BY embedding <=> '[0.1, 0.2, ...]'
LIMIT 5;
import numpy as np
import psycopg # pip install "psycopg[binary]" pgvector
from pgvector.psycopg import register_vector
conn = psycopg.connect("postgresql://postgres:pass@localhost:5432/postgres", autocommit=True)
register_vector(conn)
conn.execute("INSERT INTO chunks (source, content, embedding) VALUES (%s, %s, %s)",
("faq.md", text, np.array(vec)))
rows = conn.execute(
"SELECT content, 1 - (embedding <=> %s) AS sim FROM chunks ORDER BY embedding <=> %s LIMIT 5",
(np.array(qvec), np.array(qvec)),
).fetchall()
Available as a managed option on Azure Database for PostgreSQL (48 - Azure).
9. Qdrant (Vector DB Server)¶
A fast, open-source vector database server with rich filtering. Runs as a service (Docker or cloud); Python client creates collections, upserts points (vector + payload) and queries.
Use it for production RAG with large collections, heavy filtering, multiple apps sharing one store.
from qdrant_client import QdrantClient, models # pip install qdrant-client
client = QdrantClient(url="http://localhost:6333") # or QdrantClient(":memory:") for tests
client.create_collection(
"docs",
vectors_config=models.VectorParams(size=384, distance=models.Distance.COSINE),
)
client.upsert("docs", points=[
models.PointStruct(id=i, vector=v.tolist(), payload={"text": t, "topic": topic})
for i, (v, t, topic) in enumerate(zip(doc_vecs, docs, topics))
])
hits = client.query_points(
"docs",
query=model.encode("I can't log in").tolist(),
limit=3,
query_filter=models.Filter(must=[models.FieldCondition(key="topic", match=models.MatchValue(value="account"))]),
).points
[(h.payload["text"], h.score) for h in hits]
10. Azure AI Search and Other Managed Options¶
Hosted search services that combine vector, keyword and filter search. You push documents (with vectors or let the service vectorise them); query through an API; the provider handles scaling and backups.
Use it for enterprise apps, large data, when you do not want to run a database yourself.
| Service | Notes |
|---|---|
| Azure AI Search | Hybrid (vector + BM25) + semantic ranker, integrates with Azure OpenAI and Blob Storage (48) |
| Pinecone | Fully managed, serverless vector DB |
| Qdrant Cloud, Weaviate Cloud, Zilliz (Milvus) | Managed versions of open-source DBs |
| Elasticsearch / OpenSearch | Mature keyword search with vector support |
| MongoDB Atlas Vector Search, Cosmos DB | Vectors inside document databases |
11. Metadata Filtering¶
Restricting a vector search by structured fields. Store fields (source, date, language, customer, access level) with each vector; filter before / during the similarity search.
Use it for multi-tenant apps (users only see their documents), time ranges, document types, permissions.
col.query(query_texts=[q], n_results=5, where={"$and": [{"tenant_id": "acme"}, {"year": {"$gte": 2025}}]})
Security rule: enforce access filters in your code on every query; never rely on the LLM to hide documents it should not see.
12. Hybrid Search (Vector + Keyword)¶
Combining semantic similarity with classic keyword matching. Run both searches, then merge the rankings (for example Reciprocal Rank Fusion).
Use it for queries with exact terms (product codes, names, error messages) where pure vector search misses matches.
from rank_bm25 import BM25Okapi # pip install rank-bm25
bm25 = BM25Okapi([d.lower().split() for d in docs])
keyword_scores = bm25.get_scores("error E-4012".lower().split())
RRF_K = 60 # damping constant commonly used for Reciprocal Rank Fusion
def rrf(rankings: list[list[int]]) -> list[int]:
"""Merge several ranked lists of doc ids; items ranked high in any list rise to the top."""
scores: dict[int, float] = {}
for ranking in rankings:
for rank, doc_id in enumerate(ranking):
scores[doc_id] = scores.get(doc_id, 0) + 1 / (RRF_K + rank + 1)
return sorted(scores, key=scores.get, reverse=True)
Many vector DBs (Qdrant, Weaviate, Azure AI Search, Elasticsearch) have hybrid search built in.
13. Other Uses: Clustering, Dedup, Classification¶
Using embeddings beyond search. Treat vectors as features for classic ML (23 - Scikit-learn).
Use it for grouping feedback, finding duplicates, cheap classifiers.
from sklearn.cluster import KMeans
from sklearn.linear_model import LogisticRegression
labels = KMeans(n_clusters=8, n_init="auto", random_state=42).fit_predict(vectors) # topic clusters
DUPLICATE_THRESHOLD = 0.95 # tune on your own data
sim = vectors @ vectors.T
dupes = [(i, j) for i in range(len(sim)) for j in range(i + 1, len(sim)) if sim[i, j] > DUPLICATE_THRESHOLD]
clf = LogisticRegression(max_iter=1000).fit(train_vectors, train_labels) # fast text classifier
14. Which Vector Store When¶
Picking storage for your vectors. Start simple; move up when size, concurrency or features demand it.
Use it for designing a RAG / search system.
| Situation | Choose |
|---|---|
| Learning, < 50k chunks, one script | NumPy / FAISS |
| Prototype RAG app with persistence | Chroma |
| Already on Postgres, want SQL + vectors | pgvector |
| Large scale, heavy filtering, dedicated service | Qdrant / Weaviate / Milvus |
| Enterprise on Azure, hybrid + semantic ranking | Azure AI Search |
| No-ops managed service | Pinecone, Qdrant Cloud |
15. Performance and Cost Tips¶
Keeping indexing and search fast and cheap. Batch embedding calls, cache vectors, choose dimensions wisely.
Use it for indexing large corpora, production search.
- Embed in batches (e.g. 64 to 256 texts per call) instead of one by one.
- Cache embeddings (hash of text -> vector); never re-embed unchanged text.
- Smaller dimensions = less storage and faster search (some models let you truncate).
- Store vectors as
float32(or quantized in the DB) rather than float64. - Re-embed everything when you change the embedding model; version your index.
16. Troubleshooting¶
| Problem | Fix |
|---|---|
dimension mismatch / expected 1024 dimensions, got 384 |
Collection / column created for another model; recreate with the right size |
| Results look random | Different models for docs vs queries; missing normalisation; wrong metric |
| Exact terms (codes, names) not found | Add hybrid keyword search |
| Good matches ranked low | Chunks too big / too small (31); add a reranker; try a stronger model |
| Slow search with many vectors | Use an ANN index (HNSW); raise hardware; filter first |
| Scores all around 0.8 | Normal for some models; compare relative ranking, tune thresholds on your data |
faiss install fails on Windows |
pip install faiss-cpu (or use Chroma / Qdrant) |
| Users see other users' docs | Enforce metadata filters (tenant / permissions) in code on every query |
| Chroma / DB lost data on restart | Use PersistentClient(path=...) / Docker volumes |
17. Try It¶
Short exercises to practise this guide. Try each task yourself first, then open the solution.
Use it right after reading the guide, or later as a quick self-test.
Exercise 1: Cosine by hand¶
Embed three sentences and print the similarity matrix.
Solution
Exercise 2: Chroma with a filter¶
Add three documents with a topic metadata field and query only the account topic.
Solution
col = chromadb.PersistentClient("./db").get_or_create_collection("kb", metadata={"hnsw:space": "cosine"})
col.add(ids=["1", "2", "3"], documents=docs, metadatas=[{"topic": "account"}, {"topic": "shipping"}, {"topic": "account"}])
col.query(query_texts=["can't log in"], n_results=2, where={"topic": "account"})
Exercise 3: Same model rule¶
Why must documents and queries use the same embedding model?
Solution
Each model has its own vector space: dimensions and meaning differ, so distances between vectors from two different models are meaningless. Changing the model means re-embedding everything.
Previous: 29 - Tool Use (Function Calling) | Index: All guides | Next: 31 - RAG (Retrieval-Augmented Generation)