Embeddings Explained: Vectors, Similarity and Where It Breaks
Embeddings turn text into vectors so you can search by meaning instead of keywords. Here is how they work, what similarity really measures, and the failure modes to expect.
An embedding model takes a piece of text and returns a fixed-length list of numbers. Texts with similar meaning end up close together in that space, which is what lets you search by meaning rather than by exact keyword match.
That is the whole idea, and it is genuinely useful. The trouble starts when teams treat cosine similarity as a measure of relevance, which it is not, and build retrieval systems that fail in ways nobody can debug.
What the vector represents
The output is a dense vector — typically several hundred to a couple of thousand dimensions, every one a real number. There is no interpretable meaning to any single dimension. The information lives in the geometry of the whole space.
Embedding models are trained with contrastive objectives: pull related pairs together, push unrelated pairs apart. What counts as related is a training decision, and it is the single most important thing to understand about any given model. A model trained on question-answer pairs learns that a question and its answer should be close. A model trained on paraphrase pairs learns something quite different.
This is why an embedding model that performs well on one task can perform poorly on another with the same input text. You are not measuring similarity in the abstract; you are measuring similarity as defined by that model's training signal.
Cosine similarity and what it does not tell you
The standard comparison is cosine similarity: the cosine of the angle between two vectors, ranging from -1 to 1. Most modern embedding models produce unit-normalised vectors, in which case cosine similarity and dot product are the same operation, and Euclidean distance is a monotonic transform of it. The choice of metric matters much less than people expect.
What matters more is that the scores are not calibrated. A similarity of 0.82 means nothing on its own. Different models occupy different ranges — some cluster almost everything above 0.7, others spread more widely — so a fixed threshold tuned for one model is meaningless for another.
Two consequences follow. First, always tune thresholds empirically against your own data, and re-tune when you change models. Second, prefer relative ranking over absolute cutoffs where you can. Top-k retrieval is robust to miscalibration; a hard-coded "reject below 0.75" is not.
Dimensions and the Matryoshka trick
Higher-dimensional embeddings can encode more, at the cost of storage and search latency. Historically that was a fixed choice made by the vendor.
Matryoshka representation learning changed that. The training procedure applies the loss not only to the full vector but also to truncated prefixes, so a 768-dimensional embedding remains meaningful when cut to 256, or 128, or 64 — the earlier dimensions carry the coarser, more important structure, like nested dolls. Several current embedding APIs expose this as a dimension parameter.
The practical pattern is two-stage retrieval: search a large index using heavily truncated vectors, then rerank the survivors with full-length vectors. You get most of the accuracy at a fraction of the memory footprint.
Choosing a model without guessing
The Massive Text Embedding Benchmark (MTEB) is the standard reference point. The original release covered 58 datasets across 112 languages and eight task types — bitext mining, classification, pair classification, clustering, reranking, retrieval, semantic textual similarity and summarisation. The community expansion, MMTEB, extended it to over 500 evaluation tasks across more than 250 languages.
Read it as a filter, not an answer. A leaderboard aggregate hides the fact that you probably care about one task type — usually retrieval — in one language, on one kind of document. Filter to that column before comparing.
The things that actually decide the choice in production:
- Retrieval score specifically, not the overall average.
- Maximum input length. Many embedding models silently truncate long inputs, which means the tail of your document was never indexed.
- Dimension count, because it sets your storage and query cost at scale.
- Domain fit. Code, legal text and clinical notes all behave differently from general prose, and a specialised model often beats a higher-ranked general one.
- Self-hosted or API. Re-embedding a large corpus is a real cost; being able to run locally changes the economics of iteration.
Where embedding search actually fails
Exact identifiers. Error codes, SKUs, function names and version numbers are precisely what semantic search is worst at, because embeddings capture meaning and these carry almost none. A user searching for a specific error string wants lexical match. This is the main argument for hybrid search — combine dense retrieval with BM25 keyword search and merge the rankings.
Negation. "Contains gluten" and "does not contain gluten" embed close together. The words overlap almost entirely and the semantic difference is a single operator. If negation is load-bearing in your domain, do not rely on similarity alone.
Asymmetry. A short question and a long answer are different kinds of text. Models trained for symmetric similarity handle this badly. Many models expect task-specific prefixes on queries versus documents — omitting them quietly costs you accuracy, and it is one of the most common configuration mistakes.
Chunking. When retrieval underperforms, the embedding model is usually not the culprit. Splitting on a fixed character count severs sentences and separates claims from their qualifiers. Split on structure — headings, function boundaries, paragraphs — and overlap adjacent chunks slightly so a fact spanning a boundary appears whole somewhere.
A workable default setup
- Pick a well-rated retrieval model that supports your language and input length.
- Chunk on document structure, with modest overlap, keeping each chunk interpretable on its own.
- Apply whatever query and document prefixes the model documents. Check this; it is easy to miss.
- Run hybrid retrieval — dense plus keyword — because the failure modes are complementary.
- Rerank the top 20 to 50 candidates with a cross-encoder if accuracy matters. This is usually a larger gain than switching embedding models.
- Build a small labelled set of real queries with known correct documents, and measure recall. Without it you are tuning blind.
The last step is the one teams skip and the one that pays. Retrieval quality caps the quality of everything downstream, and it is the part of a RAG system that is straightforward to measure in isolation.
Common questions
What does a cosine similarity of 0.8 mean?
On its own, nothing. Score ranges differ per model and are not calibrated across them. Tune thresholds against your own labelled data, and prefer top-k ranking over fixed cutoffs wherever possible.
Do I need a vector database?
Not below roughly a hundred thousand vectors. Brute-force similarity over a numpy array or a Postgres extension is fast enough and far simpler. Reach for dedicated infrastructure when scale or filtering complexity forces it.
Why does semantic search miss exact matches?
Embeddings encode meaning, and identifiers like error codes or SKUs carry almost none. Run hybrid search that combines dense vectors with BM25 keyword matching so lexical queries are handled by the method suited to them.