Why We Switched from Cosine to Dot-Product Similarity for Embedding Retrieval

Normalized embeddings make dot product identical to cosine, but faster.

by
Why We Switched from Cosine to Dot-Product Similarity for Embedding Retrieval

Why We Switched from Cosine to Dot-Product Similarity for Embedding Retrieval

For years, cosine similarity was the default choice for comparing embedding vectors in retrieval-augmented generation (RAG) pipelines. It’s intuitive and well-understood. But after profiling our retrieval stack, we realized that cosine was costing us latency and complexity for zero gain. We switched to dot-product similarity and never looked back.

Cosine similarity vs dot product: which one should you use?

Cosine similarity compares only the direction of two vectors: it divides their dot product by the product of their lengths, so the score always falls between −1 and 1. The dot product also grows with vector length, so a longer vector can score higher even when it points in the same direction. When every embedding is L2-normalized to unit length, the two give the same score and the same ranking, and the dot product is cheaper because it skips the division. Use cosine when your vectors are not normalized and only meaning matters; use the dot product when they are normalized, or when the embedding model was trained for inner-product scoring and magnitude carries signal.

The Math: Cosine vs. Dot Product

Cosine similarity measures the angle between two vectors:

cosine_sim(A, B) = (A · B) / (||A|| * ||B||)

Dot-product similarity is simply:

dot_sim(A, B) = A · B = Σ A_i * B_i

The key insight: if every vector is L2-normalized (unit length), then ||A|| = ||B|| = 1, and cosine similarity reduces to dot product. The results are identical.

But why would you normalize? In most embedding models, output vectors are not normalized by default. You must explicitly normalize them after inference. Once you do, cosine and dot product become equivalent.

Why Cosine Is Slower

Cosine requires two extra operations per comparison: computing the norm of each vector (if not cached) and dividing the dot product by the product of norms. In practice, many vector databases compute cosine by first normalizing queries on the fly or storing precomputed norms. Either way, it adds CPU cycles.

Consider a search over 1 million vectors. With cosine, each comparison does an extra division and multiplication. With dot product, it’s just the inner product. On modern hardware, that difference adds up—especially when you’re doing thousands of queries per second.

Real-World Impact

We run pgvector on PostgreSQL 16 with a dataset of 10 million 768-dimensional embeddings (BGE-M3). Our benchmark:

  • Cosine: ~12ms average query time (with IVFFlat index, probes=10)
  • Dot product (normalized): ~9ms average query time

That’s a 25% reduction in latency. For our real-time RAG pipeline serving user queries, this meant dropping from 150ms to 110ms end-to-end. No change in recall—identical results because all vectors are normalized.

How We Normalized

We normalized during embedding generation. After calling the BGE-M3 model via vLLM, we apply L2 normalization to the output vector before inserting into Postgres:

import numpy as np

def normalize(vec):
    norm = np.linalg.norm(vec)
    return vec / norm if norm > 0 else vec

Then we store the normalized vector in a pgvector column and use <=> (cosine) operator—but since vectors are normalized, pgvector’s cosine implementation actually does dot product internally. We later switched to using the <#> (inner product) operator explicitly and saw identical results.

Indexing Considerations

pgvector supports three index types: IVFFlat, HNSW, and (soon) diskann. For cosine similarity, pgvector internally normalizes vectors before indexing when you specify vector_cosine_ops. But this normalization happens at query time, adding overhead. With dot product, you can use vector_ip_ops directly, skipping that step.

We rebuilt our IVFFlat index with vector_ip_ops:

CREATE INDEX idx_embeddings_ip ON embeddings USING ivfflat (embedding vector_ip_ops) WITH (lists = 1000);

This index is used for inner product searches. Querying is straightforward:

SELECT id, embedding <#> query_embedding AS distance
FROM embeddings
ORDER BY embedding <#> query_embedding
LIMIT 10;

Note: <#> returns the negative inner product (pgvector returns distance, not similarity). So lower values mean more similar. Adjust your threshold logic accordingly.

When Not to Switch

If you cannot normalize your embeddings—for example, if you need to preserve magnitude for some downstream task—stick with cosine. But in most RAG setups, you only care about direction, not magnitude. Normalization is safe.

Also, if your vector database doesn’t support inner product indexing (e.g., older versions of Faiss), cosine may be the only option. But modern tools like pgvector, Qdrant, and Weaviate all support dot product.

Conclusion

Switching from cosine to dot-product similarity is a free performance win for normalized embeddings. It eliminates redundant computations, simplifies indexing, and reduces query latency by 20–30% with zero impact on accuracy. If you’re using cosine similarity in your RAG pipeline, benchmark a switch to dot product—you might be surprised at the gains.

#embedding#lessons-learned#rag#retrieval#vector-search
Share — X / Twitter · LinkedIn · HN · Email
Damir Radulić
Founder of RiNET. On the Croatian internet since 1996 (Kvarner Net). In Amsterdam now, building autonomous AI infrastructure that runs on Monday morning when nobody's watching — sovereign stacks, agent swarms, LoRA fine-tuning, civic-intelligence platforms.

Related