Skip to main content
Concept
A distance metric is the formula a vector search uses to decide how close two embeddings are. Cosine distance compares direction and ignores length, Euclidean distance measures the straight-line gap between two points, and Manhattan distance sums the gap along each dimension. The right choice is usually the metric your embedding model was trained for, which its provider normally documents.

Learning objectives

After reading this article you will be able to:
  • Explain what cosine, Euclidean, Manhattan, and dot product measure
  • Compare how the metrics rank the same vectors when vector lengths differ
  • Describe why cosine and Euclidean distance agree on normalized vectors
  • Choose a metric that fits the embedding model and whether vectors are normalized

What are cosine, Euclidean, and Manhattan distance?

They are three ways to turn a pair of vectors into one number that says how far apart they are. Dot product is a related similarity score that many models use, so it is included for comparison. For two vectors a and b with n dimensions:
Similarity and distance point in opposite directions. Cosine similarity and dot product are similarities, so higher is closer. Cosine distance, Euclidean, and Manhattan are distances, so lower is closer. Search systems usually sort by distance, closest first.

How can the metrics rank the same vectors differently?

They weigh length and per-dimension gaps differently, so the nearest vector under one metric can be farther away under another. Take a query vector a = (1, 0) and two stored vectors, b = (0.6, 0.8) and c = (2, 0). Every column below is a distance, so lower is closer. c points in exactly the same direction as a but is twice as long. Cosine ignores the length, so c is a perfect match. Euclidean counts the extra length as distance, so b wins. Cosine and Euclidean disagree here because the vectors have different lengths. Manhattan distance is also sensitive to length, and it picks c only because of this particular geometry. b differs from a by 0.4 and 0.8 across two dimensions. Euclidean squares those gaps before adding them, so they total less than c’s single gap of 1.0. Manhattan adds them as they are, so they total 1.2.

When do the metrics agree?

When every vector is normalized to unit length, cosine and Euclidean distance produce the same ranking. For unit vectors, the squared Euclidean distance equals 2 − 2 × cosine similarity, so ordering by one orders by the other. The dot product of unit vectors also equals their cosine similarity. Manhattan distance does not have this property: it can rank normalized vectors differently from cosine and Euclidean distance. Normalizing a vector means dividing it by its length. Do it when your model’s documentation recommends it, or when you want vector length to stop influencing results.
Many embedding models return normalized vectors. If yours does, cosine is a safe default. With exact search, switching to Euclidean distance gives the same ranking; with approximate search, results can differ slightly.

How do you choose a distance metric?

Start from the embedding model, then confirm with your own data.
  1. Use the metric the embedding model was trained for. Model documentation usually says whether to compare its embeddings with cosine similarity, dot product, or Euclidean distance. Matching it gives the rankings the model was optimized to produce.
  2. Check whether the vectors are normalized. If they are, cosine, Euclidean, and dot product rank results the same way, and the choice matters little.
  3. Keep one metric per index. Every vector in an index, and every query vector, must come from the same model and be compared with the same metric. Changing the model means re-embedding every stored vector, and changing the metric typically means rebuilding an approximate index such as HNSW.
  4. Evaluate on your own queries. When in doubt, measure recall and relevance on a sample of real questions with each candidate metric.

How does HelixDB support distance metrics?

Each HelixDB vector index declares one metric (cosine, Euclidean, or Manhattan) and a fixed dimension. Every stored embedding and every query vector must match that dimension exactly.
  • Results come back closest first. Hits are ordered by distance, then by ID when distances tie.
  • Search is approximate. Vector indexes run approximate nearest neighbor search, and the docs state over 90% recall.
  • Cosine needs a non-zero vector. A cosine vector with zero length has no direction, so HelixDB rejects it with the zero_norm_cosine_vector error code.
  • Dot-product models can use cosine. For a model that recommends dot product on normalized embeddings, use cosine, which ranks those vectors the same way.
The application computes embeddings, and HelixDB stores and indexes them on nodes or edges. See What is vector search? for how nearest-neighbor search works.

Frequently asked questions

Is cosine similarity better than Euclidean distance?

Neither is better in general. Cosine compares direction only, which suits most text embeddings. Euclidean distance also accounts for length. For normalized vectors they rank results identically.

Can I change the distance metric after building an index?

Usually you build a new index. An approximate index such as HNSW is constructed by comparing vectors with its metric, so its structure reflects that metric. The stored embeddings can be reused if the model stays the same.

When is Manhattan distance useful?

Manhattan distance adds up absolute differences without squaring them, so one very large difference makes up a smaller share of the total than it does in Euclidean distance. It is sometimes preferred for sparse or very high-dimensional features. Test it against your own data before choosing it.

What is the difference between dot product and cosine similarity?

Cosine similarity is the dot product divided by both vectors’ lengths. For unit-length vectors they are equal. For vectors of different lengths, the dot product also rewards longer vectors.

What are vector embeddings?

How models turn inputs into vectors, and why a model change means re-embedding.

What is vector search?

Embeddings, nearest neighbors, and approximate search.

What is HNSW?

A layered graph index widely used for approximate nearest neighbor search.

What is filtered vector search?

Return the right top k when results must match a filter.

What is a vector database?

What a vector store does and where vector-only storage falls short.

HelixDB vector indexes

Create a vector index and run nearest-neighbor search.