Vector databases

What stores millions of embeddings, and why "find the nearest" is deliberately an approximation.

Part of the Memory and truth track on lAItest.

Finding the closest vector out of ten million by comparing against all ten million is far too slow. So nobody does that.

A vector database is an index for "nearest", not for "equals".

An ordinary database index answers "where is the row with this id". A vector index answers "which stored vectors point most like this one". It does that with approximate nearest neighbour search. A structure such as HNSW links each vector to its neighbours in a graph, and a query walks that graph towards the query point instead of scanning the whole table.

Approximate means you agree to sometimes miss.

Turn the search effort up and you find more of the genuinely nearest vectors and wait longer. Turn it down and it is faster and misses more. That trade has a name — recall — and it is a setting, not a defect. It also means a retrieval system can be quietly losing results without anything ever reporting an error.

A common misconception

Commonly believed: You need a specialist vector database before you can build anything.

Actually: Usually not. pgvector is an extension for Postgres, so vectors live in the database you already run, with the same backups and transactions as everything else, and that is enough for a lot of projects. Dedicated engines earn their place at larger scale or when you need sparse, dense and multi-vector search in one place. The search engines many companies already run do vector search too.

In one sentence

A vector database is a fast way to be approximately right about "nearest", at a size where being exactly right would be too slow.