Meaning as coordinates

An embedding turns text into a list of numbers, positioned so similar meanings land near each other.

Part of the Memory and truth track on lAItest.

"Car" and "automobile" share no letters. Anything that matches letters will never connect them.

The fix is to give every piece of text a position instead.

A vector is a list of numbers, and a list of numbers is a place.

Two numbers put a point on a map: three across, five up. Three numbers put it in a room. Embedding models use hundreds or thousands of numbers, so the point sits in a space nobody can picture. The arithmetic does not mind. Distance and direction still work exactly the same way they do on a map.

An embedding is a vector produced by a model trained to place meaning.

You hand it text, it hands back coordinates. It was trained so that text people treat as similar comes out nearby. "Refund policy" lands close to "how do I get my money back" and far from "roof repair". Nobody wrote those positions by hand. They fell out of training, which is why the same trick works on languages and phrasings nobody anticipated.

Try it

Pick a word and read off its nearest neighbours. This map has two dimensions; a real one has thousands. This step is an interactive widget; open the lesson to use it.

A common misconception

Commonly believed: Each number in an embedding stands for something you could name — one for "is it about food", one for "is it cheerful".

Actually: Almost never. The dimensions come out of training with no labels and no promise that any single one means anything alone. Meaning lives in the whole position, not in one coordinate. That is why you cannot read an embedding. You can only compare it to another one.

In one sentence

An embedding is an opinion about meaning, written down as a location.