Meaning search vs keyword search

One finds what you meant, the other finds what you typed, and each fails exactly where the other works.

Part of the Memory and truth track on lAItest.

Search your notes for error code E4192. Meaning search may hand you a thoughtful page about error handling instead.

Keyword search matches strings. Meaning search matches positions.

The classic keyword method is BM25: count how often your words appear in a document, discount words that appear in everything, adjust for document length. It has no idea that "car" and "automobile" are related. Meaning search embeds the query, embeds every document in advance, and returns the nearest. It has no idea that E4192 is a string you need character for character.

A common misconception

Commonly believed: Embeddings made keyword search obsolete.

Actually: Exact tokens are where embeddings are weakest. Product codes, error numbers, surnames, a function name, rare in-house jargon — these are the things a meaning space smears together, because the model never saw them often enough to place any of them well. Systems that work in production run both methods and merge the results.

Which query is keyword search more likely to get right than meaning search?

Answer: "SKU 88-4410-B". A code is a literal string with almost no meaning attached to it, so exact matching nails it while embedding it drops it somewhere vague among other codes. The other two are paraphrases with no exact anchor at all, which is where meaning search earns its keep.

In one sentence

Keyword search is literal, meaning search is loose, and most real questions need some of each.