Precision and quantization
Every weight costs bits. Spend fewer bits per weight and the model shrinks, speeds up, and gets slightly worse in ways you have to measure.
Part of the Cheap and fast track on lAItest.
You can shrink a model by writing its numbers less precisely. It mostly still works. That should bother you.
Precision is a dial, not a property.
Precision is how many bits you spend per number
Every weight is stored in some number of bits. More bits means finer distinctions between values, more memory to hold them, and more traffic to move them. Quantization means storing the same weights in fewer bits each. Memory drops, fetching gets faster, and the answers shift a little. The whole game is finding out how little precision a particular job can stand.
The cost is measured, not guessed
Vector search is where the tradeoff gets published. Milvus reports a quantization scheme that keeps about three percent of the original memory footprint at roughly ninety-five percent recall, and searches about four times faster. Voyage measured a 99.48 percent storage cut by shrinking embeddings to 512 binary dimensions, while still matching a competitor model at full 3072-dimension floats.
A common misconception
Commonly believed: Quantization is effectively lossless if you pick a good method.
Actually: It is lossy by construction. You are deliberately throwing away distinctions between numbers. Good methods make the losses land where they matter least on average. Whether they matter for your task is not something a benchmark can answer for you.
A model quantized to fewer bits per weight is faster mainly because:
Answer: the same weights take fewer bytes to move from memory. Quantization removes no weights. It makes each one smaller, so the same set of weights travels from memory to the chip in less time. Fetching is usually the bottleneck, so that is where the speed comes from.
In one sentence
Precision is a price you pay per weight, and most jobs are willing to pay less than the default.