Mixture of experts

Huge total parameters, small active ones — you pay memory for the size and compute for the slice.

Part of the What's inside track on lAItest.

Some of the largest models switch on roughly a thirtieth of themselves per token.

That is the design, not a defect.

One fat layer becomes many thin ones and a router.

In a mixture-of-experts layer, the feed-forward block is replaced by many parallel copies called experts. A small router network scores the experts for each token and sends that token to the top one or two. Every other expert does no arithmetic for that token. Total parameters go up steeply; work per token barely moves. The idea goes back to the Switch Transformer, arXiv:2101.03961.

Total versus active is the number to ask for.

DeepSeek V4-Pro is described as 1.6 trillion parameters in total with about 49 billion active per token. Mistral Large 3 is a sparse mixture at roughly 675 billion total and 41 billion active. Z.ai's GLM-5.2 is 753 billion total; its active count is widely reported as about 40 billion but is not confirmed in first-party documentation.

Sparse models sit at the bottom of the price sheet

Read from a live model index at page-render time, each figure linked to the vendor page it came from.
WhatValueProvenance
DeepSeek V4-Pro input$1.32 per 1Msource, verified .
DeepSeek Flash input$0.30 per 1Msource, verified .
MiniMax M3 input$0.30 per 1Msource, verified .
Median frontier input price$2 per 1Msource, verified . the middle of the frontier, for comparison
Cheapest frontier input price$0.75 per 1Msource, verified .

What does a mixture-of-experts model save, compared with a dense model of the same total size?

Answer: Arithmetic per token, because most experts stay idle for that token. Every expert has to be resident in memory in case the router picks it, so nothing is saved there. The saving is compute: only the routed experts run. That is why these models are cheap to serve at scale and awkward to run on one machine at home.

In one sentence

A mixture of experts buys capacity with memory instead of with compute, and that trade is why the cheapest capable models are sparse.