Multimodal input

Images, audio and video are converted into tokens before the model sees them, and tokens are what you are billed for.

Part of the Getting good answers track on lAItest.

A model does not look at your screenshot. It reads a few thousand tokens that used to be a screenshot.

Everything becomes tokens.

One door, several formats.

A transformer consumes a sequence of vectors. Text gets there through a tokenizer. An image gets there by being cut into patches and encoded. Audio arrives as a spectrogram, or as discrete audio tokens produced by a neural codec. Once inside, all of it sits in the same sequence and attention runs across the whole thing without caring which part used to be a picture.

A common misconception

Commonly believed: Multimodal means the model was shown the picture.

Actually: It means the picture was converted into the same kind of numbers as words and dropped into the same sequence. That is why an image costs tokens, why resolution changes the price, and why a model can miss small text in a photo you can read easily. The detail was discarded during conversion, before any thinking started.

Billing follows the conversion.

Because everything becomes tokens, everything lands on the same bill. Vendors publish separate rates for text, images, audio and video, and some price images by pixel count rather than per file. A high-resolution image can be worth more tokens than several pages of text, which is a surprise the first time it appears on an invoice.

Why does the same photo cost more when you send it at higher resolution?

Answer: More pixels become more patches, and each patch becomes tokens in the sequence. The image is cut up before the model sees it. More detail means more pieces, more pieces means a longer sequence, and sequence length is what you pay for. It is also why downscaling an image you only need the gist of is a genuine cost lever.

In one sentence

Multimodal is not a second kind of understanding. It is a second door into the same sequence of tokens.