Multimodal Models Explained: How Images Become Tokens
A multimodal model does not see an image. It sees patch embeddings projected into the same space as text tokens. What that implies for cost and accuracy.
The phrase "the model can see images" hides the mechanism, and the mechanism is what predicts behaviour. A multimodal model does not process pixels alongside text in some special visual pathway. It converts the image into a sequence of vectors that live in the same embedding space as text tokens, and then does exactly what it always does.
Once you hold that picture, most of the surprising behaviour — why small text in a screenshot goes wrong, why one image can cost more than a page of prose, why cropping helps — stops being surprising.
Patches are the unit, not pixels
The foundation is the Vision Transformer, introduced in An Image is Worth 16x16 Words (arXiv 2010.11929, ICLR 2021). The insight in the title is the whole idea: cut the image into fixed-size square patches, flatten each one, and treat the resulting sequence like a sequence of words. A pure transformer applied directly to patch sequences performs very well, no convolutions required.
So a 16 by 16 pixel patch is the visual analogue of a token. An image becomes a grid of them, plus positional information so the model knows where each sat.
This immediately explains the resolution problem. Detail smaller than a patch has no dedicated representation. Text rendered at eight pixels tall does not occupy enough patches to be reconstructible, which is why models misread small labels in dense screenshots and why cropping to the region of interest works better than sending the whole screen and asking politely.
Three stages, one pipeline
Almost every current vision-language model is the same three components in a row.
1. A vision encoder. Usually a ViT, frequently one pretrained contrastively against text. CLIP established that approach — train image and text encoders jointly so that matching pairs land close together in a shared space — and its descendants remain the standard starting point. The encoder turns patches into embeddings that already carry semantic content rather than raw appearance.
2. A projector. The encoder output has its own dimensionality, which is not the language model dimensionality. A small learned network — often just an MLP — maps one to the other. This is where a great deal of the practical alignment work happens, and it is small enough to train separately from both other components.
3. The language model. Two integration styles are common. Either the projected vectors are inserted into the token sequence as ordinary embeddings at a specific position — the dominant approach in recent open models — or dedicated cross-attention layers are added so the language model attends to visual features without them occupying sequence positions.
The inline approach is simpler and keeps one attention mechanism over everything. Its cost is that image content consumes context window, which brings us to the part that shows up on your invoice.
Images consume context
If visual embeddings sit in the sequence, they occupy positions the same way text does. A large image can therefore be equivalent to a substantial block of text, and providers bill it accordingly.
The arithmetic is unforgiving. Patch count grows with the square of resolution: double each dimension and you get roughly four times the visual tokens. A high resolution screenshot can dominate a conversation that also contains several pages of writing.
Because full resolution scales badly, production systems use tiling. Downscale the image to a base representation, split larger images into fixed-size tiles, encode each tile, and concatenate. That keeps the relationship between resolution and cost roughly linear in area rather than unbounded, at the price of some cross-tile context.
Do not guess the numbers. Providers differ in patch size, tiling thresholds and how they count, and they change these between model versions. Send a representative image, read the reported input token count, and calculate from that.
Practical consequences
- Crop before you send. A tight crop of the relevant region beats a full screenshot on both accuracy and cost. You are removing patches that contribute noise.
- Resolution matters more than file size. Compression artefacts hurt a little; downscaling below the point where text is legible to you hurts a lot.
- Split multi-part images. Several focused images usually work better than one collage, because each gets its own full patch budget.
- Text position is approximate. Patch grids plus positional embeddings give a coarse sense of layout. Expect reliable reading order in simple documents and unreliable coordinates in complex ones.
- Ask for what is visible, not what is inferred. Transcription and description are strong. Precise measurement, exact pixel positions and counting many small similar objects remain weak.
Beyond images
The same architecture generalises, which is why "multimodal" expanded from vision so quickly.
Audio follows the identical pattern with a different encoder: convert the waveform to a spectrogram, encode it into embeddings, project into the language model space. Native audio input removes the separate speech-to-text stage, which matters because a transcription step discards tone, emphasis, overlapping speakers and non-speech sound before the model ever sees it.
Video is images plus a sampling problem. Frames become patch sequences, and the design question is which frames to keep, since sending every frame at full resolution is not affordable. Most systems sample sparsely and lose fine-grained motion in exchange.
Output is a separate capability from input. A model that accepts images does not necessarily produce them; image generation usually means a different model, or a separate decoder head, invoked as a tool.
How to evaluate one for your use
Benchmark scores on academic visual question answering datasets predict very little about your documents. Test on your own material instead.
- Collect twenty representative inputs, including the ugly ones — low contrast, rotated, screenshots of screenshots.
- Write the expected answer for each by hand.
- Measure accuracy, and separately measure the input token count per image so you know the unit cost.
- Repeat at two or three resolutions. There is usually a knee where accuracy stops improving and cost keeps climbing, and that is your operating point.
- Check failure style. A model that says it cannot read something is far more usable than one that confidently invents a plausible value.
That last point deserves emphasis. In document and screenshot work, calibrated uncertainty is worth more than a few points of raw accuracy, because a wrong number you can detect costs you a retry and a wrong number you cannot detect costs you a bug.
Common questions
How does a language model process an image?
The image is cut into fixed-size patches, a vision encoder turns those patches into embeddings, and a small projection network maps them into the language model embedding space. From there they are treated much like text tokens.
Why do images cost so many tokens?
Because visual embeddings occupy positions in the sequence just as text tokens do, and patch count grows with the square of resolution. Doubling each dimension roughly quadruples the visual tokens, which is why providers tile large images rather than encoding them whole.
Why do models misread small text in screenshots?
Detail smaller than one patch has no dedicated representation, so very small characters are not reconstructible from the encoding. Cropping to the region of interest or sending a higher-resolution image of that region works better than sending the full screen.