How LLMs See Images
Vision-enabled LLMs convert images into sequences of tokens that can be processed alongside text. This typically involves dividing images into patches and encoding them with a vision transformer.
The Vision Transformer (ViT)
The Vision Transformer architecture adapts the transformer model for image processing. Instead of processing words, it processes image patches.
1. Divide into Patches
The image is split into a grid of fixed-size patches (typically 14x14 or 16x16 pixels).
2. Flatten & Project
Each patch is flattened into a vector and linearly projected into an embedding space.
3. Add Position Info
Positional embeddings are added so the model knows where each patch came from.
4. Process with Transformer
The sequence of patch embeddings is processed by standard transformer layers.
One image, many patches
The same generated reference image, divided into large patches. Sizes of 64, 128 and 256 pixels keep the grid readable. These are teaching settings, not a specific model architecture or API billing rules.
6 × 4 = 24 patches
An RGB patch with P × P pixels contains 3P² values. ViT flattens them and applies a learned linear projection into an embedding vector. This does not simply average the patch into one color.
This grid shows only the partition. It does not calculate learned image embeddings. Reading small text also depends on resolution, projection, training and task.
API image tokens use separate accounting rules
OpenAI documentation, checked 2026-09-13. Billing patches are not necessarily the vision encoder’s patches.
| Model / mode | Example or limit |
|---|---|
| GPT-5-mini | 1024² px → 1024 × 1.2 ≈ 1229 image tokens |
| GPT-5.4 / original | 6000 px; 10,000 patch resizing budget |
| GPT-5.6-sol/terra/luna / original | 65,535 px; >30,000 patches rejected |
Exact resizing, supported modes and limits depend on the model. Image tokens, text and output must fit the context budget.
Current rules and cost calculatorWhich details does the task need?
A scene description and an exact serial number need different detail. Keep the original, choose useful crops and check the result against a reference. Higher resolution may help, but does not guarantee a correct answer.