Vision & Images

Beginner

How modern LLMs process and understand visual information alongside text.

Last updated: Sep 13, 2026

How LLMs See Images

Vision-enabled LLMs convert images into sequences of tokens that can be processed alongside text. This typically involves dividing images into patches and encoding them with a vision transformer.

The Vision Transformer (ViT)

The Vision Transformer architecture adapts the transformer model for image processing. Instead of processing words, it processes image patches.

1. Divide into Patches

The image is split into a grid of fixed-size patches (typically 14x14 or 16x16 pixels).

2. Flatten & Project

Each patch is flattened into a vector and linearly projected into an embedding space.

3. Add Position Info

Positional embeddings are added so the model knows where each patch came from.

4. Process with Transformer

The sequence of patch embeddings is processed by standard transformer layers.

Dosovitskiy et al.: An Image is Worth 16x16 Words

One image, many patches

The same generated reference image, divided into large patches. Sizes of 64, 128 and 256 pixels keep the grid readable. These are teaching settings, not a specific model architecture or API billing rules.

6 × 4 = 24 patches

An RGB patch with P × P pixels contains 3P² values. ViT flattens them and applies a learned linear projection into an embedding vector. This does not simply average the patch into one color.

This grid shows only the partition. It does not calculate learned image embeddings. Reading small text also depends on resolution, projection, training and task.

API image tokens use separate accounting rules

OpenAI documentation, checked 2026-09-13. Billing patches are not necessarily the vision encoder’s patches.

Model / modeExample or limit
GPT-5-mini1024² px → 1024 × 1.2 ≈ 1229 image tokens
GPT-5.4 / original6000 px; 10,000 patch resizing budget
GPT-5.6-sol/terra/luna / original65,535 px; >30,000 patches rejected

Exact resizing, supported modes and limits depend on the model. Image tokens, text and output must fit the context budget.

Current rules and cost calculator

Which details does the task need?

A scene description and an exact serial number need different detail. Keep the original, choose useful crops and check the result against a reference. Higher resolution may help, but does not guarantee a correct answer.