Prompt Caching

Intermediate

Reuse computed KV caches across API requests to dramatically reduce cost and latency for repeated prompt prefixes.

Last updated: Sep 13, 2026

Prompt caching rewards stable prefixes.

Prompt caching is not memory, retrieval, or a database. It is a way for an inference provider to reuse the already-computed KV state for the beginning of a prompt when later requests start with the same tokens. The changing part still has to be processed; the win comes from not recomputing the repeated prefix.

The prefix is the contract

Cacheability is mostly determined before the model sees the user's actual task. If the long, expensive part begins at token 500 but token 20 changes on every request, the shared prefix is effectively broken. Design the prompt shape first, then tune provider settings.

Cache hit versus cache miss

On a miss, the provider performs the expensive prefill for the whole prompt and may write the reusable prefix. On a hit, the provider reads the cached prefix and only computes the suffix. This changes input cost and latency, but output generation still costs what it costs.

Worked token-count example: 8,000 stable prefix tokens plus 2,000 new tokens. Toggle whether the prefix is already cached. Counts describe prefill work and stored/read prefix tokens, not elapsed time or a promised speedup.

Prefill tokens
10000
Prefix tokens written
8000
Prefix tokens read
0
Cached Prefix: 8,000
New Tokens: 2,000

Try the prefix rule

The fastest way to understand prompt caching is to edit the beginning of a prompt. Tiny changes near the top can destroy the cache hit; changes after the stable prefix usually preserve it.

Prefix Matching Explorer

Edit the prompt below and watch how changes affect caching. Cache matching works character-by-character from the start — any edit invalidates everything after it.

Original Prompt (cached)
You are a helpful AI assistant specializing in code review. Always provide constructive feedback. Focus on security, performance, and readability. Use markdown formatting in your responses. --- Review the following Python function: def calculate_total(items): total = 0 for item in items: total += item.price * item.quantity return total
Your Edited Prompt
Cache Boundary Preview
You are a helpful AI assistant specializing in code review. Always provide constructive feedback. Focus on security, performance, and readability. Use markdown formatting in your responses. --- Review the following Python function: def calculate_total(items): total = 0 for item in items: total += item.price * item.quantity return total
Cached Tokens
90
New Tokens
0
Savings
100%
Cached TokensNew Tokens

Provider behavior differs

Do not build your architecture around a single marketing number. Providers differ in how caches are created, how long they live, which inputs count, and how usage is reported. The portable skill is prompt discipline: stable prefix first, volatile suffix last, measure every deployment.

Anthropic (Claude)

Cache behavior, minimum prefix length, write price, read price and lifetime depend on the Claude model and cache mode. Check cache_creation_input_tokens and cache_read_input_tokens. A cold write and a warm read have different costs.

Sources

OpenAI (GPT-4o)

Caching modes and prices depend on the model. GPT-5.6 and later support implicit and explicit breakpoints, with separate write and read charges. Earlier models use implicit caching. Check the exact model's minimum length, retention and cached-token usage.

Sources

Google (Gemini)

Gemini supports implicit caching on eligible models and explicit cached-content resources in supported APIs. Minimum lengths, TTL, storage charges and cache discounts depend on the model and mode.

Sources

Cost model

Writes and reads can have different prices. Savings depend on reuse before expiry. The calculator makes that grouping explicit; replace its hypothetical tariff with the prices for your selected model.

Hypothetical tariff, not a provider quote. A group reuses one prefix before expiry; each new group starts with a write. Include TTL expiry, routing misses or prefix edits by reducing requests per group. Output-token costs are excluded equally.

Cold writes
10
Cache reads
90
Without caching
$3.0000
With caching
$1.1160

Savings: $1.8840

Implementation checklist

Put static content first

Cache matching works on prefixes. Place your system prompt and few-shot examples before any dynamic content so the prefix stays stable across requests.

Mind the minimum token count

Check the minimum cacheable prefix length for the exact model. A prefix can be stable yet too short to qualify.

Understand the TTL

Model the provider's cache lifetime, refresh behavior and routing. A prefix reused after expiry may require a new write. Measure misses rather than assuming every repeated prompt hits.

Monitor cache hit rates

Check the usage fields in API responses (cache_creation_input_tokens vs cache_read_input_tokens) to verify caching is working. Low hit rates mean your prefix is changing too often.