VRAM Calculator

Beginner

Calculate a local LLM memory budget from explicit weight, cache and runtime assumptions.

Last updated: Sep 13, 2026

Memory budget, with the assumptions visible

Weights use total parameters, including all MoE experts. KV memory uses KV heads, not the hidden dimension. This estimates standard MHA/GQA/MQA caches; MLA, hybrid recurrent layers and runtime allocation need model-specific measurements.

Q labels are illustrative storage budgets: bits × overhead. Actual GGUF files contain mixed tensor types. Prefer the measured file size when selecting hardware.

Weights
3.75 GiB
KV cache
1.00 GiB
Runtime reserve (GiB)
1.00 GiB
Total incl. reserve
5.75 GiB

Within this budget (24 GiB)

GiB = 2³⁰ bytes; GB = 10⁹ bytes. Shared memory must also accommodate the OS. Reserve is an editable assumption, not a measured runtime overhead.

KV bytes = 2 × layers × KV heads × head dimension × tokens × batch × dtype bytes

Why memory size does not predict speed

A model fitting in memory is only a capacity check. Decode time also depends on active weights, KV traffic, compute, kernels, batch size and placement. There is no universal GPU efficiency factor. Measure prefill and decode separately with the exact model and runtime.

Smart Offloading for MoE Models

Expert routing varies by token. Weights that are inactive now may be needed later. CPU expert placement can reduce VRAM use, but requires CPU computation or transfers; it is not free.

llama.cpp offers --cpu-moe and --n-cpu-moe to place expert weights on the CPU, and --gpu-layers for layer placement. --override-kv changes model metadata. It does not enable expert prediction or prefetching.