Memory budget, with the assumptions visible
Weights use total parameters, including all MoE experts. KV memory uses KV heads, not the hidden dimension. This estimates standard MHA/GQA/MQA caches; MLA, hybrid recurrent layers and runtime allocation need model-specific measurements.
Q labels are illustrative storage budgets: bits × overhead. Actual GGUF files contain mixed tensor types. Prefer the measured file size when selecting hardware.
- Weights
- 3.75 GiB
- KV cache
- 1.00 GiB
- Runtime reserve (GiB)
- 1.00 GiB
- Total incl. reserve
- 5.75 GiB
Within this budget (24 GiB)
GiB = 2³⁰ bytes; GB = 10⁹ bytes. Shared memory must also accommodate the OS. Reserve is an editable assumption, not a measured runtime overhead.
KV bytes = 2 × layers × KV heads × head dimension × tokens × batch × dtype bytesWhy memory size does not predict speed
A model fitting in memory is only a capacity check. Decode time also depends on active weights, KV traffic, compute, kernels, batch size and placement. There is no universal GPU efficiency factor. Measure prefill and decode separately with the exact model and runtime.
Smart Offloading for MoE Models
Expert routing varies by token. Weights that are inactive now may be needed later. CPU expert placement can reduce VRAM use, but requires CPU computation or transfers; it is not free.
llama.cpp offers --cpu-moe and --n-cpu-moe to place expert weights on the CPU, and --gpu-layers for layer placement. --override-kv changes model metadata. It does not enable expert prediction or prefetching.