Quantization

Expert

Represent weights with fewer bits, inspect the resulting error, and separate storage savings from model quality.

Last updated: Sep 13, 2026

Precision changes representation

Quantization maps continuous or high-precision values to a finite set of representable values. It can reduce weight storage and memory traffic. Speed depends on kernels, hardware, batch size and whether weights or activations are quantized. Fewer bits do not imply a fixed speedup or quality loss.

Round actual weights

A uniform toy quantizer with fixed example weights. Change the bit width and clipping range to inspect rounding and clipping error. This does not predict model accuracy or perplexity.

Representable values

16

Mean squared weight error

0.01806

Original weightInteger codeReconstructedError
-2.4002-2.2000.200
-1.8003-1.8000.000
-0.7006-0.6000.100
-0.2507-0.2000.050
0.00080.2000.200
0.30080.200-0.100
0.900101.0000.100
1.600111.400-0.200
2.700142.600-0.100

Δ = (max − min) / (2ᵇ − 1); q = clamp(round((w − min) / Δ), 0, 2ᵇ − 1); ŵ = min + qΔ.

The endpoints are included, giving exactly 2ᵇ available values. Zero need not be one of them in this simplified grid. Production formats use other grids, group scales and calibration methods. Weight error alone is not a measure of task quality.

Raw weight storage: 35.00 GB

Decimal GB. This counts only packed weights: parameters × bits / 8. Scales, metadata, activations, runtime buffers and the KV cache need additional memory. CPU offloading can reduce GPU residency at a latency cost.

Different methods, different assumptions

Post-training quantization (PTQ)

Convert a trained model. Some methods use calibration data to choose scales or reduce output error. Results depend on the method and calibration distribution.

Quantization-aware training (QAT)

Expose the model to quantization effects during training or fine-tuning. Full training from scratch is not required by the definition; gradient handling and the quantizer depend on the recipe.

QLoRA

Keep the base model quantized and frozen while training low-rank adapters. This is a fine-tuning method, not a guarantee that every quantized inference model keeps its original quality.

GGUF formats

GGUF is a storage format. Names such as Q4_K_M refer to particular quantization recipes; nominal bit width is not an exact total-memory calculation. Inspect the actual artifact and runtime.

What fits in memory?

A 70B model requires 140 GB for raw 16-bit weights or 35 GB for raw 4-bit weights, before scales and runtime overhead. Thus 70B at 4 bits does not fit entirely in 24 GB of GPU memory. CPU offloading or a different compression scheme may make execution possible, with different latency and quality tradeoffs.

Measure quality for the intended task

Compare a specified base checkpoint and quantized artifact using the same prompts, decoding settings and evaluation data. Report task scores, failure cases and latency. A small change in weight error or perplexity cannot certify unchanged coding, reasoning or factual performance. There is no universal percentage of retained accuracy for INT4.

Sources and next steps

Continue to LoRA and adapter training →