Home / Chapter 7 · Production and serving
Last edited · 6 min read
Quantisation
Quantisation stores the weights, and sometimes the activations and the KV cache too, in fewer bits by rounding them to a coarser grid of values. The model takes 2–4 times less memory, and generation gets faster because decode reads fewer bytes from memory.
In plain wordsA photo saved as a JPEG instead of at full quality: it is several times smaller and looks the same on a phone screen. Artefacts appear only under heavy compression, and first in the fine detail. For a model, the fine detail is long reasoning, code and less common languages.
Lower the bit count, add an outlier, then turn on a separate scale per block. Watch the rounding error
40 weights from a normal distribution, a grid symmetric around zero, a block is 8 weights. The error leaves out the outlier itself, so you can see what happens to the rest.
Pick a model size and see what hardware the weights alone fit on. Change the bit count in the widget above
Counted as the weights plus 15% headroom, enough for one person’s short context. Serving many conversations needs much more memory for the KV cache. Block scales add a few per cent.
How it works
- Each weight is divided by a scale, rounded to a number from a small range (−8 to 7 in INT4; a symmetric grid, as in the widget, uses −7 to 7, i.e. 15 levels) and stored like that. At compute time it is multiplied back by the scale. The error is rounding noise, and the network tolerates small noise well.
- Outliers are the problem. One large value stretches the grid, and the rest of the weights land on a few points around zero. That is why the scale is computed separately for small groups, typically 16–128 weights. In the activations of models from about 6.7 billion parameters up, systematic outliers appear in a few channels (Dettmers et al., LLM.int8(), 2022).
- Post-training quantisation (PTQ) works on a finished model in minutes or hours. GPTQ and AWQ use a small sample of calibration data to round where it hurts the layer’s output least. Quantisation-aware training (QAT) copes better with 4 bits and below, but it requires training, so it is usually done by the model’s author.
Formats and what they speed up
- Weight-only quantisation, e.g. W4A16 (weights in 4 bits, activations in 16), reduces the bytes read in decode, so it speeds up generation at small batch sizes. The multiplication still runs in 16 bits, so prefill and large batches gain almost nothing.
- FP8 for weights and activations (W8A8) computes on the tensor cores in FP8: an H100 reaches about 1,979 TFLOPS dense in it versus 989 in BF16. So it also speeds up prefill and large batches. Kurtic et al. (2024) measure FP8 W8A8 as practically lossless, INT8 W8A8 at 1–3% loss, and W4A16 as closer to 8 bits than expected.
- The new 4-bit formats are floating-point numbers with a scale per block. MXFP4 uses blocks of 32 weights with a power-of-two scale, about 4.25 bits per weight; OpenAI released the expert weights of gpt-oss in it. NVFP4 uses blocks of 16 with an FP8 scale plus an extra per-tensor scale, about 4.5 bits, and is computed in hardware on Blackwell GPUs.
- The KV cache gets quantised too. An FP8 cache takes half the memory per token, so with long contexts it frees more room for the batch than squeezing the weights further (see “The generation loop and KV cache”).
What you lose and how to decide
- Rule of thumb, depending on the model and method: 8 bits is practically lossless, 4 bits with a good method is a small loss, 3 bits and below drop off fast. Losses show up first in long reasoning, code, maths and less common languages, and a benchmark average can hide them.
- For the same memory, a larger model in 4 bits usually beats a smaller one in 16. After more than 35,000 experiments, Dettmers and Zettlemoyer (2022) found 4 bits almost always optimal for a fixed total number of model bits.
- The choice depends on traffic and hardware. On Hopper GPUs such as the H100, Kurtic et al. recommend W4A16 for single requests and small batches and FP8 W8A8 for heavy traffic with continuous batching. On Blackwell, NVFP4 for weights and activations (W4A4) runs on the FP4 tensor cores at 2–3× the FP8 rate, so 4 bits can win under heavy traffic too, at a somewhat larger loss than FP8: Red Hat measures about 99% of BF16 accuracy for 70B+ models and 95–98% for 7–14B ones. Check quality on your own eval set, because the same model can be quantised differently by different hosts (see “Open-weight models” and “Evals”).
Check yourself
What does quantisation give you, what do you lose, and which format would you choose?
The weights are stored in fewer bits, from 16 down to 8 or 4, with a scale computed per small block, because single outliers stretch the grid. The model needs 2–4 times less memory, and decode gets faster because it reads fewer bytes. Weight-only quantisation such as W4A16 helps at small batch sizes. FP8 for weights and activations also speeds up the maths, so on Hopper it wins under heavy traffic; on Blackwell, NVFP4 for weights and activations is faster still. 8 bits is practically lossless, 4 bits costs a little, and below that quality drops fast, first in reasoning and code. Your own evals decide.
Po polsku
Wagi zapisuje się mniejszą liczbą bitów, z 16 do 8 albo 4, ze skalą liczoną dla małych bloków, bo pojedyncze wartości odstające rozciągają siatkę. Model zajmuje 2–4 razy mniej pamięci, a decode przyspiesza, bo czyta mniej bajtów. Kwantyzacja samych wag, jak W4A16, pomaga przy małym batchu. FP8 dla wag i aktywacji przyspiesza też liczenie, więc na Hopperze wygrywa przy dużym ruchu; na Blackwellu jeszcze szybsze jest NVFP4 dla wag i aktywacji. 8 bitów jest praktycznie bezstratne, 4 bity to niewielka strata, niżej jakość szybko spada, najpierw w rozumowaniu i kodzie. Rozstrzygają własne ewaluacje.
Follow-up questions (4)
- PTQ or QAT?
- PTQ quantises a finished model in minutes or hours, using a small sample of calibration data (GPTQ, AWQ). QAT simulates quantisation during training, so the model learns to tolerate it and holds its quality better at 4 bits and below. The cost is training, so QAT is usually done by the model’s author.
- Why are activations harder to quantise than weights?
- Weights are known in advance and can be carefully rescaled offline. Activations depend on the input and have large outliers in a few channels. Methods that shift the difficulty onto the weights (SmoothQuant) help, as does FP8, which has a wider range than INT8.
- After quantising to 4 bits the benchmarks look fine, but users complain about code. What do you do?
- A benchmark average hides losses in long reasoning chains and code. Compare both versions on your own eval set from that area. If the loss is real, go back to 8 bits or keep the sensitive layers at higher precision.
- When does quantising the KV cache give more than quantising the weights?
- With long contexts and large batches, when the cache of all conversations takes more memory than the weights. An FP8 cache holds twice as many tokens, i.e. twice as many conversations or twice the context length.
Sources
- Maarten Grootendorst: A Visual Guide to Quantization
- Kurtic et al.: “Give Me BF16 or Give Me Death”? Accuracy-Performance Trade-Offs in LLM Quantization
- Dettmers and Zettlemoyer: The case for 4-bit precision (2022)
- NVIDIA: Introducing NVFP4
- Red Hat: Accelerating large language models with NVFP4 quantization (2026)