Chapter 09 · Quantisation
Quantisation
Fewer bytes per weight and per KV value, and faster matmuls where the hardware has units for the format. Below, the mechanism as the simulator ran it, step by step; then what it measured, why it behaves that way, and the papers.
Loading the bytes-per-step animation…
Its row of the matrix
Loading the matrix…
Fewer bytes per weight, faster matmuls where the units exist
Quantisation does two different things, and the simulator keeps them apart. A narrower storage format means fewer bytes per weight or per KV value: every memory-bound pass reads less, and the same HBM holds more KV. A narrower compute format runs the matmuls on tensor-core units for that format, at twice (FP8, INT8) or four times (FP4, on B200) the BF16 rate; where a device has no such units the weights are dequantised to BF16 first, so compute-bound passes are no faster (weight-only, written W8A16 or W4A16). INT4 and FP4 bytes include their block scales (AWQ, Microscaling (MX) formats).
Why decode and prefill gain differently
A decode step reads every weight for a few tokens, so its time is its bytes over the HBM bandwidth (the roofline's left side), and halving the bytes nearly halves it. A prefill step is compute-bound, so only a faster matmul format helps it: on four H100s FP8 weights alone leave an 8,192-token prefill at 564.3 ms, while FP8 matmuls (W8A8) cut it to 282.4 ms.
| Device | Format | Weights GB | KV tokens | Decode b=64 | Prefill 8,192 | J/token at b=64 (incl. idle) |
|---|---|---|---|---|---|---|
| H100 | BF16 | 141.1 | 448,288 | 17.5 ms | 564.3 ms | 0.424 |
| H100 | FP8 weights only (W8A16) | 70.6 | 663,597 | 11.0 ms | 564.3 ms | 0.319 |
| H100 | FP8 W8A8 | 70.6 | 663,597 | 11.0 ms | 282.4 ms | 0.246 |
| H100 | INT4 weights only (W4A16) | 36.4 | 767,887 | 7.9 ms | 564.3 ms | 0.267 |
| H100 | FP8 KV cache | 141.1 | 896,577 | 15.5 ms | 564.3 ms | 0.392 |
| H100 | FP8 W8A8 + FP8 KV | 70.6 | 1,327,194 | 9.0 ms | 282.4 ms | 0.214 |
| B200 | BF16 | 141.1 | 1,546,921 | 7.6 ms | 248.3 ms | 0.281 |
| B200 | FP4 W4A4 | 37.5 | 1,863,156 | 3.6 ms | 62.5 ms | 0.110 |
What the sweep measured
FP8 weights and matmuls (W8A8) help every workload, by as much as any single lever outside prefix caching: goodput per GPU +89% on chat, +243% on the coding agent and +385% on voice, with TTFT and TPOT both falling. INT4 weights with BF16 matmuls help decode (TPOT p99 chat -48%) but not prefill (TTFT p99 +8%), as the roofline predicts. An FP8 KV cache alone does little here (chat +7%), because the sweep's KV memory never binds (chapter 2). On B200, FP4 weights and matmuls (W4A4) add +100% over B200's own BF16 baseline on chat.
Accuracy is not simulated. Every one of these numbers assumes the quantised model is good enough; whether it is depends on the model, the method and the format, which is what Numerics Explained is about.
What the sweep measured, lever by lever
FP8 weights and matmuls (W8A8)
Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live
| Workload | goodput | $/M tok | J/tok | TTFT p99 | TPOT p99 | ITL p99 |
|---|---|---|---|---|---|---|
| Chat | +89% | -46% $ | -50% | -52% | -47% | -60% |
| Coding agent | +243% E | -71% E$ | -63% E | -50% | -44% | -35% |
| Offline batch | +61% | -38% $ | -46% | -47% | -40% | -86% |
| Long-context RAG | +124% E | -55% E$ | -54% E | -47% | -51% | -35% |
| Real-time voice | +385% E | -79% E$ | -70% E | -44% | -45% | -35% |
INT4 weights, BF16 matmuls (W4A16)
Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live
| Workload | goodput | $/M tok | J/tok | TTFT p99 | TPOT p99 | ITL p99 |
|---|---|---|---|---|---|---|
| Chat | +32% | -24% $ | -18% | +8% | -48% | -43% |
| Coding agent | +41% | -29% $ | -26% | +0% | -37% | -51% |
| Offline batch | +20% | -16% $ | -9% | -6% | -3% | -87% |
| Long-context RAG | +33% | -25% $ | -19% | +6% | -37% | -50% |
| Real-time voice | +152% E | -60% E$ | -49% E | +5% | -47% | -52% |
FP8 KV cache
Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live
| Workload | goodput | $/M tok | J/tok | TTFT p99 | TPOT p99 | ITL p99 |
|---|---|---|---|---|---|---|
| Chat | +7% | -6% $ | -4% | +14% | -4% | -3% |
| Coding agent | +6% | -6% $ | -3% | -2% | -1% | -3% |
| Offline batch | +4% | -3% $ | -2% | -6% | -2% | -86% |
| Long-context RAG | +0% | +0% $ | -0% | +0% | -1% | -5% |
| Real-time voice | -6% | +6% $ | +3% | -2% | -5% | -1% |
FP8 weights, matmuls and KV
Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live
| Workload | goodput | $/M tok | J/tok | TTFT p99 | TPOT p99 | ITL p99 |
|---|---|---|---|---|---|---|
| Chat | +113% E | -53% E$ | -54% E | -51% | -48% | -61% |
| Coding agent | +266% E | -73% E$ | -64% E | -49% | -46% | -37% |
| Offline batch | +70% | -41% $ | -48% | -47% | -43% | -89% |
| Long-context RAG | +142% E | -59% E$ | -56% E | -46% | -53% | -38% |
| Real-time voice | +359% E | -78% E$ | -70% E | -44% | -45% | -36% |
FP4 weights and matmuls (W4A4)
Change from the baseline on B200 (it needs FP4 units), at capacity for goodput, cost and energy, at the reference load for latency. Run it live
| Workload | goodput | $/M tok | J/tok | TTFT p99 | TPOT p99 | ITL p99 |
|---|---|---|---|---|---|---|
| Chat | +100% EB | -50% E$B | -66% EB | -67% B | -45% B | -63% B |
| Coding agent | +210% EB | -68% E$B | -73% EB | -69% B | -48% B | -37% B |
| Offline batch | +98% B | -49% $B | -67% B | -64% B | -58% B | -23% B |
| Long-context RAG | +200% EB | -67% E$B | -71% EB | -71% B | -49% B | -36% B |
| Real-time voice | +251% EB | -71% E$B | -76% EB | -61% B | -43% B | -37% B |
Caveats on these numbers
- EHuge percentages come from a baseline at its SLO edge. Capacity is the highest load at which 90% of requests meet both SLOs. Where the baseline sits just past a latency cliff (voice and the coding agent especially), a lever that pulls latency back under the SLO multiplies capacity, so gains of hundreds or thousands of percent are real in the model but say more about the cliff than about the lever. Marked wherever capacity at least doubles; read the absolute numbers too.
- $Prices per GPU-hour are illustrative. Cost per million tokens uses round illustrative prices (H100 $3.00, H200 $3.50, B200 $5.00 per GPU-hour), not quotes. Cost scales linearly with them, so the ranking of levers on one device does not depend on them; comparisons across devices do.
- BB200 BF16 rate and power are illustrative. B200's BF16 rate is taken as half its FP8 datasheet rate, and its power coefficients (1,000 W board power, 140 W idle, energy per FLOP and per byte) are illustrative.
✓ better than the baseline, ✗ worse, by more than ±2%. Every lever on every metric and device: the matrix; on any two metrics: the explorer.
Papers and sources
- Lin, Tang, Tang et al., AWQ: Activation-aware Weight Quantization for LLM Compression and Acceleration (arXiv:2306.00978): INT4 weights in groups of 128 with a scale each (the simulator's INT4 bytes)
- Rouhani, Zhao, More et al., Microscaling Data Formats for Deep Learning (arXiv:2310.10537): MXFP4: blocks of values sharing one scale (the simulator's FP4 bytes)
- NVIDIA, H100 Tensor Core GPU: datasheet figures the simulator's H100 preset uses
- NVIDIA, DGX B200: per-GPU figures the simulator's B200 preset uses (its power coefficients are illustrative)