inference-tradeoffs-explained

Chapter 09 · Quantisation

Quantisation

Fewer bytes per weight and per KV value, and faster matmuls where the hardware has units for the format. Below, the mechanism as the simulator ran it, step by step; then what it measured, why it behaves that way, and the papers.

Loading the bytes-per-step animation…

Its row of the matrix

Loading the matrix…

Fewer bytes per weight, faster matmuls where the units exist

Quantisation does two different things, and the simulator keeps them apart. A narrower storage format means fewer bytes per weight or per KV value: every memory-bound pass reads less, and the same HBM holds more KV. A narrower compute format runs the matmuls on tensor-core units for that format, at twice (FP8, INT8) or four times (FP4, on B200) the BF16 rate; where a device has no such units the weights are dequantised to BF16 first, so compute-bound passes are no faster (weight-only, written W8A16 or W4A16). INT4 and FP4 bytes include their block scales (AWQ, Microscaling (MX) formats).

Why decode and prefill gain differently

A decode step reads every weight for a few tokens, so its time is its bytes over the HBM bandwidth (the roofline's left side), and halving the bytes nearly halves it. A prefill step is compute-bound, so only a faster matmul format helps it: on four H100s FP8 weights alone leave an 8,192-token prefill at 564.3 ms, while FP8 matmuls (W8A8) cut it to 282.4 ms.

DeviceFormatWeights GBKV tokensDecode b=64Prefill 8,192J/token at b=64 (incl. idle)
H100BF16 141.1 448,288 17.5 ms 564.3 ms 0.424
H100FP8 weights only (W8A16) 70.6 663,597 11.0 ms 564.3 ms 0.319
H100FP8 W8A8 70.6 663,597 11.0 ms 282.4 ms 0.246
H100INT4 weights only (W4A16) 36.4 767,887 7.9 ms 564.3 ms 0.267
H100FP8 KV cache 141.1 896,577 15.5 ms 564.3 ms 0.392
H100FP8 W8A8 + FP8 KV 70.6 1,327,194 9.0 ms 282.4 ms 0.214
B200BF16 141.1 1,546,921 7.6 ms 248.3 ms 0.281
B200FP4 W4A4 37.5 1,863,156 3.6 ms 62.5 ms 0.110
Llama-3-70B on four GPUs (no all-reduce, to isolate the format), recorded in results.md section 24; ✓ / ✗ against BF16 on H100. Accuracy is not simulated.

What the sweep measured

FP8 weights and matmuls (W8A8) help every workload, by as much as any single lever outside prefix caching: goodput per GPU +89% on chat, +243% on the coding agent and +385% on voice, with TTFT and TPOT both falling. INT4 weights with BF16 matmuls help decode (TPOT p99 chat -48%) but not prefill (TTFT p99 +8%), as the roofline predicts. An FP8 KV cache alone does little here (chat +7%), because the sweep's KV memory never binds (chapter 2). On B200, FP4 weights and matmuls (W4A4) add +100% over B200's own BF16 baseline on chat.

Accuracy is not simulated. Every one of these numbers assumes the quantised model is good enough; whether it is depends on the model, the method and the format, which is what Numerics Explained is about.

What the sweep measured, lever by lever

FP8 weights and matmuls (W8A8)

Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live

Workloadgoodput$/M tokJ/tokTTFT p99TPOT p99ITL p99
Chat+89% -46% $-50% -52% -47% -60%
Coding agent+243% E-71% E$-63% E-50% -44% -35%
Offline batch+61% -38% $-46% -47% -40% -86%
Long-context RAG+124% E-55% E$-54% E-47% -51% -35%
Real-time voice+385% E-79% E$-70% E-44% -45% -35%

INT4 weights, BF16 matmuls (W4A16)

Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live

Workloadgoodput$/M tokJ/tokTTFT p99TPOT p99ITL p99
Chat+32% -24% $-18% +8% -48% -43%
Coding agent+41% -29% $-26% +0% -37% -51%
Offline batch+20% -16% $-9% -6% -3% -87%
Long-context RAG+33% -25% $-19% +6% -37% -50%
Real-time voice+152% E-60% E$-49% E+5% -47% -52%

FP8 KV cache

Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live

Workloadgoodput$/M tokJ/tokTTFT p99TPOT p99ITL p99
Chat+7% -6% $-4% +14% -4% -3%
Coding agent+6% -6% $-3% -2% -1% -3%
Offline batch+4% -3% $-2% -6% -2% -86%
Long-context RAG+0% +0% $-0% +0% -1% -5%
Real-time voice-6% +6% $+3% -2% -5% -1%

FP8 weights, matmuls and KV

Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live

Workloadgoodput$/M tokJ/tokTTFT p99TPOT p99ITL p99
Chat+113% E-53% E$-54% E-51% -48% -61%
Coding agent+266% E-73% E$-64% E-49% -46% -37%
Offline batch+70% -41% $-48% -47% -43% -89%
Long-context RAG+142% E-59% E$-56% E-46% -53% -38%
Real-time voice+359% E-78% E$-70% E-44% -45% -36%

FP4 weights and matmuls (W4A4)

Change from the baseline on B200 (it needs FP4 units), at capacity for goodput, cost and energy, at the reference load for latency. Run it live

Workloadgoodput$/M tokJ/tokTTFT p99TPOT p99ITL p99
Chat+100% EB-50% E$B-66% EB-67% B-45% B-63% B
Coding agent+210% EB-68% E$B-73% EB-69% B-48% B-37% B
Offline batch+98% B-49% $B-67% B-64% B-58% B-23% B
Long-context RAG+200% EB-67% E$B-71% EB-71% B-49% B-36% B
Real-time voice+251% EB-71% E$B-76% EB-61% B-43% B-37% B

Caveats on these numbers

  • EHuge percentages come from a baseline at its SLO edge. Capacity is the highest load at which 90% of requests meet both SLOs. Where the baseline sits just past a latency cliff (voice and the coding agent especially), a lever that pulls latency back under the SLO multiplies capacity, so gains of hundreds or thousands of percent are real in the model but say more about the cliff than about the lever. Marked wherever capacity at least doubles; read the absolute numbers too.
  • $Prices per GPU-hour are illustrative. Cost per million tokens uses round illustrative prices (H100 $3.00, H200 $3.50, B200 $5.00 per GPU-hour), not quotes. Cost scales linearly with them, so the ranking of levers on one device does not depend on them; comparisons across devices do.
  • BB200 BF16 rate and power are illustrative. B200's BF16 rate is taken as half its FP8 datasheet rate, and its power coefficients (1,000 W board power, 140 W idle, energy per FLOP and per byte) are illustrative.

✓ better than the baseline, ✗ worse, by more than ±2%. Every lever on every metric and device: the matrix; on any two metrics: the explorer.

Papers and sources