/method
How the sweep was run
Every number on this site comes from one recorded sweep of Disaggregated_Inference_Sim, a discrete-event simulator of LLM serving with a roofline cost model and a power model: examples/tradeoffs.py wrote examples/tradeoffs.json at simulator commit 0caa090, 350 configurations in 56 minutes on 2 workers. This site vendors that file, the simulator's JavaScript engine and its results.md byte for byte at commit 69ffd2b (the code is unchanged since the sweep ran).
The cluster and the levers
Every configuration is the same 8 GPUs serving Llama-3-70B. The baseline is two instances of TP4 colocated, prefill-priority, reserved KV, BF16; each lever changes that (the combined rows change several things at once):
| Family | Lever |
|---|---|
| Baseline | 2 x TP4 colocated, prefill-priority, reserved KV, BF16 |
| Batching | Decode-priority batching |
| Batching | Chunked prefill, 512-token budget |
| Batching | Chunked prefill, 2,048-token budget |
| KV memory | Paged KV, preempt by recompute |
| KV memory | Paged KV, preempt by swap to host |
| Prefix caching | Prefix caching (paged KV) |
| Disaggregation | Disaggregated 1P1D (TP4 each) |
| Disaggregation | Disaggregated 2P (TP2) + 1D (TP4) |
| Disaggregation | Disaggregated 1P1D, paged decode, prefix-cached prefill |
| Parallelism | 1 x TP8 |
| Parallelism | 4 x TP2 |
| Parallelism | 2 x (TP2, PP2), 2 micro-batches |
| Parallelism | 2 x (TP2, PP2), no micro-batching |
| Quantisation | FP8 weights and matmuls (W8A8) |
| Quantisation | INT4 weights, BF16 matmuls (W4A16) |
| Quantisation | FP8 KV cache |
| Quantisation | FP8 weights, matmuls and KV |
| Quantisation | FP4 weights and matmuls (W4A4) (B200 only) |
| Speculative decoding | Speculative, MTP head, gamma 3, alpha 0.7 |
| Speculative decoding | Speculative, Llama-3.2-1B draft, gamma 4, alpha 0.7 |
| Combined | Chunked 2048 + paged + prefix cache + FP8 (W8A8, KV) |
| Combined | Disaggregated 1P1D + paged + prefix cache + FP8 |
| Combined | Modern colocated + speculative (MTP) |
What is measured
- Capacity: the highest arrival rate at which at least 90% of requests meet both SLOs (time to first token and time per output token), DistServe's goodput (arXiv:2401.09670), found by doubling and then bisection; rejected requests count as misses. At capacity the sweep records goodput and output tokens per second per GPU, cost per million output tokens (at the illustrative prices below) and joules per output token (the power model, static power included).
- Latency: TTFT, TPOT and ITL p50 and p99, and KV occupancy, at each workload's reference load: half the H100 baseline's capacity, the same offered load for every lever.
- Multi-turn workloads run closed-loop sessions; offline batch is measured saturated (every request at once). One model, one seed (1), and capacity is found to a tolerance, so results near an SLO cliff are noisy. Accuracy is not simulated: what a number format does to accuracy is on Numerics Explained.
The metrics, and which way is better:
| Metric | Meaning | Better |
|---|---|---|
| Goodput / GPU | requests per second per GPU meeting both SLOs, at capacity | higher |
| Tokens/s / GPU | output tokens per second per GPU, at capacity | higher |
| $ / M tokens | dollars per million output tokens at capacity (illustrative $/GPU-hour) | lower |
| J / token | joules per output token at capacity (the power model, static power included) | lower |
| TTFT p50 | median time to first token at the reference load | lower |
| TTFT p99 | 99th-percentile time to first token at the reference load | lower |
| TPOT p50 | median time per output token at the reference load | lower |
| TPOT p99 | 99th-percentile time per output token at the reference load | lower |
| ITL p50 | median inter-token latency at the reference load | lower |
| ITL p99 | 99th-percentile inter-token latency at the reference load | lower |
| KV peak | peak KV-cache occupancy at the reference load (lower leaves more headroom) | lower |
The workloads
Each is a named, parameterised distribution in the simulator, with its source or rationale:
| Workload | Prompt (mean) | Output (mean) | Turns | SLOs (TTFT, TPOT) |
|---|---|---|---|---|
| Chat | 161 | 338 | 3 | 1,000 ms, 50.0 ms |
| Coding agent | 512 | 160 | 8 | 1,500 ms, 40.0 ms |
| Offline batch | 2,048 | 256 | 1 | 60,000 ms, 500.0 ms |
| Long-context RAG | 7,904 | 230 | 1 | 5,000 ms, 50.0 ms |
| Real-time voice | 64 | 48 | 6 | 300.0 ms, 25.0 ms |
- Chat. ShareGPT turn lengths (means 161 in, 338 out, as vLLM's evaluation, arXiv:2309.06180 section 6.1); three turns 10 s apart behind one of eight 512-token system prompts (our choice). SLOs: 1 s to the first token, 50 ms a token (20 tokens/s, faster than reading); DistServe's chatbot SLOs on A100s were 0.25-4 s and 0.1-0.2 s (arXiv:2401.09670, Table 1).
- Coding agent. An agent loop: a 6,144-token prefix (tools, instructions, repository context) shared by every session, eight turns of tool output (mean 512) and short actions (mean 160), 2 s of tool time between turns: long prompts, short outputs, heavy prefix reuse (our choice of numbers).
- Offline batch. Bulk summarisation or labelling: nobody is waiting, so the SLOs are loose (60 s, 0.5 s a token) and throughput and cost per token decide (our choice).
- Long-context RAG. Retrieved documents make the prompt long: lengths fitted to arxiv_summarization's median and 90th percentile (7,059 / 12,985 in, 208 / 371 out; Sarathi-Serve, arXiv:2403.02310, Table 2; lognormal means 7,904 and 230, cv 0.50 and 0.48), behind one of four 1,024-token instruction prefixes. SLOs: 5 s, 50 ms (DistServe's summarisation: 15 s, 0.15 s on A100s).
- Real-time voice. A spoken conversation: short utterances and replies, six turns 3 s apart, a 1,024-token persona prompt. People minimise the silence between turns (Stivers et al., PNAS 2009, doi:10.1073/pnas.0903616106), so the first token gets 300 ms and the stream 25 ms a token to keep speech synthesis fed (our choice).
The devices
| Device | BF16 TFLOP/s | HBM TB/s | HBM GB | Faster formats | $/GPU-hour (illustrative) |
|---|---|---|---|---|---|
| H100-SXM | 989 | 3.35 | 80 | FP8 2×, INT8 2× | $3.00 |
| H200-SXM | 989 | 4.80 | 141 | FP8 2×, INT8 2× | $3.50 |
| B200 | 2,250 | 8.00 | 180 | FP8 2×, FP4 4× | $5.00 |
Caveats
These are marked wherever the numbers they affect appear (the letter beside a matrix cell, the list under each chart):
PPipeline parallelism is not overlapped
The simulator does not keep several batches in flight across pipeline stages, so PP shows only its cost (bubbles, weight re-reads, an activation hop per stage) and every PP configuration in the sweep loses. Real engines overlap micro-batches across steps.
TTP all-reduce cost is pessimistic at small batch
Tensor-parallel all-reduces use a first-order α–β ring model whose latency term (5 µs a hop, illustrative) dominates at small batch; NVSwitch hardware with in-switch reduction does better. It has not been calibrated against nccl-tests, and communication does not overlap with compute.
MPaged KV shows +0% where memory does not bind
Llama-3-70B on four GPUs per instance never runs out of KV cache at these loads, so paged allocation has nothing to win and the sweep shows no change. Paged KV matters where memory binds: the simulator's results.md section 23 shows it on OPT-13B.
EHuge percentages come from a baseline at its SLO edge
Capacity is the highest load at which 90% of requests meet both SLOs. Where the baseline sits just past a latency cliff (voice and the coding agent especially), a lever that pulls latency back under the SLO multiplies capacity, so gains of hundreds or thousands of percent are real in the model but say more about the cliff than about the lever. Marked wherever capacity at least doubles; read the absolute numbers too.
$Prices per GPU-hour are illustrative
Cost per million tokens uses round illustrative prices (H100 $3.00, H200 $3.50, B200 $5.00 per GPU-hour), not quotes. Cost scales linearly with them, so the ranking of levers on one device does not depend on them; comparisons across devices do.
αSpeculative acceptance α = 0.7 is assumed
Speculative decoding's gain depends on how often the target accepts a drafted token, which depends on the draft, the target and the text. The sweep fixes α at 0.7 as a parameter; it is not measured.
BB200 BF16 rate and power are illustrative
B200's BF16 rate is taken as half its FP8 datasheet rate, and its power coefficients (1,000 W board power, 140 W idle, energy per FLOP and per byte) are illustrative.
Provenance
The vendored files, their commit and their SHA-256 (src/lib/tradeoffs/vendor/VENDORED.json, checked by the unit tests and re-read from the simulator's repository in CI):
- web/sim_engine.js
c7c1d2bdb0b19df050c0a0c79feec7b183703d3622da5fb3b97c1fa8708e0478 - examples/tradeoffs.json
25202e8781b301c5551ee949017b19be4ce0d59a8b537cbd352d8458c9ed37a5 - examples/results.md
4d288746ebecc3d0882be2edc9a2a412cf9d4333985efd9349034effbb08dba5