inference-tradeoffs-explained

/method

How the sweep was run

Every number on this site comes from one recorded sweep of Disaggregated_Inference_Sim, a discrete-event simulator of LLM serving with a roofline cost model and a power model: examples/tradeoffs.py wrote examples/tradeoffs.json at simulator commit 0caa090, 350 configurations in 56 minutes on 2 workers. This site vendors that file, the simulator's JavaScript engine and its results.md byte for byte at commit 69ffd2b (the code is unchanged since the sweep ran).

The cluster and the levers

Every configuration is the same 8 GPUs serving Llama-3-70B. The baseline is two instances of TP4 colocated, prefill-priority, reserved KV, BF16; each lever changes that (the combined rows change several things at once):

FamilyLever
Baseline2 x TP4 colocated, prefill-priority, reserved KV, BF16
BatchingDecode-priority batching
BatchingChunked prefill, 512-token budget
BatchingChunked prefill, 2,048-token budget
KV memoryPaged KV, preempt by recompute
KV memoryPaged KV, preempt by swap to host
Prefix cachingPrefix caching (paged KV)
DisaggregationDisaggregated 1P1D (TP4 each)
DisaggregationDisaggregated 2P (TP2) + 1D (TP4)
DisaggregationDisaggregated 1P1D, paged decode, prefix-cached prefill
Parallelism1 x TP8
Parallelism4 x TP2
Parallelism2 x (TP2, PP2), 2 micro-batches
Parallelism2 x (TP2, PP2), no micro-batching
QuantisationFP8 weights and matmuls (W8A8)
QuantisationINT4 weights, BF16 matmuls (W4A16)
QuantisationFP8 KV cache
QuantisationFP8 weights, matmuls and KV
QuantisationFP4 weights and matmuls (W4A4) (B200 only)
Speculative decodingSpeculative, MTP head, gamma 3, alpha 0.7
Speculative decodingSpeculative, Llama-3.2-1B draft, gamma 4, alpha 0.7
CombinedChunked 2048 + paged + prefix cache + FP8 (W8A8, KV)
CombinedDisaggregated 1P1D + paged + prefix cache + FP8
CombinedModern colocated + speculative (MTP)

What is measured

The metrics, and which way is better:

MetricMeaningBetter
Goodput / GPUrequests per second per GPU meeting both SLOs, at capacityhigher
Tokens/s / GPUoutput tokens per second per GPU, at capacityhigher
$ / M tokensdollars per million output tokens at capacity (illustrative $/GPU-hour)lower
J / tokenjoules per output token at capacity (the power model, static power included)lower
TTFT p50median time to first token at the reference loadlower
TTFT p9999th-percentile time to first token at the reference loadlower
TPOT p50median time per output token at the reference loadlower
TPOT p9999th-percentile time per output token at the reference loadlower
ITL p50median inter-token latency at the reference loadlower
ITL p9999th-percentile inter-token latency at the reference loadlower
KV peakpeak KV-cache occupancy at the reference load (lower leaves more headroom)lower

The workloads

Each is a named, parameterised distribution in the simulator, with its source or rationale:

WorkloadPrompt (mean)Output (mean)TurnsSLOs (TTFT, TPOT)
Chat16133831,000 ms, 50.0 ms
Coding agent51216081,500 ms, 40.0 ms
Offline batch2,048256160,000 ms, 500.0 ms
Long-context RAG7,90423015,000 ms, 50.0 ms
Real-time voice64486300.0 ms, 25.0 ms

The devices

DeviceBF16 TFLOP/sHBM TB/sHBM GBFaster formats$/GPU-hour (illustrative)
H100-SXM9893.3580FP8 2×, INT8 2×$3.00
H200-SXM9894.80141FP8 2×, INT8 2×$3.50
B2002,2508.00180FP8 2×, FP4 4×$5.00

Caveats

These are marked wherever the numbers they affect appear (the letter beside a matrix cell, the list under each chart):

PPipeline parallelism is not overlapped

The simulator does not keep several batches in flight across pipeline stages, so PP shows only its cost (bubbles, weight re-reads, an activation hop per stage) and every PP configuration in the sweep loses. Real engines overlap micro-batches across steps.

TTP all-reduce cost is pessimistic at small batch

Tensor-parallel all-reduces use a first-order α–β ring model whose latency term (5 µs a hop, illustrative) dominates at small batch; NVSwitch hardware with in-switch reduction does better. It has not been calibrated against nccl-tests, and communication does not overlap with compute.

MPaged KV shows +0% where memory does not bind

Llama-3-70B on four GPUs per instance never runs out of KV cache at these loads, so paged allocation has nothing to win and the sweep shows no change. Paged KV matters where memory binds: the simulator's results.md section 23 shows it on OPT-13B.

EHuge percentages come from a baseline at its SLO edge

Capacity is the highest load at which 90% of requests meet both SLOs. Where the baseline sits just past a latency cliff (voice and the coding agent especially), a lever that pulls latency back under the SLO multiplies capacity, so gains of hundreds or thousands of percent are real in the model but say more about the cliff than about the lever. Marked wherever capacity at least doubles; read the absolute numbers too.

$Prices per GPU-hour are illustrative

Cost per million tokens uses round illustrative prices (H100 $3.00, H200 $3.50, B200 $5.00 per GPU-hour), not quotes. Cost scales linearly with them, so the ranking of levers on one device does not depend on them; comparisons across devices do.

αSpeculative acceptance α = 0.7 is assumed

Speculative decoding's gain depends on how often the target accepts a drafted token, which depends on the draft, the target and the text. The sweep fixes α at 0.7 as a parameter; it is not measured.

BB200 BF16 rate and power are illustrative

B200's BF16 rate is taken as half its FP8 datasheet rate, and its power coefficients (1,000 W board power, 140 W idle, energy per FLOP and per byte) are illustrative.

Provenance

The vendored files, their commit and their SHA-256 (src/lib/tradeoffs/vendor/VENDORED.json, checked by the unit tests and re-read from the simulator's repository in CI):