Inference Trade-offs Explained
Which lever helps which metric, for which workload
Serving a language model is a stack of decisions: how to batch, how to hold the KV cache, whether to split prefill from decode, how to split the model across GPUs, which number formats, whether to draft tokens speculatively. Each lever helps some measures and costs others, and the answer changes with the workload. This site is organised around those decisions.
Nothing here is asserted. Every number comes from a recorded sweep of a discrete-event simulator: 23 levers and the baseline, 5 workloads and 3 devices, 350 configurations of the same 8 GPUs serving Llama-3-70B. For example, chunked prefill with a small budget changes goodput per GPU on H100 by +28% for chat but by -24% for a coding agent.
The simulator is Disaggregated_Inference_Sim. Its JavaScript engine runs in your browser for the what-if, and reproduces the sweep's recorded latencies exactly. How the sweep was run, and what is illustrative: the method.
Loading the Pareto explorer…
Pareto explorer
Pick two metrics and a workload: every configuration as a point, the front drawn, and watch it move as the workload or the SLO changes.
Lever × metric matrix
Every lever against every metric, better or worse than the baseline and by how much, for each workload: see which cells flip.
Live what-if
Start from a measured configuration, toggle levers and re-run the simulator in your browser on the very requests the sweep used.
One page per lever, with its row of the matrix animated across the workloads: the levers.
Part of a family of companion sites: the Transformer Decoder Explainer (one forward pass), LLM Inference Explained (how serving works), LLM Architectures Explained (how the models differ), GPU Kernels Explained (how a GPU runs the maths), Numerics Explained (the number formats, and what they do to accuracy) and Systolic Arrays Explained (the silicon underneath). This one is about the decisions. How it was built: about.