inference-tradeoffs-explained

Inference Trade-offs Explained

Which lever helps which metric, for which workload

Serving a language model is a stack of decisions: how to batch, how to hold the KV cache, whether to split prefill from decode, how to split the model across GPUs, which number formats, whether to draft tokens speculatively. Each lever helps some measures and costs others, and the answer changes with the workload. This site is organised around those decisions.

Nothing here is asserted. Every number comes from a recorded sweep of a discrete-event simulator: 23 levers and the baseline, 5 workloads and 3 devices, 350 configurations of the same 8 GPUs serving Llama-3-70B. For example, chunked prefill with a small budget changes goodput per GPU on H100 by +28% for chat but by -24% for a coding agent.

The simulator is Disaggregated_Inference_Sim. Its JavaScript engine runs in your browser for the what-if, and reproduces the sweep's recorded latencies exactly. How the sweep was run, and what is illustrative: the method.

Loading the Pareto explorer…

One page per lever, with its row of the matrix animated across the workloads: the levers.

Part of a family of companion sites: the Transformer Decoder Explainer (one forward pass), LLM Inference Explained (how serving works), LLM Architectures Explained (how the models differ), GPU Kernels Explained (how a GPU runs the maths), Numerics Explained (the number formats, and what they do to accuracy) and Systolic Arrays Explained (the silicon underneath). This one is about the decisions. How it was built: about.