inference-tradeoffs-explained

/learn

The levers

One chapter per lever. Each opens with its row of the matrix, animated across the workloads, and lists what the sweep measured. How each mechanism works is explained step by step on LLM Inference Explained; the chapters here are about the trade-off. The chapters not linked yet draw on the simulator's earlier results and are being written.

  1. Which requests share a forward pass: prompts first, running decodes first, or prompts cut into chunks that ride along with the decodes.

  2. Reserving each request's whole KV cache up front, or handing out blocks as it grows and evicting a request when they run out.

  3. Keeping the KV of shared prompt prefixes so the next request with the same prefix skips their prefill.

  4. Running prefill and decode on separate GPU pools, handing each request's KV cache across a link.

  5. 05Heterogeneous pools and the optical prefill pool (being written)

    Different hardware for prefill and decode, including a hypothetical optical transform engine.

  6. 06KV hand-off links and in-transit compression (being written)

    What the link between the pools costs, and compressing the KV cache on its way across.

  7. 07Encoder-only prefill (CED) (being written)

    A causal encoder–decoder split that changes what the prefill pool has to compute.

  8. Splitting one model across GPUs: by layer slices, by stages, or by experts, and what each costs in communication.

  9. Fewer bytes per weight and per KV value, and faster matmuls where the hardware has units for the format.

  10. A cheap draft proposes several tokens and the target model checks them in one pass.

  11. 11Power and energy (being written)

    Where the joules per token go, and what power caps and clock scaling trade.

  12. 12Hardware choice and cost (being written)

    H100, H200 or B200: what each buys on each workload, per GPU and per dollar.

  13. The levers together, as modern serving engines run them, against each one alone.