inference-tradeoffs-explained

Chapter 13 · Combined

Combining levers

The levers together, as modern serving engines run them, against each one alone. Below, the mechanism as the simulator ran it, step by step; then what it measured, why it behaves that way, and the papers.

Loading the combining animation…

Its row of the matrix

Loading the matrix…

The levers together

Serving engines do not pick one lever: they chunk prefills, page the KV cache, cache prefixes and run narrow formats at once, and some add speculation or split the pools. The sweep measures three such combinations against the same baseline: a modern colocated configuration (chunked prefill with a 2,048-token budget, paged KV, prefix caching, FP8 weights, matmuls and KV), the same levers disaggregated into one prefill and one decode instance, and the colocated one with an MTP speculative head.

Why they multiply, and where they do not

The levers mostly attack different costs (prompt work, memory, bytes per pass, passes per token), so their gains compound: the modern colocated configuration raises goodput per GPU +361% on chat, against +102% for prefix caching and +89% for FP8 alone. Adding speculation helps some workloads and not others: on the coding agent +3827% against +2730% without it, on chat +344% against +361%.

With every lever available to both, disaggregation no longer wins on any workload's goodput: modern disaggregated gains +16% on offline batch against the colocated one's +73%, and less on chat too. It still buys the flattest decode tail (ITL p99 on chat -94% against -89%), so it remains the choice when that tail is the SLO that binds. The very large percentages on voice and the coding agent are real measurements against a baseline that sits at a cliff of their tight SLOs; the matrix marks them.

What the sweep measured, lever by lever

Chunked 2048 + paged + prefix cache + FP8 (W8A8, KV)

Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live

Workloadgoodput$/M tokJ/tokTTFT p99TPOT p99ITL p99
Chat+361% E-78% E$-78% E-75% -60% -89%
Coding agent+2730% E-96% E$-95% E-85% -65% -37%
Offline batch+73% -42% $-48% -42% -65% -49%
Long-context RAG+227% E-69% E$-63% E-47% -56% +330%
Real-time voice+6448% E-98% E$-97% E-87% -57% -36%

Disaggregated 1P1D + paged + prefix cache + FP8

Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live

Workloadgoodput$/M tokJ/tokTTFT p99TPOT p99ITL p99
Chat+260% E-72% E$-74% E-78% -61% -94%
Coding agent+2314% E-96% E$-94% E-90% -64% -35%
Offline batch+16% -14% $-39% +5% -95% -92%
Long-context RAG+142% E-59% E$-58% E-41% -69% -36%
Real-time voice+5395% E-98% E$-96% E-90% -54% -6%

Modern colocated + speculative (MTP)

Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live

Workloadgoodput$/M tokJ/tokTTFT p99TPOT p99ITL p99
Chat+344% Eα-77% E$α-75% Eα-75% α-81% α-94% α
Coding agent+3827% Eα-97% E$α-95% Eα-85% α-83% α-34% α
Offline batch+95% α-49% $α-49% α-35% α-84% α-48% α
Long-context RAG+348% Eα-78% E$α-66% Eα-42% α-77% α-37% α
Real-time voice+7323% Eα-99% E$α-97% Eα-86% α-75% α-33% α

Caveats on these numbers

  • EHuge percentages come from a baseline at its SLO edge. Capacity is the highest load at which 90% of requests meet both SLOs. Where the baseline sits just past a latency cliff (voice and the coding agent especially), a lever that pulls latency back under the SLO multiplies capacity, so gains of hundreds or thousands of percent are real in the model but say more about the cliff than about the lever. Marked wherever capacity at least doubles; read the absolute numbers too.
  • $Prices per GPU-hour are illustrative. Cost per million tokens uses round illustrative prices (H100 $3.00, H200 $3.50, B200 $5.00 per GPU-hour), not quotes. Cost scales linearly with them, so the ranking of levers on one device does not depend on them; comparisons across devices do.
  • αSpeculative acceptance α = 0.7 is assumed. Speculative decoding's gain depends on how often the target accepts a drafted token, which depends on the draft, the target and the text. The sweep fixes α at 0.7 as a parameter; it is not measured.

✓ better than the baseline, ✗ worse, by more than ±2%. Every lever on every metric and device: the matrix; on any two metrics: the explorer.

Papers and sources