inference-tradeoffs-explained

Chapter 10 · Speculative decoding

Speculative decoding

A cheap draft proposes several tokens and the target model checks them in one pass. Below, the mechanism as the simulator ran it, step by step; then what it measured, why it behaves that way, and the papers.

Loading the speculative decoding animation…

Its row of the matrix

Loading the matrix…

Guess several tokens, check them in one pass

A small draft (here an extra layer of the target that shares its embedding and output head: a multi-token prediction head, DeepSeek-V3) proposes γ tokens; the target model checks all of them in one forward pass and keeps the run it agrees with, plus one token of its own (Leviathan et al.). If each drafted token is accepted with probability α, a pass yields the expected number of tokens in the equation. In the animation, with three drafts per pass and α = 0.7, the two requests kept 2.690 tokens per pass on average, and TPOT p99 fell from 6.4 ms to 3.1 ms.

Why it helps at low batch and hurts at high

A decode pass at small batch is memory-bound: it reads every weight for a handful of tokens, and checking γ + 1 positions per row costs almost the same bytes as checking one (the roofline's left side). At large batch the pass is already compute-bound, the extra positions cost their full FLOPs, and rejected drafts are wasted work:

DeviceContextBatchms/token plainms/token speculativeverify boundSpeed-up
H100512 1 6.126 3.076 memory 1.99x
H100512 32 0.216 0.107 memory 2.02x
H100512 128 0.073 0.053 power 1.37x
H100512 512 0.037 0.048 compute 0.77x
A100512 1 9.743 4.927 memory 1.98x
A100512 128 0.117 0.153 compute 0.76x
Time per output token from the cost model, Llama-3-8B with an MTP draft, three drafts per pass and the same α, recorded in results.md section 25. Below a speed-up of one, speculation loses.

The simulator reproduces Leviathan et al.'s closed forms: the paper's Table 1 speeds to the printed digits (for example 2.53x against 2.53x), and at saturation, where the verify pass is compute-bound, it measures throughput falling to 0.93x of plain decoding.

What the sweep measured

With α fixed at the sweep's parameter (not a measurement), the MTP head cuts TPOT p99 on chat by -51% and raises goodput per GPU on every workload, most on real-time voice (+142%), whose tight TPOT SLO is what limits it, and least on offline batch (+16%), whose large batches leave the least idle compute; there its ITL p99 even rises (+39%). A separate small draft model (Llama-3.2-1B) does about as well, at the cost of its own weights and KV in memory.

What the sweep measured, lever by lever

Speculative, MTP head, gamma 3, alpha 0.7

Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live

Workloadgoodput$/M tokJ/tokTTFT p99TPOT p99ITL p99
Chat+27% α-21% $α-8% α+1% α-51% α-43% α
Coding agent+29% α-23% $α-17% α+6% α-40% α+5% α
Offline batch+16% α-14% $α-2% α+4% α-1% α+39% α
Long-context RAG+39% α-28% $α-15% α+9% α-42% α+2% α
Real-time voice+142% Eα-59% E$α-42% Eα+5% α-45% α+6% α

Speculative, Llama-3.2-1B draft, gamma 4, alpha 0.7

Change from the baseline on H100, at capacity for goodput, cost and energy, at the reference load for latency. Run it live

Workloadgoodput$/M tokJ/tokTTFT p99TPOT p99ITL p99
Chat+25% α-20% $α-5% α+14% α-55% α-49% α
Coding agent+41% α-29% $α-21% α+3% α-39% α+5% α
Offline batch+14% α-12% $α-0% α+9% α-9% α+177% α
Long-context RAG+45% α-31% $α-16% α+2% α-43% α+3% α
Real-time voice+174% Eα-64% E$α-45% Eα+6% α-48% α+5% α

Caveats on these numbers

  • EHuge percentages come from a baseline at its SLO edge. Capacity is the highest load at which 90% of requests meet both SLOs. Where the baseline sits just past a latency cliff (voice and the coding agent especially), a lever that pulls latency back under the SLO multiplies capacity, so gains of hundreds or thousands of percent are real in the model but say more about the cliff than about the lever. Marked wherever capacity at least doubles; read the absolute numbers too.
  • $Prices per GPU-hour are illustrative. Cost per million tokens uses round illustrative prices (H100 $3.00, H200 $3.50, B200 $5.00 per GPU-hour), not quotes. Cost scales linearly with them, so the ranking of levers on one device does not depend on them; comparisons across devices do.
  • αSpeculative acceptance α = 0.7 is assumed. Speculative decoding's gain depends on how often the target accepts a drafted token, which depends on the draft, the target and the text. The sweep fixes α at 0.7 as a parameter; it is not measured.

✓ better than the baseline, ✗ worse, by more than ±2%. Every lever on every metric and device: the matrix; on any two metrics: the explorer.

Papers and sources