Chapter 03 · Prefix caching
Prefix caching
Loading the matrix…
Keeping the KV of shared prompt prefixes so the next request with the same prefix skips their prefill.
The chapter text (the mechanism animated, why it behaves as measured, and the papers) is being written. This page already shows what the sweep measured.
Prefix caching (paged KV)
Change from the baseline on H100, at capacity for goodput and cost, at the reference load for latency. Run it live
| Workload | goodput | $/M tok | TTFT p99 | TPOT p99 | ITL p99 |
|---|---|---|---|---|---|
| Chat | +102% E | -50% E$ | -57% | -27% | -63% |
| Coding agent | +868% E | -90% E$ | -74% | -42% | -1% |
| Offline batch | +0% | +0% $ | -6% | +0% | -82% |
| Long-context RAG | +9% | -8% $ | -5% | -9% | +1% |
| Real-time voice | +656% E | -87% E$ | -77% | -29% | -0% |
Caveats on these numbers
- EHuge percentages come from a baseline at its SLO edge. Capacity is the highest load at which 90% of requests meet both SLOs. Where the baseline sits just past a latency cliff (voice and the coding agent especially), a lever that pulls latency back under the SLO multiplies capacity, so gains of hundreds or thousands of percent are real in the model but say more about the cliff than about the lever. Marked wherever capacity at least doubles; read the absolute numbers too.
- $Prices per GPU-hour are illustrative. Cost per million tokens uses round illustrative prices (H100 $3.00, H200 $3.50, B200 $5.00 per GPU-hour), not quotes. Cost scales linearly with them, so the ranking of levers on one device does not depend on them; comparisons across devices do.
✓ better than the baseline, ✗ worse, by more than ±2%. Every lever on every metric and device: the matrix; on any two metrics: the explorer.