Chapter 02 · KV memory
Paged KV and preemption
Loading the matrix…
Reserving each request's whole KV cache up front, or handing out blocks as it grows and evicting a request when they run out.
The chapter text (the mechanism animated, why it behaves as measured, and the papers) is being written. This page already shows what the sweep measured.
Paged KV, preempt by recompute
Change from the baseline on H100, at capacity for goodput and cost, at the reference load for latency. Run it live
| Workload | goodput | $/M tok | TTFT p99 | TPOT p99 | ITL p99 |
|---|---|---|---|---|---|
| Chat | +0% M | +0% M$ | +0% M | +0% M | +0% M |
| Coding agent | +0% M | +0% M$ | +0% M | +0% M | +0% M |
| Offline batch | +0% M | +0% M$ | -6% | +0% M | -82% |
| Long-context RAG | +0% M | +0% M$ | +0% M | +0% M | +0% M |
| Real-time voice | +0% M | +0% M$ | +0% M | +0% M | +0% M |
Paged KV, preempt by swap to host
Change from the baseline on H100, at capacity for goodput and cost, at the reference load for latency. Run it live
| Workload | goodput | $/M tok | TTFT p99 | TPOT p99 | ITL p99 |
|---|---|---|---|---|---|
| Chat | +0% M | +0% M$ | +0% M | +0% M | +0% M |
| Coding agent | +0% M | +0% M$ | +0% M | +0% M | +0% M |
| Offline batch | +0% M | +0% M$ | -6% | +0% M | -82% |
| Long-context RAG | +0% M | +0% M$ | +0% M | +0% M | +0% M |
| Real-time voice | +0% M | +0% M$ | +0% M | +0% M | +0% M |
Caveats on these numbers
- MPaged KV shows +0% where memory does not bind. Llama-3-70B on four GPUs per instance never runs out of KV cache at these loads, so paged allocation has nothing to win and the sweep shows no change. Paged KV matters where memory binds: the simulator's results.md section 23 shows it on OPT-13B.
- $Prices per GPU-hour are illustrative. Cost per million tokens uses round illustrative prices (H100 $3.00, H200 $3.50, B200 $5.00 per GPU-hour), not quotes. Cost scales linearly with them, so the ranking of levers on one device does not depend on them; comparisons across devices do.
✓ better than the baseline, ✗ worse, by more than ±2%. Every lever on every metric and device: the matrix; on any two metrics: the explorer.