Chapter 05
Heterogeneous pools and the optical prefill pool
Different hardware for prefill and decode, including a hypothetical optical transform engine. Below, the mechanism as the simulator ran it, step by step; then what it measured, why it behaves that way, and the papers.
Loading the pool animation…
What the simulator recorded
The sweep does not vary pool hardware; these are the simulator's recorded runs (results.md sections 10 and 11), in the matrix's colours against the all-H100 pools.
| Pools | TTFT p99 | TPOT p99 | Goodput (req/s) | J / token | tok/s per $1000 |
|---|---|---|---|---|---|
| H100 prefill + H100 decode (--device h100) | 353.6 ms | 8.8 ms | 7.69 | 0.326 | 2,850 |
| H100 prefill + A100 decode | 353.6 ms | 16.8 ms | 7.69 | 0.299 | 2,714 |
| A100 prefill + H100 decode | 47.1 s | 7.7 ms | 0.00 | 0.448 | 1,952 |
| 2x A100 prefill + H100 decode | 1,093.5 ms | 8.9 ms | 7.46 | 0.412 | 1,859 |
| Configuration | TTFT p99 | Goodput (req/s) | J / token | tok/s per $1000 |
|---|---|---|---|---|
| 1P1D H100 | 5.4 ms | 7.69 | 0.185 | 2,846 |
| 1P1D H100, GPU FFT at 1/16 | 40.0 ms | 7.69 | 0.210 | 2,846 |
| optical-fft prefill (defaults) + H100 decode | 322.1 s | 0.00 | 0.568 | 689 |
| optical-fft prefill (optimistic) + H100 decode | 90.5 ms | 7.69 | 0.194 | 2,772 |
Put each phase on the device that suits it
Once prefill and decode run on separate pools, they need not run on the same hardware (Splitwise). Decode is memory-bound, so an older or cheaper part with good bandwidth may do; prefill is compute-bound, so it wants FLOPs. In the animation an A100 decode pool behind an H100 prefill pool raises TPOT p99 from 6.8 ms to 10.8 ms, still inside the TPOT SLO, while an A100 prefill pool cannot keep up with the prompts: TTFT p99 goes from 204.6 ms to 1,044 ms.
The optical prefill pool asks the same question of a device that does not exist: a Fourier-optical transform engine co-packaged with an H100-class part, which could only help models whose mixing is an FFT (not attention). With optimistic coefficients its TTFT p99 is 36.0 ms, against 3.1 ms on an H100 that runs the FFT itself; with the default coefficients prompts queue without bound (3,799 ms).
Why the optical pool mostly loses
A transform engine speeds up only the FFT part of a prefill, and in a GPU that part is already fast unless the GPU runs FFTs far below its matmul rate. Each optical pass also pays for conversions between the analogue and digital worlds and for rewriting the optical mask, and a static power for lasers that the GPU does not. The recorded break-even: against a GPU running FFTs at full rate the optical pool never matches its TTFT at any ENOB or mask rate tried (results.md section 13: never), and it matches its energy only below 2.3 W of static power. The Fourier Optics for Inference series (its hub) works through why, and reaches the same mostly negative verdict.
Hardware choice inside one pool is chapter 12; where the FLOPs and bytes of each device come from, the roofline and other ways to build a matrix engine.
Papers and sources
- Patel, Choukse, Zhang et al., Splitwise: Efficient generative LLM inference using phase splitting (arXiv:2311.18677): different hardware for the prompt and token phases
- Zhong, Liu, Chen et al., DistServe: Disaggregating Prefill and Decoding for Goodput-optimized Large Language Model Serving (arXiv:2401.09670): goodput under TTFT and TPOT SLOs, the metric this site measures capacity by