inference-tradeoffs-explained

Chapter 05

Heterogeneous pools and the optical prefill pool

Different hardware for prefill and decode, including a hypothetical optical transform engine. Below, the mechanism as the simulator ran it, step by step; then what it measured, why it behaves that way, and the papers.

Loading the pool animation…

What the simulator recorded

The sweep does not vary pool hardware; these are the simulator's recorded runs (results.md sections 10 and 11), in the matrix's colours against the all-H100 pools.

PoolsTTFT p99TPOT p99Goodput (req/s)J / tokentok/s per $1000
H100 prefill + H100 decode (--device h100)353.6 ms 8.8 ms 7.69 0.326 2,850
H100 prefill + A100 decode353.6 ms 16.8 ms 7.69 0.299 2,714
A100 prefill + H100 decode47.1 s 7.7 ms 0.00 0.448 1,952
2x A100 prefill + H100 decode1,093.5 ms 8.9 ms 7.46 0.412 1,859
Heterogeneous pools, Llama-3-8B shape, eight requests per second. Silicon dollars are illustrative (die area and yield, not prices).
ConfigurationTTFT p99Goodput (req/s)J / tokentok/s per $1000
1P1D H1005.4 ms 7.69 0.185 2,846
1P1D H100, GPU FFT at 1/1640.0 ms 7.69 0.210 2,846
optical-fft prefill (defaults) + H100 decode322.1 s 0.00 0.568 689
optical-fft prefill (optimistic) + H100 decode90.5 ms 7.69 0.194 2,772
An FFT-mixing model (Hyena-2 with circulant mixing, last-token head: the most transform-heavy variant) with a hypothetical optical prefill pool. Every optical coefficient is illustrative.

Put each phase on the device that suits it

Once prefill and decode run on separate pools, they need not run on the same hardware (Splitwise). Decode is memory-bound, so an older or cheaper part with good bandwidth may do; prefill is compute-bound, so it wants FLOPs. In the animation an A100 decode pool behind an H100 prefill pool raises TPOT p99 from 6.8 ms to 10.8 ms, still inside the TPOT SLO, while an A100 prefill pool cannot keep up with the prompts: TTFT p99 goes from 204.6 ms to 1,044 ms.

The optical prefill pool asks the same question of a device that does not exist: a Fourier-optical transform engine co-packaged with an H100-class part, which could only help models whose mixing is an FFT (not attention). With optimistic coefficients its TTFT p99 is 36.0 ms, against 3.1 ms on an H100 that runs the FFT itself; with the default coefficients prompts queue without bound (3,799 ms).

Why the optical pool mostly loses

A transform engine speeds up only the FFT part of a prefill, and in a GPU that part is already fast unless the GPU runs FFTs far below its matmul rate. Each optical pass also pays for conversions between the analogue and digital worlds and for rewriting the optical mask, and a static power for lasers that the GPU does not. The recorded break-even: against a GPU running FFTs at full rate the optical pool never matches its TTFT at any ENOB or mask rate tried (results.md section 13: never), and it matches its energy only below 2.3 W of static power. The Fourier Optics for Inference series (its hub) works through why, and reaches the same mostly negative verdict.

Hardware choice inside one pool is chapter 12; where the FLOPs and bytes of each device come from, the roofline and other ways to build a matrix engine.

Papers and sources