inference-tradeoffs-explained

Case study

Real-time voice

SLOs: time to first token at most 300.0 ms and time per output token at most 25.0 ms, met by at least 90% of requests at capacity. The workload, as the sweep generates it:

A spoken conversation: short utterances and replies, six turns 3 s apart, a 1,024-token persona prompt. People minimise the silence between turns (Stivers et al., PNAS 2009, doi:10.1073/pnas.0903616106), so the first token gets 300 ms and the stream 25 ms a token to keep speech synthesis fed (our choice).

The frontier

Goodput per GPU against cost per million tokens, starting on this workload (play to compare the others; pick other metrics below).

Loading the Pareto explorer…

Levers that change sign

Levers whose effect on goodput per GPU (H100) points one way here and the other way on at least one other workload: + more goodput than the baseline, − less, 0 within ±2%.

LeverReal-time voiceChatCoding agentOffline batchLong-context RAG
Chunked prefill, 512-token budget++−−+
Chunked prefill, 2,048-token budget++−++
Disaggregated 1P1D (TP4 each)+−−−−
Disaggregated 1P1D, paged decode, prefix-cached prefill+++−+
1 x TP8−−+−−
FP8 KV cache−+++0

Other case studies: Chat · Coding agent · Offline batch · Long-context RAG.