inference-tradeoffs-explained

Case study

Long-context RAG

SLOs: time to first token at most 5,000 ms and time per output token at most 50.0 ms, met by at least 90% of requests at capacity. The workload, as the sweep generates it:

Retrieved documents make the prompt long: lengths fitted to arxiv_summarization's median and 90th percentile (7,059 / 12,985 in, 208 / 371 out; Sarathi-Serve, arXiv:2403.02310, Table 2; lognormal means 7,904 and 230, cv 0.50 and 0.48), behind one of four 1,024-token instruction prefixes. SLOs: 5 s, 50 ms (DistServe's summarisation: 15 s, 0.15 s on A100s).

The frontier

Goodput per GPU against cost per million tokens, starting on this workload (play to compare the others; pick other metrics below).

Loading the Pareto explorer…

Levers that change sign

Levers whose effect on goodput per GPU (H100) points one way here and the other way on at least one other workload: + more goodput than the baseline, − less, 0 within ±2%.

LeverLong-context RAGChatCoding agentOffline batchReal-time voice
Chunked prefill, 512-token budget++−−+
Chunked prefill, 2,048-token budget++−++
Disaggregated 1P1D (TP4 each)−−−−+
Disaggregated 1P1D, paged decode, prefix-cached prefill+++−+
1 x TP8−−+−−

Other case studies: Chat · Coding agent · Offline batch · Real-time voice.