inference-tradeoffs-explained

Case study

Coding agent

SLOs: time to first token at most 1,500 ms and time per output token at most 40.0 ms, met by at least 90% of requests at capacity. The workload, as the sweep generates it:

An agent loop: a 6,144-token prefix (tools, instructions, repository context) shared by every session, eight turns of tool output (mean 512) and short actions (mean 160), 2 s of tool time between turns: long prompts, short outputs, heavy prefix reuse (our choice of numbers).

The frontier

Goodput per GPU against cost per million tokens, starting on this workload (play to compare the others; pick other metrics below).

Loading the Pareto explorer…

Levers that change sign

Levers whose effect on goodput per GPU (H100) points one way here and the other way on at least one other workload: + more goodput than the baseline, − less, 0 within ±2%.

LeverCoding agentChatOffline batchLong-context RAGReal-time voice
Chunked prefill, 512-token budget−+−++
Chunked prefill, 2,048-token budget−++++
Disaggregated 1P1D (TP4 each)−−−−+
Disaggregated 1P1D, paged decode, prefix-cached prefill++−++
1 x TP8+−−−−
FP8 KV cache+++0−

Other case studies: Chat · Offline batch · Long-context RAG · Real-time voice.