inference-tradeoffs-explained

/matrix

Lever × metric matrix

Each row is one lever, each column one metric, and each cell the lever's measured change from the baseline (two colocated instances of four GPUs, prefill-priority batching, reserved KV, BF16) on the same workload and device. Blue is better and vermillion worse, whichever way the metric runs; the arrows say which way it moved and roughly how far. Capacity metrics (goodput, throughput, cost, energy) are at each configuration's own capacity; latencies at the workload's reference load.

Loading the matrix…

Levers that change sign

The same lever can help one workload and hurt another. These are the levers whose effect on goodput per GPU changes sign between workloads on H100 (+ more goodput than the baseline, − less, 0 within ±2%):

LeverChatCoding agentOffline batchLong-context RAGReal-time voice
Chunked prefill, 512-token budget+−−++
Chunked prefill, 2,048-token budget+−+++
Disaggregated 1P1D (TP4 each)−−−−+
Disaggregated 1P1D, paged decode, prefix-cached prefill++−++
1 x TP8−+−−−
FP8 KV cache+++0−