Case study
Chat
SLOs: time to first token at most 1,000 ms and time per output token at most 50.0 ms, met by at least 90% of requests at capacity. The workload, as the sweep generates it:
ShareGPT turn lengths (means 161 in, 338 out, as vLLM's evaluation, arXiv:2309.06180 section 6.1); three turns 10 s apart behind one of eight 512-token system prompts (our choice). SLOs: 1 s to the first token, 50 ms a token (20 tokens/s, faster than reading); DistServe's chatbot SLOs on A100s were 0.25-4 s and 0.1-0.2 s (arXiv:2401.09670, Table 1).
The recommended configuration
Goodput per GPU is capacity under both SLOs, so every configuration here meets them by construction. Same eight GPUs and model for every row; prices illustrative.
| Choice | Configuration | Goodput / GPU | $ / M tokens | J / token | TTFT p99 | TPOT p99 |
|---|---|---|---|---|---|---|
| Most goodput per GPU | Modern colocated + speculative (MTP) on B200 | 12.742 | $0.34 | 0.18 J | 46.9 ms | 4.0 ms |
| Cheapest per M tokens | Modern colocated + speculative (MTP) on B200 | 12.742 | $0.34 | 0.18 J | 46.9 ms | 4.0 ms |
| Best on H100 | Chunked 2048 + paged + prefix cache + FP8 (W8A8, KV) on H100 | 6.543 | $0.40 | 0.26 J | 88.5 ms | 13.4 ms |
| Best on H200 | Chunked 2048 + paged + prefix cache + FP8 (W8A8, KV) on H200 | 6.902 | $0.44 | 0.24 J | 91.9 ms | 11.1 ms |
| Best on B200 | Modern colocated + speculative (MTP) on B200 | 12.742 | $0.34 | 0.18 J | 46.9 ms | 4.0 ms |
| Baseline | 2 x TP4 colocated, prefill-priority, reserved KV, BF16 on H100 | 1.420 | $1.84 | 1.15 J | 355.5 ms | 33.3 ms |
The frontier
Goodput per GPU against cost per million tokens, starting on this workload (play to compare the others; pick other metrics below).
Loading the Pareto explorer…
Levers that change sign
Levers whose effect on goodput per GPU (H100) points one way here and the other way on at least one other workload: + more goodput than the baseline, − less, 0 within ±2%.
| Lever | Chat | Coding agent | Offline batch | Long-context RAG | Real-time voice |
|---|---|---|---|---|---|
| Chunked prefill, 512-token budget | + | − | − | + | + |
| Chunked prefill, 2,048-token budget | + | − | + | + | + |
| Disaggregated 1P1D (TP4 each) | − | − | − | − | + |
| Disaggregated 1P1D, paged decode, prefix-cached prefill | + | + | − | + | + |
| 1 x TP8 | − | + | − | − | − |
| FP8 KV cache | + | + | + | 0 | − |
Other case studies: Coding agent · Offline batch · Long-context RAG · Real-time voice.