/workloads
Workload case studies
The same levers rank differently for different traffic. Each case study gives a workload's SLOs, the configuration the sweep measured serving it best, its Pareto front, and the levers that help it but hurt another workload (or the reverse).
- Chat
SLOs: TTFT 1,000 ms, TPOT 50.0 ms. Most goodput per GPU: Modern colocated + speculative (MTP) on B200.
- Coding agent
SLOs: TTFT 1,500 ms, TPOT 40.0 ms. Most goodput per GPU: Modern colocated + speculative (MTP) on B200.
- Offline batch
SLOs: TTFT 60,000 ms, TPOT 500.0 ms. Most goodput per GPU: FP4 weights and matmuls (W4A4) on B200.
- Long-context RAG
SLOs: TTFT 5,000 ms, TPOT 50.0 ms. Most goodput per GPU: Chunked 2048 + paged + prefix cache + FP8 (W8A8, KV) on B200.
- Real-time voice
SLOs: TTFT 300.0 ms, TPOT 25.0 ms. Most goodput per GPU: Chunked 2048 + paged + prefix cache + FP8 (W8A8, KV) on B200.