inference-tradeoffs-explained

Chapter 08 · Parallelism

Tensor, pipeline and expert parallelism

Loading the matrix…

Splitting one model across GPUs: by layer slices, by stages, or by experts, and what each costs in communication.

The chapter text (the mechanism animated, why it behaves as measured, and the papers) is being written. This page already shows what the sweep measured.

1 x TP8

Change from the baseline on H100, at capacity for goodput and cost, at the reference load for latency. Run it live

Workloadgoodput$/M tokTTFT p99TPOT p99ITL p99
Chat-24% T+32% T$-30% T+19% T-4% T
Coding agent+6% T-6% T$-45% T-2% T-1% T
Offline batch-15% T+18% T$+35% T-28% T+101% T
Long-context RAG-15% T+18% T$-32% T+10% T+1325% T
Real-time voice-36% T+55% T$-28% T+5% T+402% T

4 x TP2

Change from the baseline on H100, at capacity for goodput and cost, at the reference load for latency. Run it live

Workloadgoodput$/M tokTTFT p99TPOT p99ITL p99
Chat-87% T+659% T$+75867% T+4% T-53% T
Coding agent-100% T– T$+526% T-21% T+48% T
Offline batch-100% T– T$+675% T-87% T-84% T
Long-context RAG-100% T– T$+1011% T-26% T+40% T
Real-time voice-100% T– T$+91% T+40% T+53% T

2 x (TP2, PP2), 2 micro-batches

Change from the baseline on H100, at capacity for goodput and cost, at the reference load for latency. Run it live

Workloadgoodput$/M tokTTFT p99TPOT p99ITL p99
Chat-85% PT+568% PT$+69% PT+152% PT+136% PT
Coding agent-93% PT+1364% PT$+45% PT+116% PT+128% PT
Offline batch-38% PT+61% PT$+41% PT+44% PT+45% PT
Long-context RAG-83% PT+500% PT$+58% PT+131% PT+2828% PT
Real-time voice-100% PT– PT$+58% PT+106% PT+864% PT

2 x (TP2, PP2), no micro-batching

Change from the baseline on H100, at capacity for goodput and cost, at the reference load for latency. Run it live

Workloadgoodput$/M tokTTFT p99TPOT p99ITL p99
Chat-64% PT+183% PT$+162% PT+138% PT+209% PT
Coding agent-100% PT– PT$+121% PT+122% PT+59% PT
Offline batch-43% PT+74% PT$+82% PT+81% PT+77% PT
Long-context RAG-72% PT+257% PT$+132% PT+129% PT+70% PT
Real-time voice-100% PT– PT$+136% PT+74% PT+933% PT

Caveats on these numbers

✓ better than the baseline, ✗ worse, by more than ±2%. Every lever on every metric and device: the matrix; on any two metrics: the explorer.