Chapter 08 · Parallelism
Tensor, pipeline and expert parallelism
Loading the matrix…
Splitting one model across GPUs: by layer slices, by stages, or by experts, and what each costs in communication.
The chapter text (the mechanism animated, why it behaves as measured, and the papers) is being written. This page already shows what the sweep measured.
1 x TP8
Change from the baseline on H100, at capacity for goodput and cost, at the reference load for latency. Run it live
| Workload | goodput | $/M tok | TTFT p99 | TPOT p99 | ITL p99 |
|---|---|---|---|---|---|
| Chat | -24% T | +32% T$ | -30% T | +19% T | -4% T |
| Coding agent | +6% T | -6% T$ | -45% T | -2% T | -1% T |
| Offline batch | -15% T | +18% T$ | +35% T | -28% T | +101% T |
| Long-context RAG | -15% T | +18% T$ | -32% T | +10% T | +1325% T |
| Real-time voice | -36% T | +55% T$ | -28% T | +5% T | +402% T |
4 x TP2
Change from the baseline on H100, at capacity for goodput and cost, at the reference load for latency. Run it live
| Workload | goodput | $/M tok | TTFT p99 | TPOT p99 | ITL p99 |
|---|---|---|---|---|---|
| Chat | -87% T | +659% T$ | +75867% T | +4% T | -53% T |
| Coding agent | -100% T | – T$ | +526% T | -21% T | +48% T |
| Offline batch | -100% T | – T$ | +675% T | -87% T | -84% T |
| Long-context RAG | -100% T | – T$ | +1011% T | -26% T | +40% T |
| Real-time voice | -100% T | – T$ | +91% T | +40% T | +53% T |
2 x (TP2, PP2), 2 micro-batches
Change from the baseline on H100, at capacity for goodput and cost, at the reference load for latency. Run it live
| Workload | goodput | $/M tok | TTFT p99 | TPOT p99 | ITL p99 |
|---|---|---|---|---|---|
| Chat | -85% PT | +568% PT$ | +69% PT | +152% PT | +136% PT |
| Coding agent | -93% PT | +1364% PT$ | +45% PT | +116% PT | +128% PT |
| Offline batch | -38% PT | +61% PT$ | +41% PT | +44% PT | +45% PT |
| Long-context RAG | -83% PT | +500% PT$ | +58% PT | +131% PT | +2828% PT |
| Real-time voice | -100% PT | – PT$ | +58% PT | +106% PT | +864% PT |
2 x (TP2, PP2), no micro-batching
Change from the baseline on H100, at capacity for goodput and cost, at the reference load for latency. Run it live
| Workload | goodput | $/M tok | TTFT p99 | TPOT p99 | ITL p99 |
|---|---|---|---|---|---|
| Chat | -64% PT | +183% PT$ | +162% PT | +138% PT | +209% PT |
| Coding agent | -100% PT | – PT$ | +121% PT | +122% PT | +59% PT |
| Offline batch | -43% PT | +74% PT$ | +82% PT | +81% PT | +77% PT |
| Long-context RAG | -72% PT | +257% PT$ | +132% PT | +129% PT | +70% PT |
| Real-time voice | -100% PT | – PT$ | +136% PT | +74% PT | +933% PT |
Caveats on these numbers
- PPipeline parallelism is not overlapped. The simulator does not keep several batches in flight across pipeline stages, so PP shows only its cost (bubbles, weight re-reads, an activation hop per stage) and every PP configuration in the sweep loses. Real engines overlap micro-batches across steps.
- TTP all-reduce cost is pessimistic at small batch. Tensor-parallel all-reduces use a first-order α–β ring model whose latency term (5 µs a hop, illustrative) dominates at small batch; NVSwitch hardware with in-switch reduction does better. It has not been calibrated against nccl-tests, and communication does not overlap with compute.
- $Prices per GPU-hour are illustrative. Cost per million tokens uses round illustrative prices (H100 $3.00, H200 $3.50, B200 $5.00 per GPU-hour), not quotes. Cost scales linearly with them, so the ranking of levers on one device does not depend on them; comparisons across devices do.
✓ better than the baseline, ✗ worse, by more than ±2%. Every lever on every metric and device: the matrix; on any two metrics: the explorer.