inference-tradeoffs-explained

/about

About this site

Inference Trade-offs Explained is about the decisions in serving a language model: for each lever, what it improves, what it costs, and under which workload. It is the seventh of a family of companion sites, with the Transformer Decoder Explainer, LLM Inference Explained, LLM Architectures Explained, GPU Kernels Explained, Numerics Explained and Systolic Arrays Explained. The simulator behind it is introduced in the Inference Simulators slide series.

One number source

Every number is read from the trade-off sweep of Disaggregated_Inference_Sim, vendored byte for byte with the simulator's JavaScript engine and its results.md at commit 69ffd2b (scripts/vendor-tradeoffs.ts records each file's SHA-256). Prose numbers are looked up from the data when the page is built, and a unit test fails if a page spells out a number by hand. The method page says how the sweep was run and lists every caveat.

The checks

Built with

Next.js and React, plain SVG and HTML for the charts, the simulator's engine in a Web Worker. Copied from the newest site of the family (Systolic Arrays Explained) and adapted. The source is on GitHub.