/what-if
Live what-if
Choose a workload, a device and a starting configuration from the sweep, then toggle levers. Your browser runs the simulator's own engine on both configurations, on the very requests the sweep generated for that workload at its reference load (a few thousand for the multi-turn workloads), in a Web Worker. When a configuration is one of the sweep's, its live latencies are identical to the recorded ones, to the last bit: the engine is the same code, bit-exact with the simulator's Python package.
The live run measures one load. Capacity (the highest load meeting the SLOs, which sets goodput, cost and energy per token in the explorer and the matrix) takes a search of several runs, so for those the page quotes the sweep's record. Speculative decoding assumes an acceptance rate of 0.7.
Loading the live simulator…