THEORY × COMPUTATION
01 / COMPUTE & COMMUNICATION
Explore the theoretical budget
Projection, not a measured GPU requirement
Theory × experiment · computation × communication
All four panels follow the model, resolution and H200 count above. Measured panels also follow the setting, repetition and P95/P99 controls below. Blue is modeled; green/amber bars retain their evidence status.
THEORY × COMMUNICATION
Traffic, collective count and latency
EXPERIMENT × COMPUTATION / RUNTIME
Observed tails and component costs
EXPERIMENT × COMMUNICATION
Measured communication evidence
Formula, units, and assumptions
02 / REAL RUNS VS THEORY
What the fresh runs establish
Leave-one-process-out calibration
Fit efficiency on two U1 repeats; predict the third at the same resolution. This measures repeat-to-repeat consistency, not independent validation of FLOPs or distributed scaling.
03 / RUNTIME COMPONENTS & SETTINGS
Where the time goes
Each cell is one accepted process repetition. No averaging of P95/P99 quantiles, no mixing resolutions, no substitution of partial results for a three-repeat conclusion.
Component tails are not additive. Instrumented profiles can change timing; compare them only with the same instrumented setting. Cold-start timings have just three process-level observations, so cold-start P95/P99 are insufficient.
Every planned experiment at this resolution
04 / BOTTLENECKS & NEXT STEPS
What to optimize—and what to verify
05 / OUTPUT EXAMPLES
Nine examples, unchanged across model pages
Historical examples — preserved from the previous canvas. These clips are not regenerated fresh-campaign outputs and do not establish quality equivalence for the new timings.
Controls affect all nine videos. Refresh reloads videos only.
Older evidence is superseded
Pre-campaign latency tables, fixed MFU inputs, standalone decoder timings, and previous optimization verdicts are excluded from current calculations. The former report remains available only as a historical archive.
Open historical report ↗