01 · THEORY
Compute, communication, and H200 sizing
The explorer starts at 13 fps—the first whole-frame target above 12. Switch the shared matched resolution, target, real-run efficiency calibration, parallel strategy, and GPU count. The model recomputes token geometry, FLOPs, collective traffic, block latency, and the minimum power-of-two H200 count. Values in this section are projections; measured values appear in the next section.
Selected configuration: compute vs communication
Per-forward FLOP composition
Scaling and communication by H200 count
02 · EVIDENCE
Real runs compared with theory
480×832 vs 704×1280 at one timing boundary
03 · RUNTIME
Where real run time goes
Matched-resolution block boundary
Observed time by component and setting
04 · ACTION
Bottlenecks and required optimization
05 · OUTPUTS
Shared real-run output gallery
The gallery is deliberately independent of the selected model page. Every page keeps all three scenes and all three model columns visible for a like-for-like qualitative comparison.