SemiAnalysis finds H100 still dominates frontier-scale training, GB200 NVL72 not ready

The August 2025 SemiAnalysis report states that, at that time, no front-end laboratory or major cloud provider succeeded in executing a mega-scale training on GB200 NVL72. H100, H200 and Google’s TPU remained the only hardware that actually closed frontier-scale training runs. Blackwell may appear impressive on paper, but the software stack has not matured and reliability leads to downtime that erodes the performance-per-dollar advantage.
The report is based on benchmark runs across more than 2,000 H100 processors, with analysis of model FLOPs utilization (MFU), total cost of ownership (TCO) and cost to train one million tokens. Data are spread across cluster sizes from 128 to 2,048 processors and across different Nvidia software versions, allowing observation of how stack maturation improves efficiency over time. An energy metric—joules per token—is added, calibrated to the average annual electricity consumption of a U.S. household, to give social context to power cost.
In the second part, GB200 NVL72 benchmarks on Llama4 400B MoE and DeepSeek 670B MoE are presented alongside earlier H100 results. Raw performance appears promising, but the report incorporates downtime from poor reliability and lost engineering time into perf-per-TCO. The result is that GB200’s performance-per-dollar advantage disappears when operational reality is considered. Backplane failures and general instability are identified as the primary causes.
According to SemiAnalysis, the ramp-up of GB200 NVL72 is slightly slower than its predecessors, but not by a large margin. The report estimated that by the end of 2025 the software would improve significantly, and that, combined with model architectures designed for larger scale-up, notable efficiency gains are expected. On the reliability side, challenges remain substantial and will require tighter collaboration between Nvidia and its partners, but the ecosystem is expected to allocate resources for a rapid fix.
The report concludes with a hiring notice from SemiAnalysis itself: a new-graduate engineer for the engineering team, with emphasis on large-scale benchmarks across multiple vendors (AMD, NVIDIA, TPU, Trainium), reproducible CI/CD development and system reliability. Requirements include strong Python skills, an SRE background and modern DevOps expertise, reflecting the infrastructure level needed to generate the data presented in the report.