AIHardwareResearch
NVIDIA GB300 NVL72 Sets World Record for MoE Pre-Training with DeepSeek-V3 671B
NVIDIA's GB300 NVL72 achieved a world record in pre-training the DeepSeek-V3 671B model using mixture of experts (MoE), reaching 1,648 TFLOPs per GPU and advancing large-scale AI training efficiency.
Frontier model pre-training has converged on mixture of experts (MoE), fundamentally changing the limits of large-scale AI training. As compute per token falls, communication increasingly determines how efficiently models scale across thousands of GPUs. NVIDIA GB300 NVL72 set a world record for pre-training DeepSeek-V3 671B at 1,648 TFLOPs per GPU, showing how advances across the entire AI hardware and software stack enable improved performance and scalability for massive AI models.