Train Models Faster with JAX and MaxText Using NVFP4 on NVIDIA Blackwell
NVIDIA introduces NVFP4 precision support in JAX and MaxText to accelerate large language model pretraining on Blackwell GPUs, improving throughput and reducing training time and costs.
Pre-training frontier large language models (LLMs) comes down to throughput. When training spans trillions of tokens across thousands of accelerators, every percentage point of step time can add up to days of training and substantial compute costs. Numerical precision is one of the highest-leverage knobs available, but low-bit mixed-precision pretraining is hard to get right. To address this, NVIDIA introduces support for NVFP4 precision in JAX and MaxText frameworks running on NVIDIA Blackwell GPUs. This enables faster training of large language models by improving throughput and efficiency, reducing overall training time and compute expenses.