"FP4 All the Way: Fully Quantized Training of LLMs" (NeurIPS 2025 Spotlight): the first end-to-end FP4 (NVFP4) training of a 7B LLM that matches BF16, run on 256 Gaudi2 accelerators. It pushes the low-precision-training frontier from FP8 to 4-bit weights, activations, and gradients throughout the forward and backward passes. Joint with Intel, university-led by Daniel Soudry.

Paper

Venue NeurIPS 2025 (Spotlight)
researchefficiencyquantizationtechnique

Related