🛰️ Daily AI Frontier
‹ back to 2026-09-03

UE5M3 FP4 Block Scaling for Stable Language Model Pretraining

Research Efficiency & Systems

Ranking

Overall 81
Content 100
Popularity 37

Observed public metrics from 1 member.

Representative image for UE5M3 FP4 Block Scaling for Stable Language Model Pretraining

Merged summary

TL;DR - This paper presents an FP4 language-model pretraining recipe that pairs E2M1 values with wide-range UE5M3 block scales, simplifying quantization while improving reported loss and benchmark aggregates. It demonstrates the approach on an 8B Nemotron-H model trained for nearly 190 billion tokens.

  • UE5M3 block scales enable periodic tensor scaling without randomized Hadamard transforms.
  • The recipe selectively applies stochastic rounding to backward gradients and uses FP4 for all eligible internal linear layers.
  • A block-16 configuration achieved lower final-window training loss and quantized-inference validation loss than NVIDIA Transformer Engine’s NVFP4 recipe.
  • Removing RHT and the BF16 final-block exemption increased measured model-body token throughput by 21.2% in a native NVFP4 execution ablation.

Sources (1)

UE5M3 FP4 Block Scaling for Stable Language Model Pretraining

arXiv cs.LG Robert Hu, Carlo Luschi, Paul Balanca 2026-09-02 arXiv:2609.02846
Public signals Semantic Scholar citations 0 · Semantic Scholar influential citations 0
Providers: Hugging Face · N/A OpenAlex · N/A Publisher · N/A Semantic Scholar · Citations 0 · Influential citations 0 X · N/A Fetched 2026-09-15 14:22:40.003585 UTC

TL;DR - This paper presents an FP4 language-model pretraining recipe that pairs E2M1 values with wide-range UE5M3 block scales, simplifying quantization while improving reported loss and benchmark aggregates. It demonstrates the approach on an 8B Nemotron-H model trained for nearly 190 billion tokens.

  • UE5M3 block scales enable periodic tensor scaling without randomized Hadamard transforms.
  • The recipe selectively applies stochastic rounding to backward gradients and uses FP4 for all eligible internal linear layers.
  • A block-16 configuration achieved lower final-window training loss and quantized-inference validation loss than NVIDIA Transformer Engine’s NVFP4 recipe.
  • Removing RHT and the BF16 final-block exemption increased measured model-body token throughput by 21.2% in a native NVFP4 execution ablation.
item →