The Ling Team shows that conventional hyperparameter scaling laws break for ultra-sparse mixture-of-experts models: the optimal learning rate and batch size shift with the activation ratio, and neither total nor active parameter count explains the shift alone. The evidence is 1,800 pretraining runs across six activated-parameter scales and models up to 6B non-embedding parameters, about 20 trillion tokens and 200,000 H800 GPU-hour equivalents, from which the paper fits laws that reconcile earlier conflicting findings and make sparsity an explicit input to hyperparameter transfer. Follows the group's July efficient-MoE scaling laws.

Paper

scalingmoetrainingresearch

Related