Mixtures-of-Experts Overfit More to Repeated Data
paper Your tags
Your notes
With human-written text running out, repeating pretraining data is standard practice, but the evidence for it comes from dense models. This Stanford and University of Washington study varies repetition across single- and multi-domain mixes and across MoE settings, expert count, and granularity, and finds consistently, from 80M to 1B active (8.5B total) parameters, that MoEs degrade faster under repetition, with the effect growing with sparsity as set by total rather than active parameters. A direct input to data-budget decisions for sparse frontier models.