SMELT (Scaling Laws for Compute-Matched MoE Looped Transformers)
paper Your tags
Your notes
Tsinghua with ByteDance Seed, M-A-P, and TokenWave.AI loops the middle half of a MoE transformer's layers twice at matched FLOPs, parameters, and KV cache, scales the design to 54B non-embedding parameters, and fits per-architecture Chinchilla-style laws: 6.8 to 18.0% FLOP savings on the compute-optimal frontier, largest on code.