IO- and tile-aware kernels for Mixture-of-Experts training and inference (ICLR 2026; 731★), delivering ~1.86× throughput by rethinking the memory movement in MoE layers — systems work aimed squarely at the architecture now dominant across frontier models.

Paper

infrastructureefficiencymoe