Pruning and Distilling Mixture-of-Experts into Dense Language Models
paper Your tags
Your notes
First systematic MoE→dense conversion framework: a 350-configuration sweep on Qwen3-30B-A3B with diversity-aware DO-ACP expert scoring, beating dense-to-dense pruning by +6.3pp at 1.6× faster wall-clock. KRAFTON + KAIST (affiliations verified on the PDF); code released.