Switch Distillation (Knowledge Distillation During Mid-Training)
paper Your tags
Your notes
"Knowledge Distillation During Mid-Training Favors Reasoning over Factual Recall." FAIR authors with UW and Princeton (Jacqueline He, Howard Yen, Pang Wei Koh, Luke Zettlemoyer, Wen-tau Yih) show that forward-KL distillation from post-trained teachers behaves differently when applied in mid-training: it lifts reasoning while eroding factual recall. Switch Distillation routes the lowest-entropy fraction of tokens to reverse-KL distillation and trains the rest with cross-entropy. Code released under MIT.