From Reasoning Traces to Reusable Modules
paperYour notes
"Understanding Compositional Generalization in Language Model Reasoning" — a theory of why the standard SFT→RL post-training pipeline produces reasoning that generalizes. The authors (an MBZUAI–CMU collaboration including Eric Xing, Kun Zhang, Ruslan Salakhutdinov, and Zhengzhong Liu; supported by the MBZUAI-WIS Joint Program) propose a hierarchical latent selection model in which SFT supplies the raw module materials in compositional traces, and RL decomposes those traces to identify latent atomic modules and recombine them — proving identifiability conditions for the latent module structure and when identified modules recombine correctly out of distribution. Controlled experiments confirm that RL extracts reusable skills from compound traces, that compositional exposure during RL is essential, and that the optimal recipe has SFT covering all atomic modules while RL targets novel compositions beyond the SFT support — a theoretical grounding for the "SFT elicits, RL lifts" pattern observed empirically across frontier post-training pipelines.