Can Scale Save Us From Plasticity Loss in LLMs?
paper Your tags
Your notes
Zyphra study of plasticity loss — the declining ability of networks to learn from new data over the course of training — showing it appears even in stationary pretraining (not just continual learning), and that scaling parameters has diminishing returns as a remedy. Relevant to continual-pretraining and long-run training-dynamics research.