Loss-to-Loss Prediction
paper Your tags
Your notes
A scaling-laws result from the Kempner Institute (Brandfonbrener, Vyas, Malach, and Sham Kakade): a shifted power law predicts loss across different pretraining datasets and from pretraining loss to downstream-task loss, extrapolating to roughly 20× the FLOP budget used to fit it. Code is released as KempnerInstitute/loss-to-loss-olmo.