"How to Train Long-Context Language Models (Effectively)" from Danqi Chen's group: a principled data/recipe ablation yielding ProLong-8B-512K, state-of-the-art long-context performance at the 10B scale. The training-side companion to the filed HELMET long-context evaluation.

Model Details

Base model llama-3.1

Paper

open-weightpretrainingresearch

Related