DataComp-LM is a data-centric benchmark for LLM pretraining: a standardized 240T-token CommonCrawl pool with fixed training recipes, so filtering and curation strategies can be compared head-to-head. Its DCLM-Baseline set became a heavily adopted open pretraining corpus. Note: a large multi-institution collaboration led by UW/Apple/TTIC — TUM appears as a peripheral co-author affiliation, not a central contributor; filed here to record the association. NeurIPS 2024 Datasets & Benchmarks.

Dataset

Size 240T tokens (pool)
datasettraining-dataresearch