CuatroLLM / TransWebLLM
model Your tags
Your notes
UCL's clean multilingual pretraining line (Stenetorp): CuatroLLM/TransWeb-Edu showed a 1.3B model pretrained on 300B machine-translated tokens matches SOTA multilingual models with ~6% of Llama-3.2's data; the follow-up TransWebEdu/TransWebLLM corpus (EMNLP 2025 main) scales the recipe to 1.7T machine-translated tokens across 9 languages — a practical recipe for non-English pretraining.
Model Details
Architecture DENSE
Parameters 1.3B
Paper
Dataset
Size 1.7T tokens
Format machine-translated web/edu corpus
Languages: 9 languages