The generation documented in the SEA-LION paper (April 2025, 31 authors led by Raymond Ng with Leslie Teo as senior author): Llama-SEA-LION-v3-8B-IT and Gemma-SEA-LION-v3-9B-IT, continued-pretrained from Llama 3.1 8B Instruct and Gemma 2 9B on a 200B-token mixture the paper tuned in small-scale sweeps (roughly 55% Southeast Asian languages, 25% English, 20% code), then put through multi-stage instruction tuning, alignment, and model merging, with a 70B Llama variant released the same month. The paper claims state-of-the-art results among LLMs supporting the eleven SEA languages and releases the training data recipe, scripts, and checkpoints; a v3.5 reasoning variant (Llama-SEA-LION-v3.5-70B-R) followed in April 2025. Not pretrained from scratch.

Model Details

Architecture DENSE
Parameters 70B

Variants

Name Parameters Notes
Llama-SEA-LION-v3-8B-IT 8B CPT of Llama 3.1 8B Instruct, 200B tokens.
Gemma-SEA-LION-v3-9B-IT 9B CPT of Gemma 2 9B.
Llama-SEA-LION-v3-70B-IT 70B December 2024; v3.5-70B-R reasoning variant April 2025.

Paper

multilingualopen-weightresearch

Related