Nemotron-SEA-LION-v4.8
modelYour notes
The first mixture-of-experts models in the SEA-LION family, built with NVIDIA on the Nemotron 3 hybrid Mamba-2 architecture and released under MIT with base and post-trained checkpoints plus FP8, NVFP4, and GGUF builds. Nemotron-SEA-LION-v4.8-120B-A12B starts from Nemotron 3 Super (88 layers, 512 routed experts with 22 active plus a shared expert, multi-token prediction) and 30B-A3B from Nemotron 3 Nano (128 experts, 6 active). The 30B received continued pretraining on 150B tokens of Southeast Asian instruction, reasoning, code, and parallel data in NeMo Megatron Bridge; the 120B on 33.5B tokens in NeMo AutoModel; both keep the Nemotron tokenizer and the 1M-token pretraining context, with post-training run at 128K and the cards stating a 262,144-token context. Post-training combines supervised fine-tuning with online on-policy distillation from an FP8 Nemotron 3 Ultra 550B-A55B teacher, followed by merging with Nemotron 3.5 Lightning (30B) or Nemotron 3 Super (120B). The CPT corpus weights Indonesian, Malay, Burmese, Tamil, and Vietnamese instruction data, parallel data across eleven regional languages down to Javanese and Sundanese, and Thai medical reasoning.
On SEA-HELM the 120B rises from Nemotron 3 Super's 49.30 to 63.44 overall and the 30B from Nemotron 3.5 Lightning's 46.06 (Nano's 46.89) to 51.57; the largest moves are in the languages the base models handled worst, Burmese 4.98 to 31.35 and Tamil 21.83 to 56.35 at 120B. The 48-author report (16 September 2026) also documents where the adapted models remain no better than their Nemotron parents, and the SEA-HELM scores are point estimates from the lab's own leaderboard. Not scored on Artificial Analysis.
Model Details
Benchmark Scores
| Benchmark | Score | Mode |
|---|---|---|
| SEA-HELM (overall SEA, 120B-A12B) | 63.44 | — |
| SEA-HELM (overall SEA, 30B-A3B) | 51.57 | — |
Variants
| Name | Parameters | Notes |
|---|---|---|
| Nemotron-SEA-LION-v4.8-120B-A12B | 120B | From Nemotron 3 Super 120B-A12B (LatentMoE hybrid, 88 layers, 512 experts + 1 shared, 22 active); 33.5B-token CPT in NeMo AutoModel; SFT + on-policy distillation from Nemotron 3 Ultra, merged with Super. SEA-HELM 63.44 (from 49.30). |
| Nemotron-SEA-LION-v4.8-30B-A3B | 30B | From Nemotron 3 Nano 30B-A3B (hybrid MoE, 128 experts, 6 active); 150B-token CPT in NeMo Megatron Bridge; merged with Nemotron 3.5 Lightning. SEA-HELM 51.57 (from 46.06). |