SK Telecom's from-scratch 688B-total / 33B-active MoE (256+1 experts) with Think-Fusion hybrid reasoning, Sparse Gated Attention + MLA, and — notably — native FP8 (MXFP8) pretraining with an FP8 checkpoint. 256K context, Apache 2.0; claims AIME26 97.1 and KMMLU-Pro 80.5. The second Korean frontier open release in the same week as K-EXAONE 2.0; pretraining data cites Nemotron-CC and fineweb-2. Tech report is a GitHub PDF (no arXiv). Companions: A.X K2 ALM (Korean speech in/out on the frozen K2 Light 20B-A2.7B, "h-research" license) and A.X VE (413M from-scratch generative ViT, Apache 2.0) — a full A.X VLM is promised "soon".

Model Details

Architecture MOE
Parameters 688B
Active params 33B
Experts 256
Context window 256,000
License Apache 2.0
frontiermoeopen-weightefficiency

Related