SK Telecom's from-scratch 688B-total / 33B-active MoE (256+1 experts) with Think-Fusion hybrid reasoning, Sparse Gated Attention + MLA, and — notably — native FP8 (MXFP8) pretraining with an FP8 checkpoint. 256K context, Apache 2.0; claims AIME26 97.1 and KMMLU-Pro 80.5. The second Korean frontier open release in the same week as K-EXAONE 2.0; pretraining data cites Nemotron-CC and fineweb-2. Tech report is a GitHub PDF (no arXiv). Companions: A.X K2 ALM (Korean speech in/out on the frozen K2 Light 20B-A2.7B, "h-research" license) and A.X VE (413M from-scratch generative ViT, Apache 2.0) — a full A.X VLM is promised "soon".

Model Details

Architecture MOE
Parameters 688B
Active params 33B
Experts 256
Context window 256,000
AA Intelligence 23 was 27 on v4.2
License Apache 2.0

Paper

frontiermoeopen-weightefficiency

Related