The efficiency companion to Inkling, released as open weights two weeks after the flagship (it was Tinker-only at launch): 276B total / 12B active MoE (42 layers, 6 of 256 routed experts + 2 shared, hybrid local/global attention), natively multimodal (text, image, audio in) with up to 1M-token context. "Comparable performance to Inkling at a quarter of its size" — AA Intelligence Index v4.1: 40 vs the 975B flagship's 41, at $1.20/M output tokens vs $4.05.

Own pretraining run with a revised data mix and recipe (trained on NVIDIA GB300 NVL72 systems), then post-trained from an earlier preview checkpoint partly via on-policy distillation with Inkling as teacher, plus two further weeks of agentic-coding RL — it surpasses the flagship on reasoning and agentic coding (SWE-bench Verified 80.2 vs 77.6) while Inkling keeps the edge on knowledge and factuality. Apache 2.0; NVFP4 checkpoint alongside; fine-tunable on Tinker; day-0 vLLM/SGLang support.

Model Details

Architecture MOE
Parameters 276B
Active params 12B
Experts 256 (top-6)
Context window 1,000,000
AA Intelligence 40
License Apache 2.0

Benchmark Scores

Benchmark Score Mode
SWE-Bench Verified 80.2 effort 0.99
Terminal-Bench 2.1 64.7 effort 0.99
HLE 31.6 effort 0.99
IFBench 82.2 effort 0.99
open-weightmoemultimodalreasoningefficiency

Related