Inkling-Small
modelYour notes
The efficiency companion to Inkling, released as open weights two weeks after the flagship (it was Tinker-only at launch): 276B total / 12B active MoE (42 layers, 6 of 256 routed experts + 2 shared, hybrid local/global attention), natively multimodal (text, image, audio in) with up to 1M-token context. "Comparable performance to Inkling at a quarter of its size" — AA Intelligence Index v4.1: 40 vs the 975B flagship's 41, at $1.20/M output tokens vs $4.05.
Own pretraining run with a revised data mix and recipe (trained on NVIDIA GB300 NVL72 systems), then post-trained from an earlier preview checkpoint partly via on-policy distillation with Inkling as teacher, plus two further weeks of agentic-coding RL — it surpasses the flagship on reasoning and agentic coding (SWE-bench Verified 80.2 vs 77.6) while Inkling keeps the edge on knowledge and factuality. Apache 2.0; NVFP4 checkpoint alongside; fine-tunable on Tinker; day-0 vLLM/SGLang support.
Model Details
Benchmark Scores
| Benchmark | Score | Mode |
|---|---|---|
| SWE-Bench Verified | 80.2 | effort 0.99 |
| Terminal-Bench 2.1 | 64.7 | effort 0.99 |
| HLE | 31.6 | effort 0.99 |
| IFBench | 82.2 | effort 0.99 |