Thinking Machines' first model, trained from scratch: a natively multimodal MoE Transformer — 975B total / 41B active parameters, 256 routed experts per MoE layer (6 active plus shared experts), sliding-window and global attention interleaved 5:1 — pretrained on 45T tokens of text, images, audio, and video with no separate encoders (images as 40×40 pixel patches, audio as dMel spectrograms). 1M context. Post-training: SFT bootstrapped from open-weights models (including Kimi K2.5) followed by 30M+ RL rollouts, with a controllable thinking-effort dial that reaches peer scores at roughly a third of the tokens (AA measures ~25K output tokens per Intelligence Index task vs 37–43K for peers). AA Intelligence Index 25 (v4.3, measured at xhigh effort) — the leading U.S. open-weights model at release, ahead of Nemotron 3 Ultra (23). Apache 2.0; fine-tunable on Tinker day one; served by Together, Fireworks, Modal, Databricks, and Baseten. The companion Inkling-Small (276B/12B, AA 40) got its own open-weights release on July 30. Announcement-post only so far — no technical report yet.

Model Details

Architecture MOE
Parameters 975B
Active params 41B
Experts 256 (top-6)
Context window 1,000,000
Training tokens 45T
AA Intelligence 25 was 26 on v4.3
License Apache 2.0

Benchmark Scores

Benchmark Score Mode
HLE (with tools) 46.0 effort 0.99
AIME 97.1 effort 0.99
GPQA Diamond 87.2 effort 0.99
SWE-Bench 77.6 effort 0.99
Terminal-Bench 63.8 effort 0.99

Variants

Name Parameters Notes
Inkling 975B —
Inkling-Small 276B Open weights July 30, 2026 — tracked as its own output
Inkling-NVFP4 — NVFP4-quantized checkpoint on HuggingFace
frontieropen-weightmoemultimodalreasoning

Related