Mercury (Diffusion LLM)
modelYour notes
"Mercury: Ultra-Fast Language Models Based on Diffusion." Introduces diffusion-based LLMs (dLLMs) that forecast multiple tokens simultaneously via iterative denoising, rather than sequential autoregressive generation. Mercury Coder Mini achieves 1,109 tokens/sec on H100.
Mercury 2 (February 2026) adds reasoning capability with AA Intelligence Index v4.3: 14 at ~929 tok/s — roughly 10x faster than comparable autoregressive models at similar quality. A genuinely novel architecture paradigm. By Khanna, Kharbanda, Li, Ermon, Grover, Kuleshov et al.
Provenance: initialization is undisclosed (closed weights, no parameter counts), but the tech report frames Mercury as extending the founders' MDLM from-scratch masked-diffusion pretraining line, "trained on the order of trillions of tokens" of web + proprietary data, and cites no AR-to-diffusion adaptation work — consistent with an in-house pretrain rather than an open-checkpoint conversion.
Model Details
Variants
| Name | Parameters | Notes |
|---|---|---|
| Mercury 2.5 | — | Sep 8 2026: >1,100 tokens/s in production, 260K context (from 128K), $0.20/$0.75 per M tokens ($0.04/$0.15 launch promo); +10 points over Mercury 2, pitched against GPT-5.6 Luna, Gemini 3.5 Flash-Lite, Claude Haiku 4.5; tunable reasoning, native and parallel tool use, JSON mode; API, OpenRouter, Baseten; Mercury Voice (<170 ms TTFT) and Mercury Router previewed; not yet on AA |
| Mercury Coder | — | — |
| Mercury 2 | — | Reasoning, AA Intelligence Index v4.3: 14, 929 tok/s |