Meta Superintelligence Labs' successor to Muse Spark 1.2, billed by Mark Zuckerberg as delivering "frontier performance at low cost" and by Meta as its largest improvement in coding and agentic work to date. "Trained for agentic workflows and optimized for competitive coding performance," it is tuned for long-horizon coding: tracking context and prior results, working through messy or conflicting inputs, and asking for input when needed — while using ~20% fewer tool calls and ~25% fewer tokens than 1.2. Meta also cites improved instruction-following and constraint preservation, adversarial robustness and prompt-injection resistance, and better calibration on irreversible actions. Natively multimodal over video, images, and documents, with visual reasoning run "through a real execution environment instead of scripted steps."

Meta's launch scorecard (against Muse Spark 1.2, GPT-5.6 Sol (max), and Opus 5 (max)) leads its comparison group on coding and long context: Terminal-Bench 2.1 88.8 (up from 82.9), DeepSWE v1.1 75.4 (up from 59.3), SWEAtlas CodeBase QnA 59.4, and MRCR 98.5 / 98.1 (256K–512K / 512K–1M); agentic results include GDPval-AA v2 1754 Elo (vs Opus 5's 1824), OSWorld 2.0 66.9, DeepSearchQA 89.4, and Agentic IF Index 57.8. Artificial Analysis scored it the same day (53 on v4.2); on AA Intelligence Index v4.3 it reads 48 in max reasoning mode (45 at xhigh) — Meta's highest score to date, up from 1.2's 40, third among labs behind Anthropic and OpenAI.

Proprietary (parameters undisclosed), with a 1M-token context, live in Muse Code and the Meta Model API at launch ($1.25 / $4.25 per 1M input/output tokens, feedback-gated Contributor Tier from $0.10 input) plus OpenRouter, and rolling out to Meta AI across Instagram and Facebook. Zuckerberg says Meta plans to release open-weights versions of the Muse line in the near future.

Model Details

Context window 1,000,000
AA Intelligence 48 was 53 on v4.2
License Proprietary

Benchmark Scores

Benchmark Score Mode
Terminal-Bench 2.1 88.8 —
DeepSWE v1.1 75.4 —
SWEAtlas CodeBase QnA 59.4 —
MRCR 512K-1M 98.1 —
GDPval-AA v2 1754 —
OSWorld 2.0 66.9 —
frontiercodingagentsmultimodalreasoningproprietary

Related