Major generation leap with Opus and Sonnet variants. Claude 4 Sonnet introduced 1M token context. Both models achieved record SWE-bench Verified scores: Opus 72.5%, Sonnet 72.7%. Architecture undisclosed.

Claude 4 was widely regarded as the best coding model at launch, with particular strength in agentic tasks, extended reasoning, and instruction following. Launched alongside Claude Code GA and the Claude Agent SDK. AA Intelligence Index v4.3: 21 (Opus with extended thinking; 17 without), 19 (Sonnet with extended thinking; 17 without). The score here follows AA's extended-thinking Opus page since 2026-09-08; earlier readings in the score trail came from the non-reasoning Opus page. Proprietary.

Model Details

Context window 1,000,000
AA Intelligence 21 was 17 on v4.3

Variants

Name Parameters Notes
Claude 4 Opus ~ 1.4T 200K context; AA Intelligence Index v4.3: 21 (extended thinking)
Claude 4 Sonnet — 1M context; AA Intelligence Index v4.3: 19 (extended thinking)
frontiercodingreasoningagents

Related