Google's third Flash release in six weeks and its most intelligent workhorse model for coding and agents, built on Gemini 3.7 Flash and shipped September 2, 2026 alongside the restricted Gemini 3.8 Flash Cyber. The change Google describes is diligence rather than architecture: the model "works harder" on complex tasks, executing extra reasoning steps and calling tools iteratively, at the cost of more output tokens; low, medium, and high thinking levels let developers trade that overhead against cost and latency. Natively multimodal (text, images, video, audio, PDFs in; text out) with a 1,048,576-token context window and 65,536 output tokens, function calling, code execution, search grounding, and computer use (preview). Knowledge cutoff March 2026 per the model card.

Google's evaluation sheet (September 2026, high thinking, pass@1) reports 73.7% on DeepSWE v1.1 with a mini-swe-agent harness (3.7 Flash 65.3%, Claude Opus 5 74.0%, GPT-5.6 Sol 72.7%), 61.4% on Vals Finance Agent V2 (3.7 Flash 59.0%, Opus 5 58.6%), 10.0% on Harvey's Legal Agent Benchmark (3.7 Flash 8.8%, Opus 5 6.7%), 54.9% on HLE-Verified (GPT-5.6 Sol 54.5%, Opus 5 54.4%), Terminal-Bench 2.1 89.4%, CharXiv Reasoning 86.2%, LVBench 87.8%, LABBench2 86.2%, and a GDPval-AA v2 Elo of 1545 (Opus 5 1824). The gaps to the frontier are on the general-agent and computer-use rows: Terminal-Bench 4.0 19.1% (Opus 5 51.8%, Sol 37.3%) and OSWorld 2.0 59.0% (Opus 5 75.4%). Google also cites more than three times as many completed tasks as 3.7 Flash on long-running document-heavy workflows. Launch coverage put it at 59 on the then-current AA Intelligence Index v4.1.1, level with GPT-5.6 Sol and Grok 4.6; on AA Intelligence Index v4.3 it scores 41 at high effort (40 medium, 33 low), #32 of 202 at last check and Google's highest score on the index, and the fastest model AA measures at 328.8 output tokens/s. Pricing holds at $0.75 per million input tokens and $3.75 per million output tokens through December 31, 2026, doubling to $1.50/$7.50 on January 1, 2027. Available in the Gemini API, AI Studio, Antigravity, Android Studio, Stitch, Gemini Enterprise, and the Gemini app for Pro and Ultra subscribers; proprietary, with architecture, parameter count, and training scale undisclosed.

Model Details

Context window 1,048,576
AA Intelligence 41
License Proprietary
Base model gemini-3.7-flash

Benchmark Scores

Benchmark Score Mode
DeepSWE v1.1 73.7% high thinking, mini-swe-agent harness
GDPval-AA v2 1545 Elo —
Vals Finance Agent V2 61.4% —
Harvey Legal Agent Benchmark 10.0% all-pass rate
Terminal-Bench 2.1 89.4% Terminus 2 harness
Terminal-Bench 4.0 19.1% —
GDP.PDF 35.0% all-pass rate
CharXiv Reasoning 86.2% no tools
LVBench 87.8% agentic (87.1% static)
HLE-Verified 54.9% —
OSWorld 2.0 59.0% partial score, batched tool calls
BioMysteryBench 88.8% human-solvable (56.5% human-difficult)
LABBench2 86.2% —

Variants

Name Parameters Notes
Gemini 3.8 Flash (high) — Default AA anchor; AA Intelligence Index v4.3: 41
Gemini 3.8 Flash (medium / low) — AA Intelligence Index v4.3: 40 / 33
frontiermultimodalreasoningcodingagentic

Related