Sapiens AI's third-generation Flash model ships in two forms that the model card is careful to separate. The API model is a proprietary reasoning model with a 1M-token context, priced at $0.05 per million input tokens and $0.15 per million output; Artificial Analysis lists it from 11 September 2026 and scores it 36 on Intelligence Index v4.3, the highest of any Sapiens model and one point above the Agnes 2.5 Pro Beta checkpoint. The Agnes-3.0-Flash Preview is an earlier open-weight checkpoint released the same day under Apache 2.0: a 33B dense multimodal model (text, image, and video input) with a 262,144-token context, adjustable reasoning effort, and tool calling. The card states that the Preview is a different checkpoint and configuration from the API model and that the AA result should not be attributed to the open weights.

Provenance: Sapiens does not name a base model. The Preview's config.json reads hidden size 5120, 24 query and 4 KV heads, intermediate size 17,408, a 248,320-token vocabulary, and a 3:1 pattern of delta-attention to global-attention layers, which matches Qwen3.5-27B exactly except for depth (72 layers against 64). Together with the founder's earlier description of post-training from Qwen and DeepSeek checkpoints (see Agnes 2.5 Pro), the entry is flagged as not pretrained from scratch; the shape match is evidence of a depth-extended derivative, not proof.

Model Details

Architecture DENSE
Parameters 33B
Context window 1,000,000
AA Intelligence 36
License Apache 2.0 (Preview weights); API model proprietary

Variants

Name Parameters Notes
Agnes 3.0 Flash (API) — Proprietary reasoning model, 1M context, $0.05/$0.15 per million tokens; AA Intelligence Index v4.3: 36 (listed 11 Sep 2026).
Agnes-3.0-Flash Preview (open weights) 33B Apache 2.0 checkpoint on HuggingFace (11 Sep 2026): 72 layers, hidden 5120, 262,144 context, image and video input; not the checkpoint AA scored.
reasoningmultimodalopen-weightagentic

Related