Agnes 3.0 Flash
modelYour notes
Sapiens AI's third-generation Flash model ships in two forms that the model card is careful to separate. The API model is a proprietary reasoning model with a 1M-token context, priced at $0.05 per million input tokens and $0.15 per million output; Artificial Analysis lists it from 11 September 2026 and scores it 36 on Intelligence Index v4.3, the highest of any Sapiens model and one point above the Agnes 2.5 Pro Beta checkpoint. The Agnes-3.0-Flash Preview is an earlier open-weight checkpoint released the same day under Apache 2.0: a 33B dense multimodal model (text, image, and video input) with a 262,144-token context, adjustable reasoning effort, and tool calling. The card states that the Preview is a different checkpoint and configuration from the API model and that the AA result should not be attributed to the open weights.
Provenance: Sapiens does not name a base model. The Preview's config.json reads hidden size 5120, 24 query and 4 KV heads, intermediate size 17,408, a 248,320-token vocabulary, and a 3:1 pattern of delta-attention to global-attention layers, which matches Qwen3.5-27B exactly except for depth (72 layers against 64). Together with the founder's earlier description of post-training from Qwen and DeepSeek checkpoints (see Agnes 2.5 Pro), the entry is flagged as not pretrained from scratch; the shape match is evidence of a depth-extended derivative, not proof.
Model Details
Variants
| Name | Parameters | Notes |
|---|---|---|
| Agnes 3.0 Flash (API) | — | Proprietary reasoning model, 1M context, $0.05/$0.15 per million tokens; AA Intelligence Index v4.3: 36 (listed 11 Sep 2026). |
| Agnes-3.0-Flash Preview (open weights) | 33B | Apache 2.0 checkpoint on HuggingFace (11 Sep 2026): 72 layers, hidden 5120, 262,144 context, image and video input; not the checkpoint AA scored. |