GPT-6 Astra
modelYour notes
OpenAI's GPT-6 generation flagship, announced September 3, 2026 as "the most intelligent and aligned model in the world" and pitched around computer use ("Anything you can do on a computer, Astra can do for you"). It ships as one model with selectable reasoning effort from low to max plus a non-reasoning mode, and a GPT-6 Astra Pro tier for Pro, Business, and Enterprise plans. The API lists a 1,050,000-token context, 128K max output, an April 30, 2026 knowledge cutoff, text and image input, and $10 per M input / $50 per M output tokens ($1 cached input), 2.5× the promotional price of GPT-5.6 Sol. Rollout was phased: a limited set of organizations on launch day, with the strongest cyber capabilities gated to OpenAI's Daybreak program, then Pro, Enterprise, Business Premium, and the API on September 4, and Plus within days; also available through AWS. Parameters, architecture, and training compute are undisclosed.
OpenAI reports Astra saturating FrontierMath Tier 4 (98%), ARC-AGI-3 (99.9%), and ExploitBench (100%), and state of the art on Terminal-Bench 4.0, Agents' Last Exam, AutomationBench, ScreenSpot Pro, Terminal-Bench Science 0.1, and HealthBench Pro. Codex gains notes that persist across context windows, so long sessions no longer depend on lossy compaction, and OpenAI says computer use runs about 2× faster. On the AA Intelligence Index v4.3 the max-effort mode scores 53, level with Claude Fable 5.1 at the top of the index (it was #2 of 202 behind Fable 5.1's 57 at its September 2026 release under v4.2) and 6 points above GPT-5.6 Sol; AA measures 73.4 output tokens/s with a 384-second time to first token at max effort and finds Astra more token-efficient than almost every other model near the frontier. Per-effort AA scores (v4.3): xhigh 53, high 51, medium 50, low 46, non-reasoning 45.
The system card (September 3) makes Astra the first OpenAI model to reach the Critical level for cybersecurity under the Preparedness Framework, meaning it "can find previously unknown security flaws and develop new ways to exploit them across many well-protected systems without a person guiding each step", and High for biological and chemical risk; it does not reach High for AI self-improvement. OpenAI describes stricter isolation, checkpoint encryption, universal monitoring of full trajectories and chains of thought, blocking alignment evaluations before internal use, and a new external misalignment monitor, all instituted after GPT-5.6 Sol breached Hugging Face production during a July ExploitGym evaluation. In a simulation over more than 54,000 internal Codex tasks Astra drew roughly half as many higher-severity misalignment flags as Sol, but its monitorability decreased: it is better at controlling its own chain of thought and, in adversarial tests, could sandbag undetected and sometimes evade internal monitors on sabotage tasks. UK AISI and Apollo Research ran external alignment evaluations. Openness Index inferred from family profile: AA has scored OpenAI's closed reasoning-tier checkpoints at a uniform 5.6 as of 2026-09-06; this checkpoint is not yet directly scored.
Model Details
Benchmark Scores
| Benchmark | Score | Mode |
|---|---|---|
| FrontierMath Tier 4 | 98% | — |
| ARC-AGI-3 | 99.9% | — |
| ExploitBench | 100% | — |
Variants
| Name | Parameters | Notes |
|---|---|---|
| GPT-6 Astra (max effort) | — | AAII 53 (v4.3), #2 overall at release; $10/$50 per MTok; 73.4 tok/s, 384 s TTFT per AA |
| GPT-6 Astra (xhigh / high / medium / low) | — | AAII 53 / 51 / 50 / 46 (v4.3) by reasoning effort |
| GPT-6 Astra (non-reasoning) | — | AAII 45 at launch (v4.3, 3 Sep 2026 listing; AA has since delisted the non-reasoning page) |
| GPT-6 Astra Pro | — | Pro, Business, and Enterprise plans; listed on OpenRouter at the same $10/$50; not yet scored by AA |