OpenAI released GPT-6 Astra on September 3, 2026, describing it in its API changelog as “our most capable model, built for the hardest end-to-end work.” In its announcement, OpenAI also says the model “meets the Critical threshold in cybersecurity under our Preparedness Framework.”
On this page
Availability
OpenAI says GPT-6 Astra was “rolling out today to a limited set of organizations and over the coming days will become available to all ChatGPT Plus, Pro, Business, and Enterprise users, as well as through the OpenAI API, Microsoft Azure, and AWS Bedrock.” The API model ID is gpt-6-astra, and the model is also available in Codex.
The changelog notes that GPT-6 Astra “does not support the none reasoning effort level,” nor custom temperature or top_p values. On the same day, OpenAI added three Responses API controls for long-running work: asynchronous tool calling, mid-turn steering over WebSockets, and changing reasoning effort mid-conversation.
Notes across context windows in Codex
The main workflow change is in Codex. OpenAI says: “In Codex, Astra can keep notes across context windows, preserving accumulated details without repeatedly compressing them into a single summary.” For long coding sessions, that replaces a single rolling summary with notes the model carries from one context window to the next.
Cybersecurity rating
The Critical rating under OpenAI’s Preparedness Framework applies to cybersecurity. On ExploitBench, OpenAI reports a score of 100.0% for GPT-6 Astra.
Benchmark results
The scores below were published by OpenAI on the Astra page.
| Benchmark | GPT-6 Astra | Note |
|---|---|---|
| FrontierMath Tier 4 (v2) | 97.6% | OpenAI’s text rounds this to 98% |
| ARC-AGI-3 | 99.9% | OpenAI’s Responses API harness |
| ExploitBench | 100.0% | — |
| GPQA Diamond | 96.0% | — |
For ARC-AGI-3, OpenAI’s footnote says the model “was run with our responses API harness, which changes two settings to better match real-world performance.”
What ARC Prize reports
ARC Prize, which runs the ARC-AGI benchmarks, published its own evaluation of GPT-6 Astra on September 3. On the ARC-AGI-3 Semi-Private set, ARC Prize reports 62.7% with its Standard harness and 99.9% for Astra (high) with the Provider Adapter harness, which it says “preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.”
ARC Prize says it will report “both Standard harness and Provider Adapter harness results on the ARC-AGI leaderboard, with each evaluation condition clearly labeled.” The 99.9% figure therefore depends on the harness: under ARC Prize’s provider-neutral Standard setup, the score is 62.7%. ARC Prize adds that saturating the benchmark “would not represent ‘proof of achieving AGI’.”
Scores from both OpenAI and ARC Prize, with their harnesses labeled, are tracked on our benchmark leaderboard.
What it changes
GPT-6 Astra is the top model of OpenAI’s GPT-6 family, which added GPT-6 Sol and GPT-6 Luna on September 22. For developers, the practical changes are the Codex notes feature for long sessions and the API restrictions on reasoning effort and sampling settings. For readers comparing benchmark claims, the ARC-AGI-3 result shows how much a headline score can move with the harness used to run it.
GPT-6 Astra is listed on our AI model release timeline.




