ARC Prize published its evaluation of OpenAI’s GPT-6 Astra on ARC-AGI-3 on September 3, 2026, reporting a score of 62.7% on the Semi-Private set with its Standard harness and 99.9% for Astra (high) with a Provider Adapter harness. The results are in the ARC Prize post OpenAI’s GPT-6 Astra on ARC-AGI-3.
On this page
Two harnesses, two results
ARC Prize ran Astra in two configurations, which differ in what the model keeps from one request to the next.
- Standard harness. ARC Prize says it “enables a model to carry forward notes it chooses to keep with it throughout the environment.” In this setup Astra scored 62.7% on ARC-AGI-3 Semi-Private.
- Provider Adapter harness. This harness “preserves opaque reasoning state between requests and uses compaction for longer conversations, allowing the model to reuse prior work.” Here Astra (high) scored 99.9%.
ARC Prize describes the Standard harness as asking “how models compare under the same minimal, provider-neutral interface.” With the Provider Adapter harness, it says, Astra’s best observed score on the Semi-Private set “increased from 62.7% to 99.9%.” ARC Prize also reports the evaluation cost of each result: $26K for the Standard harness score and $19K for the Astra (high) Provider Adapter score.
| Configuration | Score (ARC-AGI-3 Semi-Private) | Evaluation cost reported by ARC Prize |
|---|---|---|
| GPT-6 Astra, Standard harness | 62.7% | $26K |
| GPT-6 Astra (high), Provider Adapter harness | 99.9% | $19K |
Action efficiency against the human baseline
ARC-AGI-3 is built from interactive game environments, and ARC Prize compares the number of actions a model takes with a human baseline. According to the post, Astra (max) in the Provider Adapter harness “used fewer actions than the human baseline on 96.0% of levels,” and used 51.7% fewer actions per level on average.
The human reference comes from approximately 500 members of the general public whom ARC Prize tested, and who, it says, were “not selected for puzzle-solving experience or ability.” For each level, ARC Prize defines the human baseline as “the median action count among players.” It also states that “humans can solve 100% of the environments.”
What the result does and does not show
ARC Prize draws a line around its own result. The post says that “saturating the benchmark would not represent proof of achieving AGI.” It also reports the two harness results side by side rather than replacing one with the other: the Standard harness measures models through one provider-neutral interface, while the Provider Adapter harness keeps the provider’s reasoning state between requests and compacts longer conversations.
For readers comparing models, that means an ARC-AGI-3 score only makes sense with its harness attached. A 62.7% Standard-harness result and a 99.9% Provider Adapter result describe the same model under different rules.
ARC-AGI-3 scores, with their harness and source, are collected on our benchmark leaderboard; GPT-6 Astra’s release is in the AI model release timeline.
Sources
- OpenAI’s GPT-6 Astra on ARC-AGI-3 | ARC Prize — ARC Prize




