Astra’s ARC Prize-verified 98.6% score on ARC-AGI-3 is an astonishing result. It is not a clean measurement of Astra’s model weights alone.
ARC Prize has now supplied the missing denominator. Its verified results show 98.55%—the launch figure rounded to 98.6%—for GPT-6 Astra at max reasoning on the ARC-AGI-3 Semi-Private set using a Provider Adapter harness. The same model and reasoning level scored 62.71% with ARC Prize’s Standard harness. Meanwhile, NVIDIA’s 100% with Claude Opus 5 inside AVO remains a result on the 25 public demonstration environments.
That paired result makes the benchmark paradox more concrete, not less. Astra is plainly a major capability advance, and ARC Prize independently verified the run. But a near-perfect score still measures a configured agent system: the gap between 62.71% and 98.55% shows how much the exposed capability changes when reasoning continuity and compaction change around the same named model.
Kingy verdict: Astra’s 98.6% is now a verified system result: GPT-6 Astra at max reasoning, the ARC-AGI-3 Semi-Private set, ARC Prize’s Provider Adapter harness and $17,332 in recorded cost. ARC Prize also reports a 62.7% Standard-harness baseline costing $26,098, plus a higher 99.9% Provider Adapter run at high reasoning. This proves Astra can clear ARC-AGI-3’s bounded bar with extraordinary action efficiency. It does not isolate model weights from the adapter or prove AGI. NVIDIA’s 100% still proves public-set saturation by AVO plus Claude Opus 5, not private-set generalization.
Evidence update, 3 September 2026: ARC Prize has published an independent technical account and a verified Astra results page covering 12 harness-and-effort configurations. Those sources now establish the ARC-AGI-3 split, harness class, reasoning effort, recorded cost and per-environment outcomes. DataCurve has also added Astra to its official DeepSWE leaderboard, where the best configuration scores 74% ±3 in the shared mini-swe-agent setup. OpenAI’s GPT-6 Astra System Card confirms the release and deployment safeguards but does not itself mention ARC-AGI-3, DeepSWE or GDPval. Our broader Astra evidence tracker covers the model’s release trail and remaining product questions.
The system card does add relevant platform context. OpenAI says its Astra misalignment monitor operates across Codex, ChatGPT and the Responses API, reviewing the model’s chain of thought alongside its actions and conversation inputs and outputs. It also distinguishes stateless Responses API requests from persistent conversations: without persistent reasoning or a conversation identifier, the monitor cannot reconstruct a complete trajectory. That is a deployment-safeguard disclosure, not a benchmark-method disclosure, but it reinforces the central point that an Astra result can depend on more than model weights alone.
The benchmark changed before Astra arrived
OpenAI supplied the best reason to resist model-only interpretations in July. Using GPT-5.6 Sol on ARC-AGI-3’s 25 public environments, the company first ran the benchmark’s generic harness. That setup kept a record of moves and brief notes but discarded private reasoning after every action. It also used rolling truncation, dropping the oldest context after the interaction history crossed 175,000 characters.
The result was 13.3%.
OpenAI then used a Responses API harness that passed the previous response ID between actions, retaining reasoning across tool calls. It added compaction to compress long histories without simply deleting the oldest experience. The same named model reached 38.3% while producing roughly one-sixth as many output tokens. OpenAI’s conclusion was unusually direct: evaluations measure a model together with API settings, harness design and prompting. That experiment does not make the higher score illegitimate. It shows why the harness belongs in the product name of the result.

ARC Prize now calls Astra’s native configuration the Provider Adapter harness. It preserves opaque reasoning state between requests and uses compaction for longer conversations. At max reasoning, the adapter scored 98.55% on Semi-Private versus 62.71% for the Standard harness—a 35.84-point absolute gap with the same named model and effort. Across the 167 game-reasoning pairs solved by both harnesses, ARC Prize reports that Provider Adapter runs were about 3.66 times faster by aggregate recorded elapsed time and used 49% fewer total tokens. The public record still cannot expose Astra’s opaque reasoning or isolate every causal component, but it now quantifies the system effect directly.
NVIDIA’s 100% proves a system can saturate the public set
NVIDIA’s result is more specific. Its AVO technical account says Claude Opus 5 completed all 183 levels across all 25 public ARC-AGI-3 environments. It used 6,624 environment actions. The model received text-only 64-by-64 observations; AVO added persistent memory, a supervisor and feedback-and-recovery loops.
NVIDIA contrasts that with ARC Prize’s roughly 30% Claude Opus 5 result and with VISTA, another system using the same model family that required 7,542 actions. The comparison is valuable, but it is not a controlled ablation. NVIDIA explicitly notes differences in backend, observation format, memory and context. AVO’s result proves that this particular model-and-architecture combination can solve the public set. It does not isolate the number of percentage points caused by persistent memory, the supervisor or any single component.
Nor is NVIDIA’s 100% automatically stronger evidence than Astra’s 98.6%. We now know the percentages do not share a denominator: NVIDIA saturated the public demonstrations, while Astra’s 98.55% was verified on Semi-Private environments using the Provider Adapter. Astra therefore carries stronger evidence of first-contact generalization. NVIDIA’s packet remains useful evidence about what explicit persistent memory, supervision and recovery can accomplish on the public set.
Public, semi-private and private are different claims
ARC Prize is explicit about this boundary. Its ARC-AGI-3 technical report defines beating the benchmark around human-level action efficiency averaged over private environments encountered for the first time. The public environments are demonstrations: useful for learning the interface, debugging a harness and exploring agent design, but not sufficient evidence of general progress if a system has been tuned to them.
The verified ARC Prize record now makes that distinction inspectable. It lists six reasoning levels under each of two harness classes. On Semi-Private, Standard-harness scores range from 17.45% to 62.71%; Provider Adapter scores range from 96.72% to 99.95%. The launch’s 98.6% maps to the 98.55% max-reasoning Provider Adapter run, while the best observed result is 99.95% at high reasoning.
ARC Prize’s open benchmarking repository now makes another important separation. The standard harness uses provider-neutral text history and manual rolling context. Provider-adapter harnesses may use native reasoning-state continuity and compaction. The games, actions, limits and scoring can stay fixed while context handling changes. The repository says these result classes should be labeled separately: the standard harness is better for controlled provider comparisons; adapters are better for measuring a provider’s native agent stack.
ARC Prize has now adopted exactly that taxonomy. Astra’s Provider Adapter score should not be dismissed because the harness is capable, and it should not be presented as a weights-only result. The correct label is GPT-6 Astra plus Provider Adapter on Semi-Private, with the Standard-harness score beside it.

The 98.6% score now has the missing labels
System. GPT-6 Astra ran at max reasoning through ARC Prize’s Provider Adapter, which preserves opaque reasoning between requests and uses compaction. The matched Standard harness lets the model carry forward visible notes it chooses to keep but does not preserve provider-native reasoning state.
Split. The 98.55% result is on ARC-AGI-3 Semi-Private, not the 25 public demonstrations. ARC Prize’s verified record separately publishes public-demo and Semi-Private outcomes.
Budget. ARC Prize reports $17,332 for the max-reasoning Provider Adapter run, versus $26,098 for Standard at max; the highest score, 99.95% at high reasoning, cost $18,817. It also reports that Astra max used fewer actions than the median tested human on 96% of completed levels and 51.7% fewer actions per level on average.
Reproduction. ARC Prize has validated the entry, published per-environment results and linked replays through its verified record. Fully open reproduction remains constrained by a closed model and a Semi-Private set, but the result is no longer only a vendor launch claim.
DeepSWE confirms frontier performance, not a clear lead
DataCurve now lists GPT-6 Astra at 74% ±3 with xhigh reasoning on DeepSWE v1.1. The displayed score is consistent with OpenAI’s 74.1% launch figure after rounding. It ties Gemini 3.8 Flash High at 74% ±1 and Claude Opus 5 Max at 74% ±4 by rounded point estimate, while GPT-5.6 Sol Max sits at 73% ±3. The uncertainty ranges overlap, so Astra has no statistically clear accuracy lead over that frontier group.
DeepSWE is designed to make comparisons cleaner. Its benchmark paper describes repository-level software tasks with executable verifiers, and every leaderboard model runs through mini-swe-agent with the same Bash tool and shared prompt. DataCurve’s 3 September changelog says it added Astra results across low, medium, high, xhigh and max reasoning at the expected launch rate card.
The efficiency figures are more distinctive than the pass-rate gap. Astra’s best row averages $6.52, 30,000 output tokens and 29 agent steps per trial. Gemini 3.8 Flash averages $2.36, 143,000 tokens and 166 steps; Claude Opus 5 averages $11.84, 118,000 tokens and 99 steps; GPT-5.6 Sol averages $6.46, 60,000 tokens and 61 steps. The shared setup makes DeepSWE a cleaner model comparison than ARC-AGI-3. Astra is a top-tier coding model with unusually low token and step counts, not an accuracy category of its own.
Why the missing GDPval result matters—and why it would not settle AGI
The launch framing reportedly included OpenAI president Greg Brockman’s declaration, “Welcome to the AGI era.” That makes the absence of a GDPval result conspicuous. OpenAI built GDPval to evaluate economically valuable work products across 44 occupations and nine industries. A strong result would be relevant evidence for a claim about broadly useful intelligence, rather than competence inside games, repositories and specialist technical tasks.
But GDPval would not be a magic counterweight. OpenAI describes the initial benchmark as a one-shot evaluation; it does not capture the iterative, long-horizon collaboration where Astra is supposed to excel. The fair conclusion is narrower: no GDPval number means the launch packet, as publicly visible at the cutoff, did not quantify Astra’s breadth on OpenAI’s own economic-work benchmark. It does not mean Astra performed poorly or was never tested.
What each result actually proves
| Result | What the public evidence supports | What remains unproven |
|---|---|---|
| Astra: 98.6% | ARC Prize verifies 98.55% at max reasoning on ARC-AGI-3 Semi-Private with Provider Adapter for $17,332; Standard at max scored 62.71% for $26,098. Provider Adapter reached 99.95% at high. | The weights-only contribution, provider-internal reasoning state and fully open reproduction on the non-public set. |
| AVO + Claude Opus 5: 100% | NVIDIA documents complete coverage of 183 levels in the 25 public environments using 6,624 actions. | Private-set generalization and a controlled estimate of how much AVO—not other configuration differences—caused the gain. |
| GPT-5.6 Sol: 13.3% → 38.3% | OpenAI demonstrates that retained reasoning and compaction can nearly triple a public-set score while reducing output tokens. | Whether the same magnitude of harness effect applies to Astra or unseen private tasks. |
| Astra: 74.1% DeepSWE | DataCurve lists GPT-6 Astra xhigh at 74% ±3 in the common mini-swe-agent setup, averaging $6.52, 30,000 output tokens and 29 steps per trial. | A statistically clear accuracy lead over Gemini 3.8 Flash or Claude Opus 5; DataCurve’s public table rounds the point estimate to 74%. |
The paradox is not that agent scaffolding makes a benchmark fake. Long-horizon intelligence may require memory, state management and recovery, just as useful computing systems require more than a processor. The mistake is presenting a system score as though it cleanly identifies the capability of one component.
Astra’s result should raise expectations—and ARC Prize’s disclosure packet now meets much of the standard the launch headline initially lacked. It names the split and harness class, reports paired Standard and Provider Adapter runs, publishes costs and action-efficiency evidence, and provides verified per-environment outcomes. DeepSWE separately confirms top-tier coding performance in a shared setup, with unusually low token and step counts but no clear accuracy lead over its closest peers.
The resulting conclusion is stronger and narrower than “AGI achieved.” Astra clears a difficult bounded benchmark with human-beating action efficiency, while the paired harness results show that memory and context orchestration remain part of what the score measures. The launch number now survives scrutiny precisely because the system around the model is visible.
The Kingy Brief
Follow The Kingy Brief.
One consequential launch, one pricing, limit, or shutdown change, one hands-on test, one exact prompt or Test Pack, and one try / watch / skip verdict.
Free · Choose your subjects · Double opt-in · Unsubscribe anytime
