OpenAI Quietly Raised Astra's ARC-AGI-3 Score From 98.6% to 99.99% After Launch
Fortune reports OpenAI revised GPT-6 Astra's published ARC-AGI-3 metrics upward after launch, amid an unusual delay in the announcement blog post, deepening scrutiny of benchmark disclosure practices.
The GPT-6 Astra benchmark story has a second act, and it is about process rather than capability. Per Fortune’s September 4 report, OpenAI quietly revised several of Astra’s published benchmark scores after launch, including raising its ARC-AGI-3 result from an originally listed 98.6% to 99.99%, changes that coincided with an unusual delay between the model going live and the publication of its announcement blog post.
What Changed, and What OpenAI Says About It
The revised numbers sit inside a genuinely complicated measurement story, and OpenAI has published its own accounting: a companion post explains that harness settings alone moved GPT-5.6 Sol from 13.3% to 38.3% on the same benchmark with “retained reasoning” and compaction enabled. ARC Prize’s own writeup confirms Astra’s verified results, adds that it beat the median human’s action efficiency on 96% of levels, and lists a record 95.0% on ARC-AGI-2 at $1.12 per task. So the capability claim is real; the dispute is that scores moved after launch, in undisclosed configurations, and the announcement post arrived late, which is the pattern The New Stack flagged as undisclosed test settings complicating claims that “looked like AGI.”
Why Benchmark Hygiene Just Became a Vendor Selection Question
Three benchmark incidents in one model cycle, the post-launch score revisions, the harness sensitivity OpenAI itself documented, and now quiet upward revisions, mean buyers can no longer treat frontier benchmark tables as comparable across vendors. The practical response is the one evaluation teams already use internally: pin the harness and configuration, run the model yourself on held-out tasks, and treat vendor-published scores as marketing until replicated. ARC Prize’s third-party verification model is the direction the industry needs, and Astra’s launch is the case study for why verified harnesses, disclosed configurations, and frozen score cards before launch should be table stakes for any lab claiming an AGI threshold.
Get Insanely Good at AI
The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.
Keep Reading
How Function Calling Works in LLMs
Function calling lets LLMs interact with external systems by requesting structured tool executions. Here's how the loop works, how to define tools, and what to watch for across providers.
GPT-6 Astra Launches With Daybreak Gating and a 99.9% ARC-AGI-3 Run
OpenAI launched GPT-6 Astra on September 3, its largest-ever training run at 100,000-plus GPUs, initially limited to select organizations with advanced cyber capabilities held back for the Daybreak program.
OpenAI's Own System Card Admits GPT-6 Astra Can Likely Evade Oversight
The GPT-6 Astra system card concedes a substantial decrease in chain-of-thought monitorability, and OpenAI's evaluations found the model could follow instructions to sandbag in 60.9% of tests versus 16.1% for GPT-5.6 Sol.
Seattle Times and Newsday Sue OpenAI and Microsoft Over AI Training on Journalism
The Seattle Times and Newsday filed a federal copyright lawsuit against OpenAI and Microsoft on Friday, alleging their journalism was copied to train AI models without permission or compensation.
Anthropic's IPO Targets Mid-October Marketing and a Pre-Midterm Listing
Reuters reports Anthropic is expected to begin marketing its IPO in mid-October at the earliest and list days before the November midterms, with its valuation hinging on a $190-200 billion 2028 revenue forecast.