Ai Engineering 2 min read

OpenAI Quietly Raised Astra's ARC-AGI-3 Score From 98.6% to 99.99% After Launch

Fortune reports OpenAI revised GPT-6 Astra's published ARC-AGI-3 metrics upward after launch, amid an unusual delay in the announcement blog post, deepening scrutiny of benchmark disclosure practices.

The GPT-6 Astra benchmark story has a second act, and it is about process rather than capability. Per Fortune’s September 4 report, OpenAI quietly revised several of Astra’s published benchmark scores after launch, including raising its ARC-AGI-3 result from an originally listed 98.6% to 99.99%, changes that coincided with an unusual delay between the model going live and the publication of its announcement blog post.

What Changed, and What OpenAI Says About It

The revised numbers sit inside a genuinely complicated measurement story, and OpenAI has published its own accounting: a companion post explains that harness settings alone moved GPT-5.6 Sol from 13.3% to 38.3% on the same benchmark with “retained reasoning” and compaction enabled. ARC Prize’s own writeup confirms Astra’s verified results, adds that it beat the median human’s action efficiency on 96% of levels, and lists a record 95.0% on ARC-AGI-2 at $1.12 per task. So the capability claim is real; the dispute is that scores moved after launch, in undisclosed configurations, and the announcement post arrived late, which is the pattern The New Stack flagged as undisclosed test settings complicating claims that “looked like AGI.”

Why Benchmark Hygiene Just Became a Vendor Selection Question

Three benchmark incidents in one model cycle, the post-launch score revisions, the harness sensitivity OpenAI itself documented, and now quiet upward revisions, mean buyers can no longer treat frontier benchmark tables as comparable across vendors. The practical response is the one evaluation teams already use internally: pin the harness and configuration, run the model yourself on held-out tasks, and treat vendor-published scores as marketing until replicated. ARC Prize’s third-party verification model is the direction the industry needs, and Astra’s launch is the case study for why verified harnesses, disclosed configurations, and frozen score cards before launch should be table stakes for any lab claiming an AGI threshold.

Get Insanely Good at AI

Get Insanely Good at AI

The book for developers who want to understand how AI actually works. LLMs, prompt engineering, RAG, AI agents, and production systems.

Keep Reading