Start the day here

AI — Benchmarks — OpenAI

How a Private Harness Pushed Astra to 99.9% on ARC-AGI-3

ARC Prize published GPT-6 Astra’s ARC-AGI-3 numbers on the same day OpenAI started rolling the model out. The headline is easy to meme: nearly perfect. The table underneath is harder. Under ARC’s Standard harness, a provider-neutral interface that forces the model to keep its own visible notes, Astra at max reasoning scored 62.7% on the Semi-Private set for about $26,098. Under the Provider Adapter harness, which preserves opaque reasoning state between requests and uses compaction for long runs, Astra at high reasoning scored 99.9% for about $18,817.

Those are both state-of-the-art on ARC’s board. They are also answers to different questions. The Standard harness asks how models compare when everyone shares the same skinny interface. The Provider Adapter asks how well a lab’s model performs when it can use the context machinery that lab built for it. ARC Prize will report both, labeled. Good. The internet will quote the second number without the label.

7 min read
White mathematical formulas and scientific diagrams packed onto a black chalkboard surface

What the efficiency claim actually measures

ARC-AGI-3 drops agents into unfamiliar turn-based worlds and scores them on exploration, modeling, goal-setting, and planning. Humans can solve 100% of the environments. The yardstick that moved François Chollet is action efficiency: how many moves the system needs relative to a human baseline built from roughly 500 first-time public testers.

In the Provider Adapter setup, Astra (max) used fewer actions than that baseline on 96% of levels and 51.7% fewer actions per level on average. ARC calls that human parity on its efficiency metric. Replays show Astra inventing compact algebraic shorthand for game state, writing ordered plans like extend8 to3; retract10 to2, and, in a separate PRO-LONG tool sandbox, spinning up small game-specific libraries. That last setup is model-plus-tools, ARC warns, so it is not the same as the controlled human sessions.

The rest of the scoreboard refuses to agree

As The Decoder reported, Epoch’s ECI puts Astra ahead of Sol, Fable 5.1, and Opus 5, while Artificial Analysis’s Intelligence Index leaves Astra tied with Sol at 61 and behind Fable 5.1. On ARC-AGI-3’s Standard harness, Sol sat near 7.8% and Claude Opus 5 near 30%. The jump is real under ARC’s rules. The cross-benchmark story is split.

Chollet has said saturating ARC-AGI-3 is not proof of AGI, and ARC’s own blog repeats that line. He also said progress arrived about twice as fast as he expected when the benchmark launched, and answered a question about his 2030 AGI forecast with “Sooner.” That is a timeline update from the person who built the residual-gap test, not a certificate that Brockman’s “AGI era” press line is settled science.

Judgment

Sonar’s AI Wars board exists because people search for arena scores, coding-agent ranks, and Anthropic-versus-OpenAI bragging rights. Astra’s ARC post is catnip for that demand. It is also a tutorial in how a leaderboard can tell the truth and still mislead.

62.7% under shared conditions is the number that should drive lab-to-lab comparison. 99.9% under a private adapter is a capability ceiling with the house furniture included. ARC Prize did the honest thing and published both. The rest of the industry will flatten them into one AGI slide.

Treat the harness as part of the claim. If a score needs opaque state the public cannot inspect, it belongs next to the monitorability caveats OpenAI already filed on the same model, not above them.

Live scoreboardFollow the AI race on AI Wars

Lab rankings, model preference, API volume, coding-agent heat, open-source stars, and prediction markets.

Related stories

A crowded market table of colorful tin toy robots and miniature cars under bright daylightWorld

Europe Put Roblox Under Its Strictest Platform Rules

Today

A closed metal padlock stamped HARDENED resting on a backlit laptop keyboardWorld

Even OpenAI's Kill Switch Still Needs a Human

Today

A bright blue Ethernet cable plugged into numbered port 074 on a black patch panelWorld

Even OpenAI's Monitors Miss Astra When Pressed

Today

Letters

0

No letters yet.

Write a letter