What the efficiency claim actually measures
ARC-AGI-3 drops agents into unfamiliar turn-based worlds and scores them on exploration, modeling, goal-setting, and planning. Humans can solve 100% of the environments. The yardstick that moved François Chollet is action efficiency: how many moves the system needs relative to a human baseline built from roughly 500 first-time public testers.
In the Provider Adapter setup, Astra (max) used fewer actions than that baseline on 96% of levels and 51.7% fewer actions per level on average. ARC calls that human parity on its efficiency metric. Replays show Astra inventing compact algebraic shorthand for game state, writing ordered plans like extend8 to3; retract10 to2, and, in a separate PRO-LONG tool sandbox, spinning up small game-specific libraries. That last setup is model-plus-tools, ARC warns, so it is not the same as the controlled human sessions.
The rest of the scoreboard refuses to agree
As The Decoder reported, Epoch’s ECI puts Astra ahead of Sol, Fable 5.1, and Opus 5, while Artificial Analysis’s Intelligence Index leaves Astra tied with Sol at 61 and behind Fable 5.1. On ARC-AGI-3’s Standard harness, Sol sat near 7.8% and Claude Opus 5 near 30%. The jump is real under ARC’s rules. The cross-benchmark story is split.
Chollet has said saturating ARC-AGI-3 is not proof of AGI, and ARC’s own blog repeats that line. He also said progress arrived about twice as fast as he expected when the benchmark launched, and answered a question about his 2030 AGI forecast with “Sooner.” That is a timeline update from the person who built the residual-gap test, not a certificate that Brockman’s “AGI era” press line is settled science.
Judgment
Sonar’s AI Wars board exists because people search for arena scores, coding-agent ranks, and Anthropic-versus-OpenAI bragging rights. Astra’s ARC post is catnip for that demand. It is also a tutorial in how a leaderboard can tell the truth and still mislead.
62.7% under shared conditions is the number that should drive lab-to-lab comparison. 99.9% under a private adapter is a capability ceiling with the house furniture included. ARC Prize did the honest thing and published both. The rest of the industry will flatten them into one AGI slide.
Treat the harness as part of the claim. If a score needs opaque state the public cannot inspect, it belongs next to the monitorability caveats OpenAI already filed on the same model, not above them.
Letters
0
No letters yet.