Start the day here

AI — Labs — Recursive Self-Improvement

The Problem With Labs Racing to Automate Their Own Research

When Jack Clark returned from paternity leave in February, Anthropic colleagues barely wrote code. They managed five or six Claudes, which sometimes managed more Claudes. TIME's Harry Booth reported the scene on August 7. Clark heard an early form of recursive self-improvement: models accelerating the research cycle that builds the next models.

Anthropic's June Institute report, When AI Builds Itself, quantifies the shift. Code volume per person rose eight-fold. Claude wrote 80 percent of it. Outside skeptics, Gary Marcus among them, called the terror marketing and the metric crude. Clark admits the yardstick is rough. The institutional fact remains: the people who set growth conditions for neural nets are already outsourcing most of the typing.

7 min read
Close-up of a car engine serpentine belt threaded through metal and plastic pulleys

The race calendar

OpenAI's chief scientist, Jakub Pachocki, told TIME the company wants full automation of its AI researchers by March 2028, with a "virtual intern" target this September. GPT-5.3 Codex, released in February, had a significant hand in its own development "from start to finish," according to Amelia Glaese. By July, experiments per researcher had doubled. A months-old startup named Recursive Superintelligence raised $650 million to chase the same prize.

The pause that is not a pause

Kaplan and Pachocki both say the race should slow once models can train successors with little human help. They signed the July letter asking Washington for coordination machinery (Sonar covered that brake ask). Clark left policy to run the Anthropic Institute and spends this summer defining what a pause would require. He puts autonomous self-improvement by 2028 at 60 percent odds. He is content with brinkmanship for now. The brake lever is a contingency plan, built by the same firm flooring the pedal.

That is the problem. A lab that profits from compression cannot be the sole designer of the stop condition. Helen Toner wants consistent, multi-company metrics published on a schedule so the world can track acceleration without taking any one firm's word. Anthropic's Responsible Scaling Policy promises tighter controls once two years of research compress into one. Clark says the firm sees acceleration and still lacks a cumulative measure. "I can't give you a specific number, because we don't have a measure," he told TIME.

Bottleneck optimists still have a case. Princeton's Arvind Narayanan tested Claude on open-ended research and watched it handle engineering while stalling on judgment. Compute remains finite. Sample efficiency is unsolved. Gradual pain across institutions may matter more than a single takeoff day. Those caveats do not erase the product plans. They describe a world where society adapts under a shrinking margin while the labs keep shipping.

Dave Orr, Anthropic's head of safeguards, put the feeling in a car metaphor: the margin for error shrinks as the speed rises. Circuit breakers in markets trip without waiting for a committee vote. Bounded autonomy is the same idea for agents. A pause designed only inside the race is a brochure. The useful demand is external measurement, enforceable boundaries, and a stop condition that does not report to the same P&L as the accelerator.

Live scoreboardFollow the AI race on AI Wars

Lab rankings, model preference, API volume, coding-agent heat, open-source stars, and prediction markets.

Related stories

A hand holds a yellow-cased phone showing a calculator app over a folder of tax formsBusiness

Zuckerberg Told Trump a National AI Regulator Was Flawed

Today

An empty operating room with a black surgical table under twin ceiling lights and wall monitorsBusiness

How OpenAI Wired ChatGPT Into Epic's Patient Charts

Today

Gold dome and white Corinthian columns of the Massachusetts State House against a clear blue skyBusiness

Massachusetts's 120-Day AI Evaluators

Today

Letters

0

No letters yet.

Write a letter