

Anthropic has started publishing a number that most AI labs have kept quiet: how much of their own research and development work is now being done by their own AI. In a report released on September 17, the company said Claude led 26 percent of its model R&D work as of August 2026, up from zero percent in February, with roughly 90 percent of all R&D now happening in some form of collaboration with the model. Around 30,000 individual Claude agents were running research and engineering tasks that month — not writing code on request, but proposing experiments, evaluating results, and shaping decisions about what Anthropic builds next.
From Zero to 26 Percent in Six Months
The headline figure is the trajectory, not just the snapshot. Anthropic says Claude’s share of R&D leadership sat at 0 percent in February 2026 and climbed to 26 percent by August — a six-month curve that the company is treating as a metric worth tracking in public, the way a cloud provider might report uptime. The company has separately disclosed that more than 80 percent of the code merged into its own production codebase in May 2026 was authored by Claude rather than by human engineers, a figure that was in the “low single digits” before Claude Code reached research preview in February 2025. Anthropic says its engineers are now merging roughly eight times as much code per quarter as they did between 2021 and 2025.
Those two figures describe different things — one is code volume, the other is R&D leadership — but they point in the same direction. Claude is no longer just a tool that Anthropic’s researchers use; increasingly, it is deciding what gets tried.
Thirty Thousand Agents at Work
The scale is the part that is easy to skim past. Roughly 30,000 Claude agents were active on research and engineering work in August, according to Anthropic’s disclosure — a number that dwarfs the size of Anthropic’s human research staff and gives a concrete sense of what “AI-assisted R&D” looks like in practice inside a frontier lab today. It is not a handful of copilots sitting next to individual engineers; it is a standing workforce of software agents running experiments, triaging results, and feeding conclusions back into the next round of model development, largely without a person in the loop for each step.
Anthropic frames the 90 percent “collaboration” figure carefully: it does not mean Claude is unsupervised on nine out of every ten tasks, but that almost all of the company’s R&D now involves the model in some capacity, from drafting code to designing evaluations. The 26 percent “leads” figure is the narrower and more consequential one — it describes work where Claude, rather than a human researcher, is setting the direction.
Why Anthropic Is Saying This Out Loud
Publishing this kind of metric is unusual, and Anthropic says that is the point. The company has argued that the industry lacks any shared, standardized way to measure how close AI labs are getting to what researchers call recursive self-improvement — a model’s ability to autonomously design and build its own successor with minimal human involvement. Without public numbers, Anthropic says, society has no way to gauge how fast that threshold is approaching at any given lab, Anthropic included. By putting its own figures out first, the company is trying to set a norm and is calling on other frontier labs to publish comparable metrics using a similar methodology.
It is a deliberately double-edged disclosure. The same numbers that Anthropic is using to demonstrate Claude’s research capability are also the numbers the company is using to argue that the industry is approaching a genuinely new kind of risk — one where the pace of progress is no longer bottlenecked by human researcher headcount.
The Self-Improvement Warning and the Case for a Pause Option
Alongside the metrics, Anthropic has been explicit about what worries it. The company has warned that if current trajectories continue, rare misalignments in an AI system today could compound across successive, self-built generations of models until they become more frequent and less understood — a slide toward a point where humans can no longer reliably control the systems they built. CEO Dario Amodei has been the most visible voice making this case publicly this year, arguing in essays and interviews that the industry should deliberately pace its own progress rather than treat speed as the only variable that matters.
Anthropic’s specific ask is narrower than an outright halt: it wants the option to slow or pause frontier development kept open, not exercised unilaterally. The company has said it would only actually pause its own work if competing labs at the frontier adopted verifiable, equivalent safeguards at the same time — an acknowledgment that a one-sided slowdown just hands the lead to whichever competitor keeps racing. That conditional framing has become a recurring theme in Anthropic’s public statements this year, echoed in its endorsement of proposed federal oversight measures like the FRONTIER Act and in the formation of its new Institute focused on AGI-related public debate.
What Happens When the Toughest Critic Is the One Building the Future
There is an obvious tension sitting underneath all of this: Anthropic is simultaneously the company most invested in Claude becoming a faster, more capable research partner and the company issuing the loudest public warnings about what happens if that trend goes unchecked. Skeptics have pointed out that a company touting an 8x productivity multiplier and a model that leads more than a quarter of its own R&D has every commercial incentive to keep pushing that number higher, whatever the accompanying safety language says. Anthropic’s counterargument is that transparency itself is the safeguard: if the company hides these numbers, outsiders including regulators, researchers, and rival labs have no way to independently assess how close any lab actually is to the handoff point Amodei has warned about.
Whether or not other labs follow Anthropic’s lead in publishing similar figures, the 26 percent number is likely to become a reference point in the broader debate over AI development speed heading into next year — the first time a frontier lab has put a specific, trending percentage on how much of its own future is already being decided by its own model.