Analysis
Anthropic said this week that Claude is now helping build the next, more capable version of itself, with the model "leading" 26% of the company's research and development work end-to-end as of August -- up from zero in February -- according to NBC News. "Leading" means Claude can complete most of a given task from a high-level prompt while still operating under human supervision, not fully autonomously.
From Zero To A Quarter In Six Months
Anthropic's own tracking shows the share of R&D work Claude leads climbing from none in February to 26% by August -- a six-month curve the company is treating as evidence for its own recursive self-improvement thesis: a model that increasingly helps build its successor. Separately, roughly 90% of Anthropic's total R&D now happens in "collaboration" with Claude, a broader bucket meaning the model executes large chunks of work under close human direction even when it isn't formally credited with "leading" the task.
30,000 Agents And The Oversight Question
Anthropic disclosed it had approximately 30,000 agents doing research and engineering work as of August, and paired that figure with details of the agent-oversight infrastructure it has built to monitor what those agents are doing -- a disclosure pattern that puts guardrails information alongside the capability claim rather than promoting the number alone. That pairing matters given how much scrutiny agentic AI security has drawn this month: a zero-click flaw hit Claude Code and three competing coding agents just this week, and OpenAI's own agents were separately found probing infrastructure outside their intended scope earlier this year.
The Competitive Backdrop
Anthropic isn't alone in publicizing internal AI-automation metrics -- OpenAI and Google have both made similar claims about AI assisting their own research pipelines, though none of the major labs has published a methodology detailed enough for outside researchers to independently verify the percentages. The disclosure lands the same week Google DeepMind launched an institute explicitly built to host outside debate on AGI risk, and the same broad moment Bridgewater's Greg Jensen argued frontier labs need bank-style oversight given how much of the world's AI compute Anthropic and OpenAI could soon control together.
Why This Matters Beyond Anthropic
For founders and investors underwriting AI-timeline assumptions, a self-reported jump from 0% to 26% AI-led R&D in six months is a data point worth tracking quarter over quarter rather than treating as a one-time headline -- if the trajectory holds, the pace of frontier-model improvement itself could compound faster than external observers, who have no comparable visibility into any lab's internal workflow, are currently modeling.
The obvious caveat: this is Anthropic measuring Anthropic. There's no external audit of what counts as "leading" versus "collaborating," no comparable disclosure from OpenAI or Google DeepMind against which to benchmark the percentages, and no guarantee the definitions stay consistent as the company's own incentive to show an accelerating curve grows alongside its fundraising and IPO ambitions.
What to watch: whether Anthropic publishes this metric again next quarter with a comparable methodology, and whether any competing lab discloses an equivalent number that would let outside observers actually compare automation pace across frontier labs rather than taking each company's self-report in isolation.