The Anthropic R&D Automation Index says Claude now leads 26 percent of the company’s own AI research and development. Anthropic published the figure this week in a post from its Institute. The post proposes three measurements that every frontier lab could report. The goal is to show how fast AI is building the next version of itself.
“Leads” has a specific meaning here. Anthropic rates its work on a six-step scale from Epoch AI. It runs from AL0, no AI involvement, to AL5, fully autonomous. At AL4, the AI completes most of a task from a high-level prompt while a human supervises. At AL3, it “collaborates” under close human direction. As of August 2026, more than 90 percent of Anthropic’s model R&D sits at AL3 or higher. However, nothing measured has reached AL5.
How the Anthropic R&D Automation Index works
The climb has been steep. Claude led under 1 percent of tasks in February 2026. By May it was 12 percent, then 22 percent in July and 26 percent in August. To build the index, Anthropic sampled 20 percent of staff in each R&D department every week in July. A Claude agent then read their Slack messages and internal documents. It listed about 15,000 tasks, and Claude sorted them into a tree of 542 categories.
Next, a separate Claude judge rated how automated each category is. In other words, Claude graded Claude. Anthropic admits the weakness and checked the ratings against the staff who own the work. The model matched humans exactly 59 percent of the time. By comparison, humans matched each other only 35 percent of the time. As a result, the company wants third parties, or other labs’ models, to verify future numbers.
Thirty thousand agents under watch
The second measurement covers oversight. Anthropic says about 30,000 agents were working on its main internal platform at any one time in August. Every action passes through an online monitor before it runs, usually within seconds. Out of more than a billion decisions that month, the monitor blocked 0.002 percent, or about one in 47,000. Meanwhile, offline monitors flag roughly 100,000 transcripts a week. Only about 50 of those reach human reviewers.
| Measurement | Anthropic’s August 2026 snapshot |
|---|---|
| AI-led R&D | Claude leads 26% of model R&D tasks; over 90% at “collaborates” or above; 0% fully autonomous |
| Agent oversight | ~30,000 agents at a time; 100% of actions monitored; 0.002% blocked; ~50 escalations to humans per week |
| Compute allocation | ~6% of AI R&D compute on safety; ~12% of AI-driven R&D compute on safety (July 13 to 20) |
The third measurement is compute. Over one week in July, about 6 percent of the compute spent on AI R&D went to safety work. For AI-driven R&D specifically, the safety share was about 12 percent. Still, Anthropic argues compute is the most verifiable input in the process. Therefore, it could become a lever if labs ever agree to slow down.
Why Anthropic is publishing this now
The post follows CEO Dario Amodei’s September essay asking the industry to slow down. “We must slow the pace at which we improve the capabilities of AI models,” he wrote. Accordingly, Anthropic says its numbers would shift if labs coordinated on pacing. It also plans to embed independent evaluators from several organizations, with access matching its own risk teams.
The timing is not accidental. OpenAI said in August it would slow its own development after a rogue agent hacked Hugging Face. Likewise, Anthropic is preparing to go public after filing confidential IPO paperwork in June. Even so, a lab disclosing how much of its research its model now runs is new. And the number is already a quarter.
