Anthropic: Claude now leads 26% of R&D work
Serge Bulaev
Anthropic says that its AI, Claude, now leads 26% of its research and development work, up from less than 1% in February 2026. This may suggest a move toward more agent-led research in the company, though humans still review important steps. Industry experts note that Anthropic tracks this progress in its new R&D Automation Index, and that no work is yet fully unsupervised. Questions remain about how Anthropic checks the safety and accuracy of these AI outputs, and how much computing power is spent on safety as agent use grows. More information may be shared when Anthropic releases its next Automation Index update.

Anthropic reports its AI, Claude, now leads 26% of the company's R&D work, a stark increase from under 1% in early 2026. This milestone, reported by Reuters and Bloomberg, signals a major shift toward AI systems taking operational leadership in their own development.
According to industry reports, the company also runs a significant number of AI agents simultaneously, which make many autonomous decisions monthly. With monitoring systems blocking only a small fraction of actions, the announcement highlights the growing scale and reliability of autonomous AI. Below are five key questions this development raises for the AI industry.
What does "leads" actually mean in Anthropic's R&D context?
In Anthropic's R&D, "leads" signifies that the Claude AI can complete a task from start to finish based on a single high-level prompt. While a human researcher still reviews the final output, this moves beyond simple assistance to a model of delegated, supervised autonomy.
Anthropic defines "leads" as the ability for Claude to execute a task end-to-end from a single, high-level prompt, all while under human supervision. This marks a crucial evolution from assistive AI, where humans micromanage steps, to delegated autonomy, where the AI handles the execution. The company tracks this progress with an internal "R&D Automation Index," which shows that while a significant portion of R&D tasks involve substantial AI collaboration ("Level 3 or higher"), no work has yet reached fully unsupervised "Level 5" autonomy. Humans retain ultimate accountability, but their role is shifting from direct control to high-level oversight.
How does Anthropic's internal AI agent scale compare across the industry?
The scale of concurrent AI agents represents one of the largest internal deployments publicly disclosed. For context, rival lab OpenAI reported in September 2026 that its researchers log 3.1 agent-workdays for every human workday. Both labs indicate they have surpassed a key benchmark, creating "automated research interns" that perform significant independent work rather than just assisting researchers. This industry-wide trend shows multi-agent systems evolving from experimental projects into core production infrastructure for R&D workflows.
What governance and safety measures accompany this scale of autonomous operation?
With this massive scale comes significant governance challenges. Anthropic revealed that a significant portion of its total AI R&D compute is allocated to safety, with an even higher percentage for compute used specifically for AI-on-AI development. This highlights a potential tension between the rapid scaling of agentic capabilities and the resources dedicated to ensuring their safety. While automated monitoring systems intervene in only a small fraction of actions, the sheer volume of monthly decisions means many events still require human review. Managing large numbers of agents requires new frameworks for orchestration and monitoring, as enterprises widely report struggling to implement effective guardrails for trust, accuracy, and security at scale.
What ethical and accountability concerns arise when AI leads substantial R&D work?
Delegating R&D leadership to AI raises profound ethical and accountability questions. As a Stanford HAI industry report notes, AI risk is not just about isolated failures but also accumulates through overreliance and the erosion of human judgment. Key concerns include:
- Accountability Diffusion: When an AI suggests experiments and writes code, it becomes harder to trace responsibility, even though humans must remain accountable for the outcomes.
- Automation Bias: Researchers may over-trust the fluent and confident outputs generated by AI, particularly when working under tight deadlines.
- Verifiability Gaps: LLMs often lack robust source verification, a critical flaw in evidence-based research environments.
- Homogenization of Research: Over-reliance on a few dominant AI models could narrow creative problem-solving and reduce intellectual diversity.
The scientific principle that LLMs cannot be co-authors because they cannot be held accountable extends directly to these new leadership roles in R&D.
What does this signal about competitive dynamics and regulatory attention?
Anthropic's announcement, which follows OpenAI's similar transparency on internal agent use, confirms that operational AI autonomy is now a competitive necessity. The ability to safely delegate R&D to AI agents is a direct path to accelerating development cycles. This rapid progress is also attracting regulatory scrutiny, shifting the focus from technical capability ("can the model do it?") to governance ("who is responsible?"). Anthropic's internal metrics show 26% AI-led R&D at Anthropic, providing insight into how AI systems are increasingly taking on leadership roles in their own development processes.