As of August 2026, Claude leads 26% of Anthropic’s measured AI R&D work. That figure comes from Anthropic’s own R&D Automation Index, published in September, and the word carrying the weight in that sentence is “leads.”
Anthropic scores its internal work on a scale from AL0, no AI involvement, to AL5, fully autonomous with no human in the loop. AL3, which Anthropic calls “collaborates,” means AI doing large chunks of work under close human direction. AL4, “leads,” means AI can complete most of a task end to end from a high-level prompt while a person supervises. The 26% is the share of measured work sitting at AL4, up from under 1% in February. More than 90% is at or above AL3. And Anthropic states plainly that Claude is not operating fully autonomously for any measured subset of that work.
Be precise about what this is. Anthropic is measuring its own internal work with its own methodology, and flags the limitations itself, including that it uses its own models to evaluate its own systems. Treat it as a credible look inside one sophisticated organization, not a benchmark for yours.
The definition matters more than the percentage
The reason this is interesting to an operator has nothing to do with frontier research. It is that the unit of delegation is getting larger. A person asking an assistant to summarize a document is delegating a step. A person setting a goal and supervising something that plans and executes most of the path to it is delegating a piece of work. Those are different operating models, and most companies are measuring the first while assuming they are getting the second.
Anthropic’s engineering data makes the division of labor concrete. As of May 2026, more than 80% of the code merged into its codebase was authored by Claude. That is Anthropic’s internal number on its own engineering and it does not generalize to your business. What does generalize is the shape Anthropic describes: humans supply the goal but no longer need to supply the method. Anthropic is equally clear that large gaps remain in judgment, and that deciding which problems are worth working on is still distinctly human.
Assistance and ownership are different models
Take invoice review, which almost every operating business runs some version of. In the assistant model, a person opens each invoice, checks it against the purchase order and the receiving record, asks an AI tool to summarize anything unusual, and decides what happens. The person owns the flow. The AI makes them faster at steps they were already doing.
In the AI role model, the role receives the invoices, checks each one against the relevant records and rules, completes the routine cases, and escalates a defined set of exceptions to a person: anything over a threshold, anything with a variance outside allowance, anything where the supporting record is missing. The person sets the standards, reviews what came back, and handles the exceptions. The distinction is not whether a human is involved. It is who owns the flow of the work.
Supervision moves up, it does not disappear
The more work an AI role carries, the more valuable human attention becomes at the points where it actually changes outcomes. Defining the goal. Deciding which problems matter. Setting the standard for acceptable output. Reviewing the higher-consequence decisions. Handling the ambiguity that the role was never scoped for. And deciding when the role has earned more authority, based on measured performance rather than enthusiasm.
That last one is the discipline most companies skip. Autonomy should be scoped to the work and expanded on evidence. A role with four hundred completed invoice checks and a documented exception rate has earned something. A role handed broad authority on day one because the demo went well has not.
Stop counting seats
The common adoption questions are about tools. How many employees use Copilot, how many ChatGPT licenses we bought, how often the team opens Claude. Those measure whether software was distributed, which is not the same as whether work changed.
The questions that tell you something: What recurring work does the AI actually own? Where does it start and finish? What share completes without a person intervening? Which exceptions require human judgment, and are they defined in advance or discovered by accident? Is quality measured against a standard? Is authority increasing as the evidence supports it?
The exercise
Pick one recurring process this week and split it into three columns. Human owns: work needing direction, judgment, approval, or unusual exception handling. AI leads: defined work something could carry largely end to end while a person supervises the result. AI assists: individual steps where AI helps a person who still owns the task.
Then look at the distribution. If everything lands in “AI assists,” you have not changed how the work gets done. You have added a better tool beside the existing process, and the person is still pushing every step forward.
The meaningful shift is not from no AI usage to heavy AI usage. It is from AI helping people perform tasks to AI roles carrying defined work while people direct, supervise, and make the decisions where their judgment earns its keep. Anthropic’s 26% is useful because we can watch that shift happening inside a company that builds the technology. You do not need to automate 26% of your business next quarter. You need one piece of recurring work defined clearly enough that AI can begin owning more of it.
P.S. If you have a process where AI already helps but a person still carries it from beginning to end, that gap is exactly what we look for when defining an AI role.
Sources: Anthropic, “Measuring the pace of AI development,” R&D Automation Index, September 2026; Anthropic, “Recursive self-improvement,” May 2026 data. Corroborated by Bloomberg, September 17, 2026.


