If you want to regulate a race, you first need some way to measure speed.
Anthropic thinks it has one.
On Thursday, the company published a new framework for measuring how quickly AI development is accelerating inside frontier labs—the small group of companies building the world’s most capable AI systems. Its headline number is already getting plenty of attention: Claude now “leads” 26% of the AI research and development work Anthropic measured. More than 90% involves Claude at least as a collaborator. Anthropic’s new measurement framework
Those numbers sound a little sci-fi. But the more consequential part of Anthropic’s announcement may be the measuring system itself.
The company has proposed three indicators for tracking the frontier AI race: how much AI is involved in building newer AI, how closely internal AI agents are monitored, and how much computing power is devoted to safety work.
In other words, Anthropic is not only telling the public how fast it thinks AI development is moving.
It is proposing the dashboard everyone else might eventually use to decide whether it is moving too fast.
And that creates a surprisingly important question: Who gets to design the speedometer before anyone agrees on the speed limit?
- First, Claude is not autonomously building Claude
- Anthropic had to invent a way to measure the work
- The second metric asks whether humans can still see what the agents are doing
- Then there is the 6% number
- A dashboard without triggers can only tell you how fast you crashed
- Whoever defines the metric gets a head start on defining the debate
First, Claude is not autonomously building Claude
That distinction matters because “Claude leads 26% of Anthropic’s R&D” is very easy to turn into “Claude is building its own successor.”
Anthropic’s actual definition is narrower.
The company adapted an automation scale developed by Epoch AI. At the bottom, AL0 means no AI involvement. At AL3, AI “collaborates,” completing large pieces of work while a human gives close direction. At AL4, AI “leads”: a human can provide a high-level assignment, and the AI completes most of it while the human supervises.
AL5 is the big one: AI operates completely autonomously, with no human in the loop.
Anthropic says none of the work it measured reached AL5. That distinction is explicit in Anthropic’s framework.
That makes the 26% figure significant without making it Skynet.
Claude apparently performs substantial chunks of the work Anthropic uses to develop future models. Humans still set goals, supervise the process, and retain authority over what actually gets deployed.
The trend is still worth watching because it creates a potential feedback loop: better AI helps researchers build better AI, which could then become better at helping researchers build the next version.
Anthropic describes fully autonomous versions of that loop as “recursive self improvement.” The company does not claim it has reached that point.
Its argument is that we should probably have instruments for measuring the distance before we get there.
Anthropic had to invent a way to measure the work
Here is where things get interesting.
There is no universally accepted unit called “percentage of AI research done by AI.”
Anthropic had to construct one.
For each week in July, the company says it sampled 20% of employees in departments involved in model R&D. Claude then reviewed internal records such as Slack messages and documentation to identify what work those employees had performed.
That produced roughly 15,000 individual tasks.
Anthropic used Claude again to organize those tasks into a hierarchy of 542 categories, then investigated how much AI currently contributes to each type of work. The company freezes that basket of tasks so future measurements can be compared against the same workload. Anthropic explains the methodology here.
This is a much more serious methodology than somebody asking engineers, “So, uh, how much Claude do you use?”
But it also exposes the core tension.
Anthropic chooses the tasks. Anthropic possesses the records. Claude helps classify the work. Anthropic determines how categories are weighted. Anthropic then reports the resulting number.
The company acknowledges that problem itself, noting that using its own models to evaluate its systems risks producing shared errors. It proposes third-party verification and eventually common methodologies across AI developers.
That is why 26% should be read as a prototype measurement, not a universal scientific constant.
There is a strong case for publishing an imperfect prototype anyway. The alternative is not a perfect neutral benchmark waiting on the shelf; much of the relevant development data is currently private, incompatible, or invisible outside individual labs. A common measurement that outsiders can inspect, challenge, and improve could still be a major governance upgrade.
What Anthropic is really doing is proposing a definition of what AI-assisted AI development looks like.
If other labs adopt something similar—and governments eventually require it—that definition could become much more powerful than the number attached to it today.
The second metric asks whether humans can still see what the agents are doing
Anthropic’s dashboard also tracks the AI agents already working inside the company.
And there are a lot of them.
Anthropic says roughly 30,000 agents were performing research and engineering work at any given time on its most-used internal agent platform in August.
The company says every action from those agents on that platform passes through an online monitor before execution and is also ingested by an offline monitoring system afterward.
The online monitor blocked 0.002% of more than one billion decisions examined during August—roughly one in every 47,000. Its offline system, meanwhile, flags around 100,000 transcripts each week for additional automated review, with roughly 50 high-priority cases reaching humans. Anthropic reports those monitoring figures in the same disclosure.
Those are impressively specific numbers.
They are also a good example of why measurement and meaning are two different problems.
A 100% monitoring rate tells us that every covered action passes through a monitoring system. It does not tell us that the system catches 100% of dangerous behavior.
Likewise, one blocked action in every 47,000 sounds reassuring until you ask the question the percentage cannot answer: How many dangerous actions did the monitor miss?
There is some outside scrutiny here. METR says one of its researchers spent three weeks red-teaming part of Anthropic’s internal agent-monitoring system earlier this year, discovering several previously unknown vulnerabilities, some of which were subsequently patched. METR published an account of that work.
That is the kind of independent testing these numbers need.
Because “everything is monitored” and “everything dangerous is detected” are very different claims.
Then there is the 6% number
Anthropic’s third metric follows the money—or, more precisely, the GPUs.
During one week in July, the company estimates that about 6% of the compute used for AI R&D went toward safety work. Among compute specifically supporting AI-driven AI research, the figure was about 12%. Anthropic describes those figures as a limited snapshot.
The obvious temptation is to look at 6% and decide whether that sounds high or low.
There is no established answer.
Anthropic itself warns that compute is an imperfect proxy for safety investment. Some safety research requires lots of human experimentation and relatively little computing power. The company also describes its classification as deliberately conservative and says some safeguards are excluded from the figure.
So why publish the percentage at all?
Because trends can become useful even when absolute numbers are messy.
Imagine Anthropic reports this every quarter. Then imagine OpenAI, Google DeepMind, xAI, and other frontier labs use comparable definitions.
Suddenly outsiders could ask whether the proportion of resources devoted to safety is rising as AI performs more of the development work—or whether capability development is accelerating while safeguards lag behind.
That is potentially useful information.
But it immediately raises another question: What happens if the numbers look bad?
A dashboard without triggers can only tell you how fast you crashed
This is the unresolved piece.
Anthropic CEO Dario Amodei recently called for the industry to “pace” frontier AI development, arguing for measures including independent evaluators embedded inside AI labs. Amodei’s “pace the frontier” proposal has already pushed the debate toward a more practical question: what can outsiders actually inspect?
Anthropic says those evaluators would receive access to internal systems, processes, and data comparable to the access available to internal risk teams. The Neuron recently examined why auditability is becoming central to the slowdown debate.
Its new metrics give those evaluators something concrete to inspect.
That adds a more operational layer to the AI-safety debate. Instead of measuring only what models can do, Anthropic is proposing measurements for what happens inside the development process itself.
Here, the questions become more tangible.
How much research is AI performing?
How much of its behavior is monitored?
How quickly are problems reviewed?
How much compute is devoted to safety?
Those are questions an auditor can theoretically check.
The Council on Foreign Relations has argued that voluntary commitments still need independent oversight and enforceable standards, including protections that make outside evaluators genuinely independent. CFR’s analysis of independent AI oversight makes the same basic distinction: visibility helps, but visibility is not enforcement.
Because an even better measurement system does not automatically produce governance.
Suppose Claude eventually leads 60% of measured R&D instead of 26%.
Or the monitoring system starts escalating dramatically more incidents.
Or safety compute falls while AI-led research rises.
Who has the authority to say: stop?
Anthropic?
An independent evaluator?
A regulator?
A cloud provider controlling access to enormous amounts of compute?
Right now, Anthropic’s framework mostly tells us what could be measured. It does not create binding consequences when those measurements cross an uncomfortable threshold. That enforcement problem is also why the question of whether AI companies can afford to slow their own progress remains unresolved.
Whoever defines the metric gets a head start on defining the debate
There is another reason this matters.
Company-defined measurements can become consequential if other labs, evaluators, or governments begin adopting them.
Once companies, researchers, journalists, investors, and regulators start quoting the same number, the assumptions behind that number become easy to forget.
That makes Anthropic’s decision to publish first consequential regardless of its intent.
The company is proposing what “AI-led R&D” means. It is proposing what agent oversight should look like. It is proposing one way to classify safety spending.
Anthropic openly says it wants common methodologies and third-party verification. That could give governments far more visibility into frontier labs than they have today.
But the design of a common standard still embeds choices about what gets measured, how it is weighted, and what counts as success.
Should an hour spent by AI automating routine engineering count the same way as an AI system choosing a major research direction?
Should safety be measured through compute, researcher time, dollars, outcomes, incidents prevented—or some combination?
What level of monitoring performance counts as adequate?
The answers could eventually determine whether a lab appears responsible, reckless, fast-moving, or safely constrained.
Publishing first matters because the first workable measurement system can become the reference point everyone else has to respond to.
That is why Anthropic’s most important contribution here may have little to do with whether Claude currently leads 26% of anything.
For years, the AI debate has wrestled with a problem hiding behind all the arguments about whether development should slow down: outsiders cannot easily see what is happening inside the labs.
Anthropic has now put three gauges on the dashboard.
Now comes the harder part: deciding who calibrates them, where the red line belongs, and what happens when someone crosses it.