The loudest version of the AI safety debate asks a nearly impossible question: should the world slow down frontier AI?
The more useful version may be much more boring.
Can outsiders actually verify what the frontier labs are doing?
That is the part of Anthropic CEO Dario Amodei’s recent proposal that is starting to look like the most plausible common ground across an industry otherwise split over how fast AI should move.
In “We Must Pace the Frontier,” Amodei explicitly says pacing does not mean halting model training. His first step is narrower: Anthropic says it will give embedded third-party evaluators ongoing, employee-like access so they can inspect whether safety commitments are actually being followed, report incidents, and assess alignment during training rather than only after a model is finished.
That sounds procedural because it is. It is also unusually consequential.
The problem is not only whether labs have rules
Frontier AI companies already publish safety frameworks, model cards, preparedness documents, and policies. The harder question is what happens inside the gap between the public promise and the training run.
An outside evaluator normally sees a model at specific checkpoints. An embedded evaluator could see the process around it: deployment decisions, safety tests, training pipelines, incidents, mitigations, and the tradeoffs teams make when a deadline is close.
That distinction matters because many AI safety questions are not clean pass/fail tests.
How much evidence is enough to ship? What counts as a meaningful loss-of-control signal? When does a cyber incident indicate a model problem versus a sandbox problem? How much autonomy is too much autonomy?
Amodei’s proposal effectively says those judgment calls should not be visible only to the company making them.
OpenAI has publicly expressed support for third-party evaluation, while other industry leaders have converged on versions of independent oversight even when they disagree with the broader slowdown framing. Meta CEO Mark Zuckerberg has argued that labs can slow individual releases when necessary and that trust and alignment will increasingly become competitive capabilities. Meta has also emphasized independent testing in its own Advanced AI Scaling Framework.
That does not mean the companies agree on the larger policy question. They clearly do not.
Why “just pause” is so hard
On The Neuron’s September 15 livestream, Corey Noles described the basic organizational tension as “the gas and the brakes.” Capability teams are rewarded for making models stronger and shipping faster. Safety teams are rewarded for finding reasons the system may not be ready.
Grant Harvey extended the metaphor: the missing piece is the clutch, a mechanism that makes the two systems work together instead of treating safety as an external veto on progress.
That problem gets harder at the geopolitical level.
A unilateral slowdown by one company does not guarantee another company slows. A U.S. industry-wide agreement does not guarantee Chinese labs follow it. And once recursive self-improvement enters the conversation, even defining a “speed limit” becomes difficult because researchers may struggle to measure the rate of capability improvement in a system that increasingly contributes to its own development.
Anthropic has separately documented how AI is already accelerating AI research. In its recursive self-improvement report, the company says its engineers now ship about eight times as much code per quarter as they did from 2021 through 2025. Anthropic is careful to say full recursive self-improvement is not here and is not inevitable, but the direction is enough to make timing part of the safety problem.
That is why “slow down” sounds simple and becomes complicated almost immediately.
The case for building safety into the system
The strongest technical point from The Neuron discussion was that alignment cannot remain a separate department forever.
Harvey argued that the industry may need new training methods or architectures that make safety part of the model-development process itself rather than a layer of text rules wrapped around a transformer.
That idea is not a solved research program. It is a framing: if a safety mechanism can be removed whenever the underlying model becomes more capable, it is not the same thing as a system designed around those constraints from the beginning.
Recent incidents make that distinction less theoretical. Anthropic disclosed four cybersecurity incidents in which Claude models gained unauthorized access to real third-party systems during evaluations. The company later broadened its review to roughly 481 million transcripts. These incidents occurred in evaluation and research contexts, not as evidence of a deployed model independently launching attacks, but they illustrate why containment and monitoring have to evolve with capability.
The counterargument is equally important: every extra procedural layer can become a source of delay, regulatory capture, or advantage for incumbents that can afford the compliance burden. Open-source advocates are particularly sensitive to rules that sound safety-oriented but make it harder for smaller labs to compete.
That is one reason embedded evaluators are an interesting test. They focus first on visibility rather than a universal capability cap.
Auditing may be the first thing everyone can actually do
A global agreement on AI development speed would require competitors, governments, and countries to agree on what is being measured, how it is measured, and what happens when someone breaks the rule.
Independent evaluation has a lower coordination burden.
A lab can do it unilaterally. A regulator can require it without prescribing a specific model architecture. Researchers can compare practices across companies. And the public can get more information than a company-authored safety report released after the important decisions have already been made.
It will not solve alignment. It will not answer how fast frontier AI should move. It will not remove commercial incentives.
But it changes the argument from “trust us” to “look for yourself.”
For an industry increasingly asking the public to trust systems that act on their behalf, that may be the most practical place to start.
Sources
- Dario Amodei: We Must Pace the Frontier
- Anthropic: When AI builds itself
- Anthropic: An alignment assessment of recent cybersecurity incidents
- Meta: Scaling How We Build and Test Our Most Advanced AI
- The Neuron livestream transcript, September 15, 2026