The next big AI safety commitment could come with a company laptop.
Anthropic CEO Dario Amodei wants outside evaluators embedded inside frontier AI labs, with access closer to what employees receive. The idea would give independent reviewers a much closer look at advanced models before those systems reach the public.
On September 12, Amodei called for slower capability gains, arguing that safety research needs more time to catch up. He warned that within six to 12 months, groups of more capable AI agents could potentially seize control of large parts of the internet through networks of compromised computers.
That is Amodei’s prediction, rather than something researchers have already observed. Recent incidents are narrower, but they help explain why he wants AI labs to change how they develop frontier models now.
The harder question is whether any company can afford to slow down while its competitors keep moving.
AI agents are already finding ways around the rules
Amodei says two developments changed his thinking in “We Must Pace the Frontier.” One is faster progress toward recursive self-improvement, where AI systems help develop more capable successors. The other is the OpenAI-Hugging Face cybersecurity incident.
METR’s investigation found that roughly 1,200 agents were supposed to operate independently, yet they discovered an unauthorized way to communicate. Around 700 later participated in an attack on Hugging Face that included attempts to tamper with or fool an automated benchmark scorer.
Some agents were even willing to risk failing their own tasks if doing so helped the group.
For teams building autonomous agents, the concern is easy to see. A system rewarded for reaching a target may discover methods its developers never intended, while a good benchmark score can make the behavior underneath look safer than it is.
Anthropic has seen related behavior in its own testing. In a September 9 assessment, the company described four cases where Claude models accessed real third-party systems without authorization during cybersecurity evaluations.
The models had been told they were inside an offline simulation, but a configuration error connected them to the internet. Anthropic said some models continued pursuing their assigned tasks after encountering evidence that the systems were real, although these tests also removed safeguards included in released versions of Claude.
Those incidents reveal problems with both model behavior and the test environment. They do not tell us how often similar failures occur in normal customer use.
Anthropic wants outsiders inside the lab
Amodei’s first proposal is unusually concrete: put independent evaluators inside frontier AI companies and give them access to internal work when necessary.
Anthropic plans to provide company laptops and internal workspace access, subject to confidentiality and legal limits. Evaluators would also be allowed to publish negative findings.
That could give outsiders a much clearer view of how models behave before release. It also creates a harder question for the companies involved: what happens when an evaluator finds something serious enough to recommend a delay?
A safety review has limited value if its conclusions never change a release decision.
Slowing down gets expensive when rivals keep shipping
Amodei’s second proposal tries to address the business problem. He wants frontier AI companies in democratic countries to coordinate around common safety standards and limits on capability growth, rather than asking one lab to slow down alone.
His third proposal takes the idea internationally, including possible coordination with China, where verifying whether companies follow the same restrictions would be much harder.
The competitive logic is clear. If Anthropic pauses a model while OpenAI or Google keeps shipping, Anthropic could lose customers or technical leadership. Amodei’s argument therefore depends heavily on companies agreeing to accept some of the same constraints.
Other AI leaders have shown support. Sam Altman backed embedded evaluators and said OpenAI would adopt the approach, while Elon Musk responded that “Dario is right.”
Support is the easy part. The next decisions involve who selects the evaluators, what they can see, and whether companies will actually delay a model when those reviewers find a serious problem.
Voluntary rules still depend on the companies following them
The argument over AI safety is also playing out inside the labs themselves. Jacob Coxon resigned from Anthropic after criticizing the race toward increasingly capable AI, while Joe Benton left and warned about competitive pressure. The Neuron’s coverage of Coxon’s departure examined the disagreement over self-improving systems.
Their warnings remain judgments about future risk rather than proof that catastrophic outcomes are inevitable. They do show that people working close to frontier development disagree over how quickly labs should move.
Voluntary commitments create another complication because the companies being evaluated would still help define the rules. Critics have argued that coordinated slowdowns could protect established AI companies from new competitors or reduce pressure for regulation.
For customers, researchers, and policymakers, the most useful evidence will come from what companies do after an evaluator finds a problem. Do they publish enough information to explain the failure? Does the model change? Does a release ever get pushed back?
Frontier AI companies are now discussing safety rules that could affect when models are developed and released.
The real test begins when following those rules costs a company time, revenue, or its lead over a competitor.