Anthropic CEO Dario Amodei has announced the company is committing to a new safety measure—giving independent evaluators permanent, employee-level access inside the company—as part of a broader three-step plan he says is needed to slow the pace of AI development.
In an essay published on Saturday, Amodei laid out a plan aimed at “pacing the frontier,” or slowing AI development. First, it calls for every frontier AI company to give independent evaluators permanent, employee-level access to verify safety practices and report incidents; second, companies in democratic countries to agree on common safety standards that limit the rate of unchecked progress; and third, democratic governments to attempt coordination with authoritarian states, starting with agreements that are in everyone’s interest, such as a ban on using AI to develop biological weapons.
Amodei has long cautioned about the pace of AI development, but he says two recent shifts have increased the need for urgent safeguards on the technology. Models, he said, are increasingly able to build their successors, which is accelerating progress further. The industry has also seen a string of safety incidents, he added, including within Anthropic itself. He believes even a couple of years of pacing model development would give researchers time to reduce the risk of something going wrong, and calls on the industry to do so now.
Anthropic is committing to the first step unilaterally, with immediate effect. Independent evaluators will work inside the company permanently, Amodei said, with the same access as its own risk-assessment teams and the right to publish their findings without Anthropic’s editorial control.
Anthropic has found itself at the center of a media storm this week after researcher Jacob Coxon publicly resigned from the lab, warning that AI companies were gambling with people’s lives. In his resignation post on X, Coxon, who spent three years doing pretraining research at both OpenAI and Anthropic, wrote: “Neither company is acting responsibly. They are racing straight to self-improving superintelligence and gambling with our lives,” adding that “these will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.”
Several other current Anthropic employees supported the post, sharing similar fears about AI — most notably safety lead Evan Hubinger, who wrote: “Jacob is correct here, we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”
The resignation lands amid a string of unsettling AI agent incidents that have rattled the industry and many in Washington. In July, OpenAI disclosed a breach in which its agents autonomously hacked the open-source repository Hugging Face. Later, researchers found OpenAI had also kept quiet about an earlier, separate episode in which rogue agents hijacked a German programming wiki, making more than 15,000 edits and turning it into a message board where agents swapped tips for evading restrictions and detection.
The incidents have fueled public and regulatory concern. U.S. politicians are now discussing urgent regulation of AI. Anthropic, for its part, was founded on the premise that safe AI development should come before speed—a mission that some former workers say has come under strain due to intense competitive pressure from OpenAI.
This story was originally featured on Fortune.com

1 hour ago
1


