Anthropic CEO Dario Amodei is urging the AI industry to slow the pace of frontier-model development enough for safety measures to keep up. His proposal does not call for a halt to AI training, but for coordinated safeguards that would constrain uncontrolled progress while preserving competition.
Anthropic says it will begin with a measure it can adopt on its own: permanently embedded independent evaluators with access close to that available to employees. These assessors would check compliance with safety rules, report incidents and examine training processes.
Amodei’s proposal has found support among rivals. OpenAI CEO Sam Altman has said OpenAI agrees with the idea of independent evaluators with internal access and will introduce a similar measure, while Elon Musk has also backed slowing AI development.
A three-part safety framework
- Embed independent evaluators at companies developing frontier AI models, with ongoing access to internal safety and training work.
- Create common standards among AI companies in democratic countries, including limits on uncontrolled advances. Amodei says governments would need to enable some coordination between competitors.
- Pursue global coordination, including with authoritarian states such as China, where compliance with agreements can be meaningfully verified.
Amodei said two developments have raised his assessment of the risk. One is recursive self-improvement, in which AI models contribute directly to developing later generations of models; Anthropic is already testing this through research agents. The other is an OpenAI-Hugging Face incident in which AI agents carried out security-related actions beyond a test’s intended objective.
The concern is that increasingly capable systems that remain poorly aligned could cause far more serious harm. Yet Amodei also argues that the United States and its allies should retain their lead over China, making coordinated action central to the proposal. The unresolved challenge is whether competing companies and governments can agree on enforceable limits without allowing a rival to continue accelerating alone.







