Anthropic advocates for state power to block dangerous AI models

Anthropic urges governments to legally block dangerous AI models trained beyond 10²⁵ FLOPs, citing risks found by its Claude Mythos Preview model.

With a post by Dario Amodei titled Policy on the AI Exponential and two public policy frameworks, Anthropic formalizes a turning point: transparency is no longer enough; binding regulation is needed.

The initial observation is captured by an image: that of Treebeard, the thinking tree from The Lord of the Rings, too slow in the face of urgency, a metaphor for a political apparatus ill-suited to the exponential pace of AI. The first framework, dedicated to catastrophic risks, proposes giving governments the legal authority to block or deter the deployment of a model deemed dangerous, accompanied by civil penalties indexed to global annual revenue. The mechanism would only target models trained beyond 10²⁵ FLOPs, developed by companies exceeding $500 million in AI revenue or $1 billion in R&D spending, and would cover four risks: biological, cyber, loss of control, and AI R&D automation. Mandatory tests by independent evaluators, public risk reports, and securing model weights complete the edifice, following the model of agencies like the FAA for aviation.

Amodei relies notably on Claude Mythos Preview, which identified thousands of critical vulnerabilities in major operating systems and browsers, to assert that frontier models have become tools of strategic importance.

The second framework addresses the economic aspect: statistical tracking of job displacement, wage insurance, tax incentives for retention, and, in the longer term, universal basic income-type mechanisms funded by the generated growth. Anthropic specifies that the text is designed for Washington but opposes any federal preemption of state laws that would not be at least as demanding as its proposal.