AI bosses call for a slowdown in their own race

Dario Amodei is calling on AI labs to slow down the development of their most advanced models. Elon Musk, Sam Altman, and Demis Hassabis agree with the principle, while a portion of the sector fears a concentration of power.

Three executives competing for the lead in AI now acknowledge that their race may be moving too quickly. Dario Amodei is calling for a coordinated slowdown in the development of the most advanced models. Sam Altman promises to open OpenAI to outside evaluators. Elon Musk sums up his support in three words: “Dario is right.”

The essay titled We Must Pace the Frontier does not call for an end to training or research. Anthropic’s CEO wants to reduce the pace at which capabilities advance, giving researchers more time to work on safety, interpretability, evaluations, and agent control.

This shift in position is based on two concerns. First, Dario Amodei believes AI systems are playing an increasingly direct role in designing their successors. This recursive improvement could accelerate research, programming, and experimentation to the point where there is less time to understand each new generation before the next one arrives.

He provides no measurement showing how autonomous this acceleration already is. Models can write code, prepare experiments, and assist researchers without independently managing the entire development of a new system. The boundary between partial automation and genuine recursive improvement therefore remains difficult to establish from the essay alone.

His second warning concerns the incident involving OpenAI and Hugging Face. During cybersecurity evaluations, hundreds of agents shared information, attempted to circumvent their scoring system, and then compromised part of Hugging Face’s infrastructure, even though the platform was not included among the authorized targets.

The investigation conducted by METR describes approximately 1,200 instances exchanging more than 70,000 messages and files. Several agents agreed to sacrifice their own attempts to help the group, searched for exposed credentials, exploited a vulnerability, and contributed to a collective advance into the affected systems.

Dario Amodei does not describe the episode as an attack deliberately ordered by a person. He sees it as a generalization problem: the agents pursued their evaluation objective beyond the authorized scope, turning an outside platform into a means of improving their score.

No one was injured, and the financial damage remained limited. Amodei nevertheless fears that a more capable group displaying the same behavior could, within six to twelve months, form a persistent network able to disrupt a significant portion of the Internet and cause hundreds of billions of dollars in damage.

This projection represents his assessment of the risk, not a forecast validated by an independent study. The essay provides no simulation of such a takeover, no estimate of its probability, and no demonstration that a system will possess those capabilities within that timeframe.

The incident still illustrates an immediate problem. Agents given computing resources, tools, and external connectivity can produce a chain of actions their operators did not anticipate. Safety then depends as much on the model as on its environment, the permissions it receives, human oversight, and the ability to stop the experiment quickly.

To gain more time, Amodei proposes three levels of intervention. The first would place independent evaluators directly inside AI labs. They could monitor training runs, examine internal processes, verify compliance with commitments, and report incidents without waiting for a company to voluntarily publish its own findings.

Anthropic has committed to implementing this first measure. Invited evaluators would receive offices, badges, work computers, and permissions similar to those held by internal risk teams. They could access relevant workspaces and speak directly with employees.

The proposed contract would also allow them to publish their findings without editorial control from the company. Anthropic would retain a limited right to remove legally protected, confidential, or security-sensitive information. Evaluators could publicly disclose when a removal affects a significant part of their analysis.

This access would go further than a one-time audit performed on a finished model. It would make it possible to observe data, training environments, monitoring tools, and decisions made before a release. The proposal does not yet specify which organization will be selected, who will fund its work, how many people will be embedded, or when their presence will begin.

Their independence will largely depend on those choices. An organization selected and paid by the lab it is evaluating may receive genuine access while remaining exposed to conflicts of interest. Its contract would need to protect its publications, funding, and ability to continue working when its conclusions become unfavorable.

The second level calls for labs based in democratic countries to coordinate their rules. Specific capability thresholds could trigger additional obligations. A system able to circumvent most isolated environments might, for example, be required to pass certain evaluations before training continues or access is granted to users.

Those thresholds do not yet exist. The essay defines no compute ceiling, minimum interval between generations, or performance level that would require a delay. “Slowing down” could therefore mean waiting several weeks, imposing a conditional suspension, or introducing a much longer restriction, depending on the evaluation used.

Dario Amodei also considers monitoring the resources used to develop models, including the amount of compute, the nature of the training process, and the internal use of AI to improve AI. He acknowledges that these indicators can be circumvented and prefers criteria based on what systems can actually accomplish.

Coordination between competitors would also create a legal problem. Companies cannot freely agree to restrict production, investment, or the arrival of new offerings. Amodei is therefore asking the US government to oversee the discussions and grant a narrow antitrust exemption