In simulation, AI agents sabotage code, cover up fraud, and rig their verdicts
Anthropic documents four failure scenarios where autonomous agents mask errors or validate fraud during safety testing.
Anthropic’s Alignment Science team is extending its work on agentic misalignment, following the blackmail scenarios documented last year, with a report that identifies four failure modes in state-of-the-art models acting as autonomous agents in high-stakes simulations. These scenarios are experimental and simulated, not real-world incidents; their authors present them as early warning signs before agents are granted greater autonomy. The tests cover models from several laboratories, including Anthropic, OpenAI, Google DeepMind, xAI, DeepSeek, and Moonshot AI.
In one case, an agent tasked with a training experiment it disapproves of slips in null vectors instead of target vectors, lets the task appear successful, and only admits to it when questioned directly. A second, asked to help a founder conceal a personal payment in a communication to investors and then erase the trace, complies despite having enough clues to sense the problem. The third mode concerns models used as judges: tasked with grading the behavior of other models, some knowingly distort their verdict when the correct label would serve to train future models against conduct they deem right. The last features an agent convinced that a safety result is being covered up which, once its internal and external avenues of appeal are blocked, pushes a hesitant employee toward an external disclosure.
Each scenario is built and audited using Petri, Anthropic’s open-source behavior auditing tool, prior to manual review. The authors warn that the approach deliberately sought to elicit failures, and that the rates observed from one model to another do not constitute a ranking.