Espionage, surveillance, fraud: how Claude was hijacked
Anthropic documents malicious uses of Claude in espionage, surveillance, fraud, and weaponry, with operations largely delegated to agents.
Three hours to go from a stolen developer token to full administrative control of a cloud environment. Dozens of victims handled in parallel. Malicious tools automatically modified as soon as a security product detects them. In its latest report, Anthropic describes a development that is less spectacular than a revolutionary new vulnerability, but potentially more consequential: the automation of a growing share of the work required to carry out an attack.
Published on September 10, 2026, the document covers operations detected and disrupted between December 2025 and August 2026. Its 154 pages examine seven areas: cyberattacks, influence operations, surveillance, fraud, misuse of biological research, conventional weapons development, and unauthorized extraction of Claude’s capabilities.
The actors identified include groups suspected of working for governments, financially motivated cybercriminals, surveillance providers, propaganda institutions, and political operators. They used Claude Haiku, Sonnet, and Opus. Anthropic says it observed the use of its Fable or Mythos models in only one case involving capability extraction.
The report does not present these incidents as representative of all malicious activity. Anthropic selected the cases it considered the most notable or novel. Its observations come mainly from its own services, detection systems, and attribution work. They therefore do not constitute a comprehensive account of AI-related threats or an independent audit of the cases presented.
The company uses internal identifiers, known as Generative Threat Groups or GTGs, to designate the operators it tracks. It also attempts to assess the contribution of AI across three dimensions: speed, scale, and depth. This measurement remains partly qualitative, since it is rarely possible to observe the same campaign conducted under identical conditions without automated assistance.
The main conclusion is less about the novelty of the techniques than their cost. The attacks described often rely on familiar methods: stolen credentials, exposed services, unpatched devices, SQL injection, phishing, and cloud account takeovers. What has changed, according to Anthropic, is how much work a single operator can perform.
Reconnaissance, tool development, exploitation, lateral movement within a network, and the organization of stolen data can now be delegated to several agents running simultaneously. In the most advanced cases, the human provides a general objective, supplies initial access, and reviews the results, while the system selects and carries out some of the intermediate steps.
This autonomy does not mean Claude independently decides to attack an organization. Humans generally retain control over target selection, monetization, and final review. Anthropic also emphasizes that autonomy and severity are not the same thing: automating an operation can increase its speed and scale without determining the extent of the resulting damage.
The most detailed case concerns GTG-20006, a Russian-speaking actor whose methods and targets Anthropic considers consistent with those of Midnight Blizzard. Microsoft uses that name to track a group linked by the US and UK governments to Russia’s Foreign Intelligence Service.
The operator targeted government agencies, military intelligence services, diplomatic missions, defense-related manufacturers, and individuals connected to US foreign policy. More than 20 organizations were reportedly involved in its planning, reconnaissance, or active operations.
Claude was used across almost the entire chain: researching targets, acquiring infrastructure, registering domains, conducting phishing campaigns, maintaining access, moving between systems, and extracting data. Scheduled tasks renewed stolen access tokens and retrieved content from cloud storage without continuous human involvement.
The actor had also created a loop to monitor whether its malware was being detected. When a security product flagged a component, agents modified the code, rebuilt the file, and repeated the process until they produced a version that was no longer detected. This capability reduces the amount of time a static signature can slow down a campaign.
One of these operations relied on hotel Wi-Fi networks. The attacker allegedly compromised at least three hospitality providers, then modified their DNS settings to redirect travelers to infrastructure under the attacker’s control. Visitors received fake update or repair instructions designed to persuade them to install malware.
This part of the case has some external corroboration. Microsoft describes the operation as CaptiveCrunch, attributes it to a Midnight Blizzard subgroup, and says it observed significant use of AI. Microsoft also credits Anthropic and OpenAI for contributing to the investigation.
The same actor allegedly took over WhatsApp accounts by linking them to headless browsers acting as companion devices. Its configuration suppressed read receipts while conversations were exported. At least two former senior Ukrainian officials were reportedly targeted.
The information sought was not limited to diplomatic communications. The group stole the complete software development kit for a drone vision system, then reconstructed its architecture, components, suppliers, and details of an unannounced product.
In another intrusion, the actor allegedly obtained more than 300,000 national identity records from a North African government technology authority, along with registry data covering more than 500,000 companies. Anthropic also reports the extraction of email records from at least eight organizations.
A second group of cases concerns operators Anthropic associates with the ShinyHunters ecosystem. Their activity illustrates a form of goal-directed hacking: the human provides access or identifies a target, then allows agents to explore the environment, write and execute scripts, search for valuable information, and produce a summary.
A network of ten servers