AI News

Anthropic Reveals How It Stopped Claude From Powering Cyberattacks and Dangerous Research

Claude Had Some Very Unwelcome Customers

Most companies publish reports to explain revenue, product updates or why a button has moved three pixels to the left. Anthropic’s latest report had a darker job: explaining how criminals, state-linked groups, surveillance operators and researchers allegedly tried to misuse Claude.

The encouraging part is that Anthropic says it found them and shut down their access.

In a wide-ranging threat intelligence report, the AI company described malicious or suspicious activity detected between December 2025 and August 2026. The cases stretched across cyberespionage, fraud, propaganda, surveillance, unauthorized model extraction and research that could potentially contribute to biological weapons.

Anthropic said it banned every account connected to the activities it identified. It also strengthened safeguards, created new detection methods and shared relevant information with government authorities and industry partners, according to the Associated Press.

Still, the report provides an unusually detailed look at how an AI laboratory can detect abuse before it disappears into the broader internet.

AI Is Moving From Adviser to Operator

The report’s central warning is not merely that criminals ask chatbots naughty questions. That would be the 2023 version of the problem.

Anthropic says some threat actors now use AI across large portions of an operation. Models can help write code, analyze information, organize campaigns and adapt when something fails. Humans increasingly supervise workflows rather than completing each technical step themselves.

In one alleged Russia-linked campaign, a group whose methods resembled those associated with Midnight Blizzard used AI during phishing, hotel Wi-Fi hijacking and WhatsApp-account takeover operations targeting Ukrainian government, military and diplomatic interests. Anthropic said the attackers used AI at nearly every stage.

The company also described a system that could notice when security software detected malware and then rewrite the code in an attempt to evade defenses. Reuters reported that Anthropic disrupted the related activity.

These are Anthropic’s findings and attributions, not independent court judgments. The Russian Embassy did not immediately respond to Reuters’ request for comment.

Even with that caveat, the pattern deserves attention. AI can compress skills that once required several specialists into a workflow available to one operator.

The Biological Research Cases Raised the Highest Alarm

Anthropic said some researchers accessed Claude from unsupported regions, concealed the purpose of their work or used intermediaries. The company could not always determine whether their intentions were legitimate. That ambiguity sits at the heart of dual-use science.

Research that helps scientists understand a virus can support vaccines and treatments. The same knowledge could potentially make a pathogen more dangerous.

One case involved assistance with a grant application concerning gain-of-function research on the chikungunya virus. According to the Associated Press, the proposed work concerned transmissibility and immune evasion and was expected to occur at a military research institute.

Anthropic blocked the request and terminated access because it could not establish a sufficiently safe explanation for the activity. The company did not publicly identify the researcher, institution or country. It also stopped short of declaring that it had uncovered a confirmed bioweapons program.

That distinction is crucial. “Could support dangerous research” is not the same as “built a biological weapon.”

Anthropic concluded that the potential consequences justified caution. It has since applied stronger restrictions to a wider range of dual-use biology queries in newer models.

Stronger Models Require Stronger Guardrails

Anthropic said most biological cases involved older systems, including Claude Opus 4 and Claude Sonnet 4.5. The company believed those models sat below the level at which they could substantially help a sophisticated scientist conduct dangerous biological work.

Newer models complicate that assurance.

Today’s systems can assist with more complex scientific tasks. They can synthesize literature, evaluate plans and maintain context across long projects. Those abilities can accelerate beneficial research, but they also increase the amount of operational help a model might provide to a dangerous user.

Anthropic therefore expanded its safeguards for newer systems, including Claude Fable 5. The company said those protections restrict a wider range of dual-use biological queries.

This illustrates an awkward truth about AI safety: yesterday’s adequate filter can become tomorrow’s decorative fence. A refusal system designed for a chatbot that explains concepts may not be strong enough for an agent capable of assisting with sophisticated research.

The response must evolve with the capability.

Scientists conducting legitimate work may occasionally hit extra friction. Anthropic’s position is that high-consequence uncertainty demands caution, especially when users disguise their identities or purpose.

Claude Became a Target, Not Just a Tool

Anthropic blocks Claude AI misuse

Some groups allegedly wanted more than Claude’s assistance. They wanted Claude’s capabilities.

Anthropic accused operators linked to several China-based AI laboratories of attempting to extract knowledge from its models through illicit distillation using outputs from a powerful system to help train another model without authorization.

The largest case described by Anthropic involved activity it attributed to Alibaba. The company said it observed more than 151 million exchanges between May and July 2026, at times approaching three million per day across more than 3,500 accounts it considered fraudulent.

Anthropic alleged that the activity aimed to improve Alibaba’s Qwen models. Reuters said Alibaba did not immediately answer a request for comment on the new claims.

Anthropic separately alleged that Moonshot and DeepSeek routed live customer conversations through Claude and used the resulting responses as training data. Some conversations reportedly contained sensitive information.

These claims require attribution because Anthropic is both the investigator and a commercial competitor. The report presents its technical conclusions; it does not substitute for independent legal findings or responses from every company named.

Still, the scale reveals a second security challenge: laboratories must prevent harmful use while stopping unauthorized efforts to copy their models.

Surveillance Became Faster and Cheaper

Anthropic also described government-linked users employing Claude to streamline surveillance.

According to Axios, cases connected to Mali, China and Iran used various Claude models to collect information, analyze social-media activity, build dossiers or select potential targets. One alleged Iranian operation used Claude while developing a malicious browser extension designed to gather identities from social networks.

The most striking change was efficiency. Tasks that previously demanded teams of analysts could reportedly be handled by a small office, an individual employee or a contractor working with AI.

Anthropic said it banned all accounts tied to the surveillance cases it uncovered and updated its safeguards. In one case, the company detected a commercial surveillance platform before it became operational.

Blocking Claude does not guarantee that a surveillance program ends. Jacob Klein told Axios that some actors moved to open-source models after enforcement created too much friction. Platform enforcement can disrupt harmful operations, but it cannot solve a political problem by itself.

AI companies are becoming early-warning systems because they may see suspicious projects while users are still building them. The opportunity is valuable. So is sharing what they learn before the same actors knock on another model’s door.

Propaganda Leaves Clues Before It Goes Public

The report also examined influence operations.

Anthropic said groups created networks of social-media accounts designed to look like ordinary users. Those accounts then amplified coordinated political messages across several regions, including Russia, Iran, Malaysia, Bangladesh and parts of Europe, Africa and South Asia.

Social platforms usually encounter an influence campaign after posts begin circulating. An AI provider may see an earlier stage: drafting personas, preparing content, translating messages and organizing the network.

That gives companies like Anthropic a potentially powerful defensive position. They can detect repeated prompts, shared infrastructure, suspicious account creation and coordinated workflows before the finished campaign reaches a large audience.

Anthropic’s report argues for using threat intelligence carefully and sharing useful patterns with other defenders. That could help social networks, security companies and governments identify campaigns faster without circulating every user conversation.

The goal is not perfect foresight. It is earlier friction.

If an influence operation requires thousands of accounts, repeated rebuilding and constant migration between tools, it becomes slower, more expensive and easier to expose.

Transparency Is Part of the Defense

Anthropic says it published the report because AI risks will grow as models become more capable. Disclosure allows competitors, governments and researchers to recognize similar behavior on their own systems.

That is the strongest positive element of the story.

Yet silence leaves every provider solving the same problem alone.

Anthropic said it shared threat indicators with authorities and industry partners. It also used the cases to build classifiers and improve detection. Its Transparency Hub describes broader enforcement and reporting practices, including millions of banned accounts and links to earlier threat-intelligence publications.

Independent scrutiny still matters. Anthropic selected the cases, controlled much of the underlying evidence and withheld identities in several sensitive examples. Outside researchers cannot fully verify every conclusion from a public summary.

Transparency therefore should be treated as a beginning, not a gold star. Useful next steps include standardized incident reports, independent audits and safe mechanisms for sharing threat indicators across companies.

Cybersecurity improved over decades because defenders learned to exchange information. AI security will need the same reflex preferably before the incident report requires its own sequel.

What Anthropic Actually Stopped

It is tempting to summarize the report with a superhero headline: Claude defeats spies, hackers and bioweapon plots before lunch.

Reality is more precise.

Anthropic says it identified and terminated access associated with the cases in its report. It disrupted the use of Claude, strengthened its own defenses and warned others. It did not claim to arrest every operator, dismantle every government program or prove that all suspicious biological work had malicious intent.

That narrower achievement still matters.

Security rarely delivers one cinematic victory. It creates layers of resistance. An account ban forces an operator to rebuild access. A classifier catches repeated tactics. Shared indicators help another platform recognize the same network. Stronger model safeguards reduce the assistance available when someone returns under a new name.

Each layer adds cost, delay and exposure.

The report also demonstrates why AI safety cannot rely only on filtering individual prompts. Sophisticated users can spread work across accounts, hide intent and combine models with external tools. Defenders must examine behavior at the system level.

Anthropic’s success, by its own account, came from connecting those patterns and acting on them. The company did not eliminate the threat. It made its platform harder to exploit and gave the rest of the industry a map of what to watch.

Claude’s Misuse Report Is a Warning With a Constructive Ending

Anthropic blocks Claude AI misuse

Anthropic’s findings show that AI-assisted cybercrime, surveillance and influence operations are no longer hypothetical. The report also shows that defenders are not standing still.

Anthropic detected suspicious activity, closed accounts, expanded biological safeguards and shared intelligence. It also published details that could help other organizations recognize similar patterns.

The next step requires collective action. AI laboratories need to exchange threat indicators without exposing innocent users’ data. Governments need rules for high-risk capabilities and credible oversight. Researchers need access to evidence that lets them test company claims. Customers need clear explanations of how providers monitor abuse.

No safeguard will work perfectly. Determined actors will switch services, use local models or invent new methods. The goal is not to make misuse mathematically impossible. It is to make dangerous activity harder, slower, costlier and more visible.

That is less dramatic than claiming the good robots defeated the bad robots. It is also how real security works.

Anthropic’s report offers a glimpse of an emerging defensive system: models spotting patterns, human investigators connecting evidence and companies acting before suspicious projects mature.

Claude did not save the world. But Anthropic says it closed several dangerous door and left the porch light on for everyone else.

Sources