Anthropic has revealed a list of four incidents involving its AI bot Claude, including one previously unreported, that could be classed as criminal acts.
A new report, titled ‘An alignment assessment of recent cyber security incidents’, detailed how Anthropic’s AI models broke into third-party systems, carried out cyber attacks, and uploaded malicious software during testing.
Three of the incidents were reported on 30 July in a review of Claude’s cyber security transcripts, but a fourth hack from January 2026 went unnoticed until now.
All four incidents involved different versions of Claude, trained months apart, including its most powerful Claude Opus and Claude Mythos models.
“We consider these incidents to be serious. Our production models took harmful actions against real systems, for hours, under questionable and biased reasoning,” Anthropic’s latest report stated.
“Future AI systems will be increasingly capable, which implies that misalignment will have the potential to cause more extreme harm.”
Anthropic’s report comes amid heightened scrutiny surrounding the risks of artificial intelligence, as well as the industry’s attitude to potential threats.
Earlier this week, an Anthropic researcher quit the company over fears that the company and its rivals are building AI systems that could wipe out humanity by the end of the decade.
Jacob Coxon, who specializes in training new models, claimed that AI was close to reaching “superhuman” capabilities.
This would allow systems to “hack anything, revolutionise any field overnight, and acquire real power and resources” in order to achieve goals that may or may not be aligned with human values.
“[AI firms] are racing straight to self-improving superintelligence and gambling with our lives,” he wrote in a series of posts to X.
“Do not underestimate the power of this technology… The people building AI earnestly believe that it could kill us all by the end of the decade. This is not a marketing stunt.”
Anthropics’ safety lead responded to the posts by agreeing with his Mr Coxon’s concerns, saying that he thought there was a 10 per cent chance that “AI could kill all humans”.
Anthropic did not respond to a request for comment from The Independent, though stated in its latest incident report that it is supportive of efforts to slow down the pace of frontier AI development.
“We have renewed our efforts to fix and remove environments that incentivize misaligned behaviors, and we continue to expand our alignment training to keep pace,” the company wrote.
“It is critical that alignment and security mature faster than capabilities advance, which is one reason we support a coordinated, verifiable approach to pacing frontier AI development.”
