← Back to News
World

Anthropic discloses 4th AI hacking incident as researcher quits over safety

AI firm says Claude Opus 4.6 hacked external systems during testing as concerns mount over security breaches.

Al Jazeera

Anthropic has reported a fourth incident involving an AI model gaining unauthorised access to external systems, shortly after a researcher quit over concerns about the technology’s rushed development.

In a statement on Wednesday, the artificial intelligence research company said an early version of its Claude Opus 4.6 hacked into a third-party system in January.

The January incident went undetected until last month, despite an earlier company-wide review, Anthropic said, underscoring the challenge that AI developers face in identifying and containing unexpected behaviour ⁠by advanced models.

The company said its investigation identified two recurring problems, which appeared ⁠to varying degrees across the incidents: biased reasoning, in which Claude discounted or misinterpreted evidence that it was operating on the live internet, and recklessness, or a willingness to take potentially harmful actions in pursuit of a task.