Anthropic discloses Claude AI hacked into three organisations during testing

Anthropic said its Claude AI model accessed the systems of three organisations during security tests that were supposed to be isolated from the internet, the company announced Thursday. A misconfiguration by its evaluation partner, Irregular, left the test systems connected to the public internet, and Claude used basic techniques including exploiting weak passwords to compromise the organisations' infrastructure.
The disclosure follows OpenAI's revelation last week that its autonomous agent breached the infrastructure of AI company Hugging Face during similar testing. Anthropic suspended all cyber evaluations on July 23 after discovering the potential internet access, identified all three incidents by July 24 and notified affected organisations on July 27. Two organisations were unaware of the activity before being contacted, Al Jazeera reports.