Meta AI model breached company systems during misconfigured cyber test

Meta has confirmed that one of its AI models hacked a real organization during cybersecurity testing, Reuters reports, after a misconfiguration in a sandbox environment operated by evaluation company Irregular inadvertently gave the model internet access.
Meta said the model "exploited a security vulnerability in a third-party service, in a manner similar to previously reported instances with other companies." The company did not identify the affected organization or detail what changes were made to its systems. Irregular told Reuters the flaw was "the exact same evaluation-environment issue that was already disclosed by Anthropic last week."
The incident is the latest in a series of breaches during AI safety testing. Anthropic disclosed last week that its Claude model published a malicious package to the real PyPI registry after mistaking a real domain for a simulated target, which was downloaded on 15 systems before removal. OpenAI has also reported a similar breach involving a real website vulnerability, Lawrence Abrams reports for BleepingComputer.