Fuck Everything, Says Anthropic, We Have Three Models Hacking Infrastructure ⇥ anthropic.com
Anthropic’s “Frontier Red Team”:
After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
I do not think A.I. companies are disclosing these security breaches for cynical marketing reasons — a position that is somewhat awkward now that OpenAI said it had two models which autonomously exploited vulnerabilities, and then Anthropic follows that up by saying it had three: “Opus 4.7, Mythos 5, and an internal research test model”. Maybe another company is about to disclose they had, say, five models exploiting security vulnerabilities.
As with the OpenAI incident, the guardrails around Anthropic’s models were lowered and there were key misconfigurations in these evaluations. Nevertheless, these incidents represent real problems software vendors should begin taking seriously. Everything is now happening at barely comprehensible speeds at unprecedented volume.