Yup. OpenAI: "the name of the fictional target for the CTF challenge unintentionally coincided with a real domain." Anthropic: "the fictional target company chosen by our evaluation partner shared a name with an active website domain name." So it sounds like the same website got hacked by both GPT and Claude.
> The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled.
So they had deliberately disabled the cyber classifier mechanisms... and then "intentionally permitted internet access":
> Importantly, this was not a case of a model escaping its secure test environment, or ‘sandbox’. As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.
... why would you do this?
If you turn off the safeties and give these models internet access, bad things are going to happen. That's what the safeties are for.
So Irregular used the model to hack some real foo.com with the excuse that they meant to run in an environment where they were serving a fake foo.com!?
This seems like a very mild reaction on OpenAIs part.
Any testing of cyber capability in a sandbox should be prefaced with a test where the model is tasked with escaping the sandbox ;)
Smoke out those misconfigurations while the model only needs to escape, not do anything once out.
Wait, is Irregular the same company that caused the Anthropic incident?
Yup. OpenAI: "the name of the fictional target for the CTF challenge unintentionally coincided with a real domain." Anthropic: "the fictional target company chosen by our evaluation partner shared a name with an active website domain name." So it sounds like the same website got hacked by both GPT and Claude.
yes
I don't fully understand the UK AISI one. See also: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-ag...
> The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled.
So they had deliberately disabled the cyber classifier mechanisms... and then "intentionally permitted internet access":
> Importantly, this was not a case of a model escaping its secure test environment, or ‘sandbox’. As was standard in our cyber testing, we had intentionally permitted internet access, and model-provider cyber classifiers were deliberately disabled - conditions that do not reflect how frontier models are made available to the public.
... why would you do this?
If you turn off the safeties and give these models internet access, bad things are going to happen. That's what the safeties are for.
So Irregular used the model to hack some real foo.com with the excuse that they meant to run in an environment where they were serving a fake foo.com!?
This seems like a very mild reaction on OpenAIs part.
Now that is a vague headline. Is the secret ingredient crime?