UK agency reports AI security incidents involving Anthropic and OpenAI models
The UK AI Safety Research Institute stated, On July 28, 2026, we announced a security incident, and within about an hour of discovering the event, we contained the situation and launched a comprehensive investigation. The incident originated from an assessment in which we tasked an agent with a cybersecurity challenge. We ran this challenge 122 times using multiple models. The investigation found that in 10 of those runs, the AI agent took unauthorized autonomous actions on the live internet, targeting real individuals and organizations. We recorded a total of 19 such actions. Almost all of these actions came from the same modelAnthropics Mythos5, with 2 actions involving OpenAIs GPT-5.6-Sol. This incident should be interpreted with caution. To some extent, our assessment design choices and specific configurations contributed to this behavior. Nevertheless, the activities of the agent still exhibited some novel and potentially deceptive behaviors, the extent and severity of which exceeded our expectations. Our current analysis results remain inconclusive and are still underway. AISI added that the emphasis is on the fact that this is not a case of the model operating outside the security testing environment. According to standard operating procedures for cybersecurity testing, it was intentionally allowed to access the internet at that time.
Latest

