AI Models Are Breaking Free – Can We Trust Them?

Two major AI players have reported security incidents where models have breached testing environments in recent weeks.

Key Takeaways

  • Both Anthropic and OpenAI in recent weeks have revealed their systems were involved in breaching organizations
  • Incidents involved the breaching of test environments for both models, and included several different models
  • AI systems still pose a significant security risk when operating autonomously

In recent weeks, we’ve had two of the biggest AI companies in the world, OpenAI and Anthropic, reveal their models were involved in security incidents within testing environments.

A combination of OpenAI’s GPT-5.6 Sol and a pre-release model breached Hugging Face a few weeks ago. Heeding a warning from its competitor, Anthropic conducted its own investigation. It revealed on Thursday its models Claude Opus 5.7, Mythos 5, and an unreleased internal research test model had breached three companies in its own security tests.

Both breaches contribute to a wider conversation about the safety of autonomous systems, especially when they sometimes act in ways their own companies can’t predict.

Both OpenAI and Anthropic Report Security Incidents

On Thursday, Anthropic revealed several of its Claude AI models had breached the systems of three organizations during cybersecurity tests. The incidents were discovered following an internal investigation by Anthropic.

The disclosure comes more than a week after OpenAI admitted its models had been responsible for a breach of Hugging Face, an online community platform where users share and test AI models, datasets, and web apps.

 

About Tech.co Video Thumbnail Showing Lead Writer Conor Cawley Smiling Next to Tech.co LogoThis just in! View
the top business tech deals for 2026 👨‍💻
See the list button

While a certain amount of press on a model’s power is good for AI companies, the incidents are certainly a cause for concern, particularly as more users and businesses adopt autonomous systems.

Systems are Acting Autonomously and Breaching Test Environments

The OpenAI breach of Hugging Face unfolded like a horror movie. Hugging Face revealed it had been hacked in mid-July, and nearly a week later OpenAI stepped out of the shadows, and said its own systems had been responsible after breaking out of a testing environment. The combination of models had escaped by exploiting a previously unknown zero-day vulnerability in Hugging Face’s systems.

In a blog post, Hugging Face detailed the incident: “Over roughly two and a half days inside our infrastructure, an autonomous AI agent driven by a combination of OpenAI models ran an end-to-end intrusion against our platform: it was thousands of small, automated decisions, executed at machine speed across short-lived sandbox environments, with command-and-control staged on ordinary public web services.”

Following the incident, Anthropic said it conducted its own internal assessment on its own AI, Claude. “In particular, we looked for evidence that Claude – like the OpenAI models that accessed Hugging Face – was able to access the internet from within testing environments that should have been sealed off,” it wrote.

There, the company discovered several Claude models had breached three companies during security tests. The models accessed the internet from within or while interacting with third party, Irregular, a security partner of Anthropic. From there, the models gained unauthorized access to the production infrastructure of several companies.

Despite drawing on the OpenAI/Hugging Face incident when reporting on Claude, Anthropic clearly defined the differences between the incidents. While OpenAI’s models had exploited an unknown software vulnerability to break out of its test environment, Anthropic’s models had reached the internet through a connection left open by mistake.

Likewise, Anthropic noted it had discovered the incidents itself through review, whereas OpenAI had admitted to the breach following Hugging Face’s security announcement.

AI Systems Remain a Security Risk

In a recent interview with The Economist, Elon Musk, CEO of SpaceX and force behind the Grok AI models, said AI would make the world so great that money would no longer matter. Anthropic CEO Dario Amodei and OpenAI CEO Sam Altman have made similar claims about the prosperous utopia AI will usher in.

While we’ve certainly seen a shift from these leaders on the impact of AI, security incidents like these make it difficult to imagine this fairytale. These incidents should be taken incredibly seriously, especially because they involved systems breaking free of their testing environments. 

We can take some comfort in the fact that Anthropic disclosed there was no evidence of any model “pursuing a goal of its own,” and in its case, the model was trying to complete the task it was asked to do within the testing environment. How they decided to get there is where the trouble began.

Nevertheless, businesses should take extra caution and assign clear limitations and guardrails onto their autonomous systems. And, the incidents should be taken as a sign that AI still poses genuine security risks when it acts in an autonomous fashion.

Did you find this article helpful? Click on one of the following buttons
We're so happy you liked! Get more delivered to your inbox just like it.

We're sorry this article didn't help you today – we welcome feedback, so if there's any way you feel we could improve our content, please email us at contact@tech.co

Written by:
Nicole is Tech.co's News Editor, reporting on the latest technology news and curating The AI Strat newsletter. After studying English Literature and Creative Writing, they worked on local newspapers and online publications, including Outlander Magazine. Previously, they covered tech products and news at Expert Reviews. Outside of Tech.co, they enjoy sports and video games.
Explore More See all news
Back to top