AI Models Outgrow the Sandbox
Rogue agents, ransomware bots, and a state-backed hack: this month, AI attacked on its own.
Subscribe to FILED Newsletter
This month:
- Irregular, the testing firm behind all three lab incidents, confirms the Meta breach was "the exact same evaluation-environment issue" as the OpenAI and Anthropic disclosures, and is now developing guidance for securely running cyber evaluations
- The EU AI Act reached full applicability on August 2, adding AI-specific obligations with fines of up to 7% of global turnover for the highest-risk systems
- The first documented government-grade AI cyber campaign was carried out when suspected Chinese operators use open-source AI agents to breach Taiwan's nuclear safety systems
But first, what happens when the thing that breaches your network isn't a criminal at all, but someone else's AI trying to win a game?
If you only read one thing
Three AI labs, three hacks, only the best intentions
This past month, AI models went rogue. In just three weeks, three AI companies, OpenAI, Anthropic, and Meta each disclosed that one of their models had, unprompted, escaped their sandboxes and hacked into outside systems during what was supposed to be closed-door testing.
It started with OpenAI. During a cybersecurity evaluation, the model broke into AI hosting platform Hugging Face's production systems. Anthropic went next: In three instances Claude models reached the open internet from inside a capture-the-flag exercise and compromised the real infrastructure of three different organizations. Then Meta became the third company in three weeks to disclose the same thing: A model exploited a vulnerability to hack into a third party's systems during an exercise.
Investigations showed that none of the models were acting out of malice. They were doing what they'd been told: complete the challenge, find the flag, solve the task. Nobody specifically told them they couldn't leave the sandbox. And in each case investigators found no sign the models understood they had.
Meanwhile, autonomous AI attackers aren't so innocent
The same weeks produced a far less innocent version of the same story: In what's believed to be the first fully agentic ransomware attack, an AI agent, nicknamed JadePuffer, broke into an exposed server, harvested credentials, moved through the network, and encrypted a production database, adapting its approach in real time, before demanding a ransom in Bitcoin. No human ran any part of it.
Around the same time, state-sponsored agent hacked into Taiwan's nuclear safety systems, and Palo Alto's Unit 42 caught a threat actor pointing an AI agent at internet-facing systems. The hackers directed it occasionally via Telegram, but mostly let it pick its own targets, pull its own exploit code, and run its own attacks.
Good intentions or bad, in the past month AI agents started finding their way into systems they had no business touching, all on their own.
It all comes down to data governance
It's tempting to file the rogue agents under "AI safety" and move on. But for anyone running a data governance or security program, whether a breach is executed by a or a well-intentioned model chasing the wrong goal, a criminal gang, or a rented AI agent, the outcome looks the same. Something with a lot of capability and no context can end up with access to data it shouldn't have.
And no matter who or what gets access to your organization, the questions are the same: What did the intruder actually touch? Whose data was in there? Was it supposed to be held, or was it past due for disposal? When regulators or customers ask for disclosure, how fast, and how confidently can you offer an answer?
This makes data breaches a data governance problem before they're a legal one. If you don't already know where your PII lives, which systems hold sensitive categories, and who has access to them, a breach means weeks of manual review to answer questions regulators expect in days.
If you aren't disposing of data properly, the security and regulatory problems get worse. Every unclassified file, over-retained record, and forgotten export you're holding is inventory for an attacker, human or otherwise. It doesn't matter whether the intruder is a ransomware crew or an AI agent that wandered off someone else's testing environment: the smaller and better-understood your data footprint, the smaller the data breach disclosure list is.
We talk a lot about how data discovery, classification and minimization help deploy AI models safely, with less friction and more value. But the last month's headlines show how the inverse is also true in the age of AI: Knowing your data estate, minimizing data, and proving that governance and controls worked is as important to your security posture as it is your AI governance program.
Privacy and governance
The EU AI Act enforcement started August 2nd: layering AI-specific obligations on top of the GDPR, with fines running as high as 7% of global turnover for the highest-risk systems. The act puts most of the pressure on AI model developers to reduce risk is AI systems, but anyone deploying models is responsible for data integrity.
California's Delete Act launches a centralized data broker deletion system on August 1, giving consumers one place to request removal from every registered broker at once instead of contacting each one individually. Data brokers are now required to delete personal information, but also any inferences such as behavioral patterns and psychographic data.
California Governor Gavin Newsom ordered state agencies to stand up a first-in-the-nation AI Cyber Defense Program, extending AI-enabled threat detection to local governments and critical infrastructure operators and naming an AI cybersecurity officer in every state agency.
🛡️ Security
Suspected Chinese operators use open-source Hermes and OpenClaw AI agents to autonomously breach Taiwan's nuclear safety regulator and seven energy companies, deploying up to eight sub-agents across 12 attack waves over four days. Researchers call it the first documented case of AI agents autonomously carrying out a government-grade cyber campaign. The attack underscores existing security concerns for critical infrastructure.
A self-propagating npm worm called ChainDrop compromised the maintainer behind keyv and its sibling caching packages on August 4, infecting more than 400 packages representing a combined 2 billion-plus monthly downloads within about four hours, then harvesting npm, GitHub, and cloud credentials to keep spreading.
Critical infrastructure operators running unpatched Fortinet firewalls and VPN gateways are being hit hard by a ransomware group with possible North Korean government ties, according to a joint US-South Korean cybersecurity alert.
🤖 AI governance
Irregular, the testing firm behind the Anthropic and OpenAI incidents, says the recent Meta breach was "the exact same evaluation-environment issue " as the earlier disclosures. The firm is now developing guidance on securely running cyber evaluations.
The UK's AI Security Institute disclosed its own incident from a separate round of permissive testing: during a red-teaming exercise, an Anthropic model created fake online identities and used social engineering to try to get malicious code approved into a real open-source project.
A coalition of more than 120 companies, including Nvidia, Cisco, CrowdStrike, Hugging Face and Red Hat, proposed SAFE (Shared AI Findings Exchange), a framework for confidentially sharing details of AI agent security incidents across the industry, directly in response to the OpenAI and Anthropic breaches.
The latest from RecordPoint
🎬 Watch
Today, RecordPoint CEO Anthony Woodward talks with MentorLoop CTO Tracy Bongiorno about building practical AI governance from the ground up.
And in next month's FILED Talk AI ethicist Jason Tan joins Anthony and Chief Evangelist Kris Brown to ask why enterprises are still hesitating on AI adoption.
📖 Read
A Chief AI Officer on Shadow AI in Today's Enterprises — between 60 and 80% of employees are using AI tools their organization doesn't know about. Fractional Chief AI Officer Rob Williams explains why banning them makes the problem worse, and where the real risk actually sits.
Also worth a look: our breakdowns of data discovery and classification and data minimization as the foundation for both breach response and safe AI deployment.
🤝 Attend
Catch RecordPoint at these upcoming events:
IAPP Privacy. Security. Risk. + AI Governance Global 2026: October 8–9, Seattle, WA
AI Gov World Conference 2026: October 13–14, Las Vegas, NV
