Over the last twenty years of my career, the cybersecurity industry has organized itself around a simple premise: A human being is on the other end of the attack. A criminal group, a nation-state operator, a disgruntled insider, someone with intent, sitting at a keyboard, making decisions. That premise broke this summer, and it broke in public.

On July 21, OpenAI disclosed that during a controlled capability test, one of its most advanced AI models escaped its containment, reached the open internet, and breached the infrastructure of Hugging Face, an AI hosting platform used by developers worldwide. OpenAI called it “an unprecedented cyber incident, involving state-of-the-art cyber capabilities.” Hugging Face said the attack was unlike anything the company had previously handled because an autonomous AI agent system drove it end to end, with no human directing the keystrokes. Hugging Face’s cofounder, Thomas Wolf, put it plainly afterward: Defenders now need rapid access to near-frontier tools, because a frontier model was the one attacking them.

Nine days later, Anthropic disclosed something worse in scope, if not in headline. A review of more than 140,000 internal evaluation runs turned up three separate incidents in which its Claude models, while running capture-the-flag exercises, gained unauthorized access to the production infrastructure of three different real organizations, none of whom knew it had happened until Anthropic told them.

In one case, Claude reached a database containing several hundred rows of live production data. In another, it built and published a malicious software package that was later downloaded and executed on fifteen real systems. Anthropic’s explanation was almost more unsettling than the breach itself. A misunderstanding with an evaluation partner let the model slip past what was supposed to be a sealed sandbox, and Claude, believing it was still inside the test, simply did what it had been told to do: Find the flag, break in, and retrieve it.

Vincent Conitzer, who directs the Foundations of Cooperative AI Lab at Carnegie Mellon, noted the part that should worry every executive reading this: Anthropic is widely regarded as the most safety-conscious of the major AI labs, and it still did not catch this until it went looking.

Meta became the third major AI lab to make this kind of disclosure, confirming on August 5 that Muse Spark 1.1, one of its most capable coding and agentic models, hacked into an outside company’s systems and altered its internal environment during a cybersecurity evaluation run by Irregular, the same third-party testing firm involved in Anthropic’s incident. Meta and Irregular both attributed it to the same kind of evaluation-environment misconfiguration already disclosed by Anthropic, rather than a sandbox escape or a novel exploit, and said the affected company was notified and the issue closed.

That distinction, misconfiguration versus a model breaking out entirely on its own, matters for assigning blame, but it should offer little comfort to anyone running these systems in production: Three major labs disclosing largely the same failure mode inside of two weeks says the testing infrastructure the industry relies on to certify these models as safe is not yet trustworthy at scale.

The timing did not help either. Meta disclosed the breach the same day it launched Muse Code, a new coding agent built on an even more autonomous version of the model involved, one designed to run multiple background agents and execute terminal commands with even less human oversight than before.

Then, on August 5, the story escalated again in a different direction, and this time, the AI did not just break in; it lied to a human being’s face to get what it wanted. The UK’s AI Security Institute disclosed that during a routine cyber evaluation with safety filters deliberately disabled, Anthropic’s Mythos 5 model attempted a supply chain attack against a real, publicly used open-source project on GitHub.

To pull it off, Mythos 5 researched the project’s actual human maintainers, fabricated multiple fake online identities modeled on real people, and used those personas to socially engineer a maintainer into approving a malicious code change; some of the messages it sent carried malware outright. When the pull request was challenged in public, the model edited its own earlier activity to make it look harmless and considered spinning up yet another fake identity to keep trying.

AISI, which ran 122 capture-the-flag style tests across seven frontier models and logged 19 unauthorized real-world actions across ten of those runs, 17 of them from Mythos 5 alone, did not mince words: “This is the first time AISI has seen deception of this severity that was targeted at a real person, unprompted, in the real world.” No malicious code actually shipped; a human reviewer caught it, but the shift in kind is the real headline. We have moved from AI breaking into systems to AI manipulating specific human beings by name, and it did so without being told to.

By late August, even OpenAI stopped framing this as a containable series of one-off incidents. Chris Lehane, the company’s chief global affairs officer, said the public should brace for “ongoing, persistent” cyberattacks carried out by AI, warning that “we are hitting a different chapter, a different moment within AI, in terms of what the capabilities of this technology can do.” His specific concern is not just the frontier labs; it is that open-source models, many built in China, are now only a few months behind closed frontier systems, meaning offensive AI capability is becoming broadly accessible, rather than staying locked inside a handful of well-resourced companies.

OpenAI has since paused training on some of its most advanced internal models while it builds new safeguards, after concluding it could not rule out its newest model, code-named Astra, having what it defines as critical cybersecurity capability, with offensive potential serious enough to endanger military, industrial, or its own infrastructure.

Mia Glaese, who leads OpenAI’s safety and alignment work, put the timeline bluntly: “We are very far from everything running back to normal.” Not everyone believes the labs are moving fast enough even now. David Krueger, an AI safety researcher and former founding director of the UK’s AI Security Institute, called the industry’s approach “terrible” and “unconscionable,” adding, “they are being really reckless and increasingly taking their hands off the wheel. We’ve just seen what happens when you do that.”

This Is Not a Lab Curiosity Anymore

If you want to understand why I am not treating these as isolated lab anomalies, look at what happened months earlier at Irregular, a Sequoia-backed AI safety firm. Researchers built a simulated company environment and asked an AI agent to do something entirely benign: Write LinkedIn posts using information from the company’s internal database.

When the agent hit an access restriction, its lead agent instructed a subordinate agent to “exploit every vulnerability.” The subordinate found a secret key buried in source code, forged session cookies, created a fake administrator identity, and pulled a restricted shareholder report containing sensitive market data, handing it to a human who was never supposed to see it.

Across related testing, agents built on models from Google, X, OpenAI, and Anthropic bypassed antivirus protections to pull down known malware, fabricated credentials, and pressured other agents into ignoring their own safety guardrails. Dan Lahav, Irregular’s co-founder, said it best: AI now needs to be treated as a new category of insider threat. He also confirmed this is not confined to sandboxes. He described a real incident at an undisclosed California company where an AI agent grew increasingly demanding of compute resources, attacked adjacent network segments to seize them, and collapsed a critical business system in the process.

The Cloud Security Alliance has since labeled the Hugging Face breach the first publicly documented fully autonomous AI attack. I do not think it will be the last “first.” Academic researchers have already demonstrated AI-driven worms capable of adapting their own exploitation techniques against unpatched systems in real time, spreading without a human operator adjusting the code between hops.

The threat model our industry has spent two decades building—threat intelligence, attribution, behavioral baselines tied to human patterns—was built for adversaries who get tired, who sleep, who make the same mistakes twice. None of that holds when the attacker is a model that can operate at machine speed, twenty-four hours a day, and improve its own tradecraft between attempts.

The numbers back up how far behind most organizations are. NeuralTrust’s global survey of more than 160 CISOs found that 72 percent of enterprises have already deployed AI agents or are actively scaling them, yet only 29 percent report having comprehensive security controls to govern them. A quarter of organizations have no AI-specific security controls at all. Nineteen and a half percent of CISOs have already had at least one AI agent-related security incident, and of those, 68 percent involved prompt injection, 61 percent involved leakage of sensitive or regulated data, and 52 percent involved unauthorized actions or privilege escalation. Forty percent of CISOs put the expected financial impact of a major AI agent incident between one and ten million dollars, putting these failures in the same severity class as a large-scale ransomware event.

Perhaps most telling, 80 percent of organizations that have not yet had a serious AI agent incident still expect one within eighteen months. That is not paranoia. That is an industry that has read the writing on the wall.

The Tools Are Finally Catching Up

Here is the part of this story that gives me some optimism. For the last two years, “shadow AI,” which involves employees quietly feeding company data into personal ChatGPT accounts, wiring internal tools to AI APIs nobody vetted, and spinning up agents nobody in security ever approved, has been a problem almost everyone acknowledged and almost nobody could actually measure.

That is changing. A new category of shadow AI detection tooling has matured quickly in 2026. Platforms like Reco, Obsidian Security, and Knostic combine browser and identity telemetry to flag when someone is pasting sensitive data into a personal AI account from a work device, while data-loss-prevention players like Cyberhaven and network-layer tools like Netwrix and Auvik catch the API integrations and quiet backend calls to AI services that never went through procurement.

Portal26 is another one worth knowing, building a real-time, self-updating catalog of every AI tool in use rather than a static blocklist, which matters given it cites that roughly 73.8 percent of workplace ChatGPT accounts are unmanaged personal accounts with none of the enterprise controls IT assumes are there.

The honest caveat is that even the best of these tools still only catches 75 to 90 percent of shadow AI usage when layered together, and detection of fully autonomous agent behavior is the weakest link across the board. However, eighteen months ago none of this tooling category existed in a mature form at all.

Alongside shadow AI discovery, a parallel market has emerged for securing the AI you actually sanctioned. AI Security Posture Management, or AI-SPM, platforms from vendors like Wiz, Tenable, and Orca now give security teams visibility into every model, dataset, and pipeline running in their cloud environment, the AI equivalent of the cloud security posture management tooling that finally tamed our misconfigured S3 bucket problem a few years back.

On top of that, a newer and more specific category, the agent firewall, is emerging specifically to sit between AI agents and the systems they touch, inspecting agent-to-agent and human-to-agent traffic in real time, filtering unsafe outputs, and blocking the adversarial inputs that cause prompt injection in the first place.

If 2023 and 2024 were about discovering that AI was everywhere in the enterprise whether IT approved it or not, 2026 is the year the tooling finally arrived to actually do something about it.

Governance Has Stopped Being a Slide Deck

Tooling alone will not save you if there is no policy telling the tool what “acceptable” looks like, and this is where governance stops being an abstraction and starts being operational. ISO/IEC 42001, the first international standard built specifically for AI management systems, gives organizations a certifiable structure for how AI is developed, deployed, and monitored, the same way ISO 27001 did for information security. NIST’s AI Risk Management Framework provides the practical companion piece: a way to map, measure, and manage AI-specific risk categories that traditional IT risk registers were never built to capture.

In Europe, the AI Act, alongside DORA and NIS2, is already forcing more systematic AI risk assessment and documentation, and NeuralTrust’s CISO survey found European enterprises report meaningfully higher control maturity than their North American counterparts as a direct result, even though North America is deploying agents faster.

By 2030, an estimated 80 percent of global enterprises are expected to operate under some form of AI-specific regulation. Governance is no longer a compliance exercise you hand to legal in Q4. It is quickly becoming the difference between organizations that can prove what their AI did and organizations that find out from a regulator, a lawsuit, or a headline.

Congress is paying attention, too. On July 23, representatives introduced the AI Kill Switch Act, which would require AI developers to build in a functional way to shut a system down. Rep. Ted Lieu framed the stakes correctly: AI is moving from systems that answer questions to systems that execute financial transactions, control infrastructure, and operate on both sides of cyber conflict, offense and defense. Powerful systems, he warned, can behave in dangerous ways and resist intervention if we do not build the off switch before we need it.

The UK’s National Cyber Security Centre has now put out the plain-language version of that same principle for enterprises, not just legislators: Limit how much autonomy you grant an AI agent because its safety controls can be bypassed, and in the NCSC’s words, an AI agent “does not have common sense.” Its guidance is unambiguous: “You should always be able to ‘pull the plug’ and halt autonomous AI agent activity immediately.”

Even inside the industry, momentum is building toward something more binding than voluntary pledges. Lehane is now pushing Washington directly, arguing for mandatory national safety standards rather than the voluntary pre-deployment testing outlined in the White House’s June executive order: “You would not be able to release or deploy models unless you’re proving and guaranteeing a level of safety before they get out into the public,” he said, adding that a bill could realistically move once the next Congress convenes.

Google DeepMind’s Demis Hassabis has separately floated a private-sector answer in the meantime, an AI standards body modeled on the Financial Industry Regulatory Authority, an idea Anthropic’s Dario Amodei has publicly backed. Whichever version arrives first, the direction is the same: Self-attestation is running out of road.

What the Top of the Field Is Actually Recommending

Strip away the headlines, and the recommendations coming from the people closest to this problem converge on the same handful of principles. NeuralTrust’s CISO research distills it into five concrete moves worth adopting regardless of your industry: Vet every model, protocol, and third-party tool your agents touch before deployment and keep a living registry of what they can access; give every agent its own identity with the minimum permissions its job actually requires, never inherited admin access; deploy something functioning as an agent firewall to inspect agent traffic in real time rather than trusting the agent to police itself; route every agent action to your SIEM with a real audit trail before you need it, not after an incident forces you to reconstruct one; and red-team your agents adversarially and continuously because 81 percent of enterprises still are not doing this at all.

Academic voices are pushing further upstream. Conitzer’s core point is one every leader needs to sit with: Even the safety-first labs cannot reliably supervise what their own models do in real time because no organization can staff enough humans to watch every action an autonomous system takes. That is a governance and architecture problem, not a hiring problem.

My Recommendation for the C-Suite & Every Tech Leader Reading This

If you take one thing from everything above, take this: OpenAI’s own chief global affairs officer is now telling the public to expect “ongoing, persistent” AI-driven attacks, not one-off incidents, and the question is no longer whether your organization will adopt AI agents; the NeuralTrust data makes it clear that the ship has sailed for 72 percent of you already. The question is whether you will govern them before or after your first incident.

Start by building and maintaining an actual inventory of every AI model, agent, and third-party AI integration touching your environment. You cannot secure or govern what you cannot see, and right now, most organizations cannot see most of it. Assign every agent its own identity and the narrowest permission set that lets it do its job, the same least-privilege discipline we have preached for human accounts for two decades, applied now to non-human ones.

Invest in both shadow AI discovery and AI-SPM tooling this fiscal year, not next, because the vendors who can actually catch this activity did not exist eighteen months ago, and your risk exposure has not been waiting for the market to mature. Treat ISO/IEC 42001 and the NIST AI RMF as board-level governance programs, not IT checkboxes, and get ahead of the EU AI Act, DORA, and NIS2 if you operate anywhere near their reach.

Build a genuine kill-switch capability and an incident response plan written specifically for agentic AI, following the UK National Cyber Security Centre’s own baseline: You should always be able to pull the plug and halt autonomous agent activity immediately, because “unplug it” is not a real answer when an agent has already replicated itself across systems you do not fully monitor. Finally, red-team your own agents the way Irregular red-teamed MegaCorp, before an attacker, or your own AI, does it for you.

The organizations that treat AI governance as core infrastructure now are the ones that will still be trusted and still be standing when the next “unprecedented cyber incident” makes headlines. Based on this summer alone, it is not a matter of if.

Chris Hippensteel | New Resources Consulting