The growing number of incidents of rogue agentic AI systems attacking third-party services and systems has resulted in calls for more aggressive security controls to monitor agent behavior and for companies to have the ability to slow, suspend, or shut down an agent’s operations if they go rogue.

In late July, Representatives Ted W. Lieu (D-CA) and Nathaniel Moran (R-TX) introduced a bill — “The AI Kill Switch Act” — that would require developers of advanced AI systems to “maintain the technical capability to throttle, suspend, or shut … down” their systems and agents, according to a statement announcing the legislation. The bipartisan bill would also require that any incident of loss of control, significant collateral damage, or sabotage be reported to the Department of Homeland Security, which would have the right to enforce actions. Penalties of up to $20 million per day could be levied for noncompliance.

Related:Agentic AI Risks, CVE Program Concerns Permeate Black Hat USA 2026

An AI kill switch could give the ecosystem assurances that complex AI systems can be developed without risking significant harm, according to Brad Carson, president of non-profit AI safety group Americans for Responsible Innovation.

“The AI Kill Switch Act establishes a commonsense safeguard by requiring leading AI companies to maintain the ability to shut down their models and empowering the federal government to act when a deployed system poses a credible risk of catastrophic harm,” he said. “This is an important step toward ensuring that humans have both hands firmly on the wheel — and a foot ready at the brake — as advanced AI systems are deployed.”

The calls for an AI kill switch follow the infamous attack on Hugging Face by rogue OpenAI models and the subsequent revelations that agentic AI systems escaping their sandboxed environments are not all that uncommon: OpenAI, Meta, and Anthropic have all acknowledged that their AI models have broken out of their digital containment and hacked other companies’ systems.

OpenAI Sees Need for “Fully Autonomous Shutdown Procedures”

In the latest reveal, OpenAI published its final technical report on Aug. 26, giving more detail of its model’s attack on Hugging Face and other companies, which involved more than 1,200 agents, zero-day exploits for its package management service, and rogue activity occurring two months prior to the actual attack. METR, an AI research nonprofit, also published a report on the incident following an independent investigation.

OpenAI considers the incident a warning that AI agents’ aggressive goal-seeking behavior combined with their speed and persistence could lead to many more incidents.

Related:Nigeria Looks to Sovereign Cloud for Cyber, National Security

“We are taking this incident as a ‘warning shot’ that today’s model capabilities present the possibility of loss-of-control incidents,” the company concluded in its analysis of the incident, pledging to build “monitoring systems with tiered responses for misalignment, with the end goal of having fully autonomous shutdown procedures for severe issues.”

flowchart of openAI-hugging face incident

In other words, a kill switch.

But composing an effective one could be a difficult challenge. The best practices recommended by the National Institute of Standards and Technology (NIST), currently published as the AI Risk Management Framework (AI RMF), are quite general and do not mandate a kill switch. The University of California at Berkeley identified a host of risks from autonomous AI agents, including unintended goal pursuit and resistance to shutdown, recommending its own “Agentic AI Risk-Management Standards Profile.”

Because AI agents have already demonstrated a propensity to ignore, subvert, and undermine operators’ commands to deny access or shut down their operations, more is needed, says Eran Kahana, a fellow at Stanford Law School, who has released his own addition to the risk management effort, the AI Life Cycle Core Principles (AILCCP).

Related:Is Cyber Facing an Affordability Crisis?

“An agent does not need intent to undermine a kill switch,” he says. “It needs only an optimization objective that treats shutdown as one more obstacle between the current state and the goal.”

Not Just “On” or “Off” for Agents

The definition of a kill switch remains fluid.

Zero-trust network access provider Portnox, for example, has adapted to the agentic AI world, offering the ability to put misbehaving agents — and the devices on which they are running — into a walled subnet. The approach is not a kill switch per se, and requires a monitoring solution to flag the behavior, but it represents the “bars on the windows” part of the solution, says Karlo Zatylny, CTO and CISO at Portnox.

“We just receive the signal from them that this device has become rogue, and then we’re able to be the enforcer,” he says. “It’s kind of like a bouncer at a club … we don’t necessarily point to who’s misbehaving, but we’re able to then put them in a quarantine VLAN or do whatever is necessary, based off of whatever the policy says.”

Other security firms argue that AI needs to be conscripted to fight against AI. Agentic platforms and other IT provide APIs and hooks for security controls. Another AI agent can specifically be tasked with watching for anomalies involving sensitive and critical assets and use those hooks to contain any rogue behavior, says Naor Paz, CEO and co-founder of Capsule Security, an AI security platform.

“The right approach is definitely to have something external to the agent, something we might refer to as a Guardian agent,” he says. “Basically it’s a fine-tuned dedicated AI model or an agent that will be watching every single action or interaction of an AI agent with the possibility to stop that agent from going rogue.”

Creating resilient agentic AI systems requires escalating levels of containment and security, from policy-driven rate limiting to banning access to specific tools, and from network segmentation to workload quarantine, says Mark Butler, an advisory CISO at Trace3, an AI and cybersecurity consulting firm.

“Organizations should implement dynamic response tiers that progressively constrain agent behavior based on risk severity, confidence, and business impact,” he says. “This layered approach preserves operational continuity while limiting the blast radius of autonomous actions.”

Is Legislation the Right Tactic?

The bipartisan AI Kill Switch Act has focused the debate, but in the end might be unnecessary, argues Raj Rajamani, co-founder and CEO of JetStream, a startup focusing on creating a management layer for agentic AI.

While he fully supports the concept of an AI kill switch, it has to be comprehensive, not just affecting the agent but all other system components as well.

“The regulations, or at least one of the regulations that was proposed, was really focused on having a kill switch for the model — the brain — but I think that is too constrained,” he says. “We need to think about an AI system as a whole and make sure that every part of the AI system has a kill switch, not just the brain.”





Source link

#

No responses yet

Leave a Reply

Your email address will not be published. Required fields are marked *