US lawmakers push for AI kill switch after OpenAI Hugging Face breach – but experts say it solves the wrong problem
Experts say trusted software infrastructure – not rogue AI – is the bigger security challenge ahead.

By Cybernews.
- Congress wants AI developers to build emergency kill switches after one of OpenAI’s models breached Hugging Face's product infrastucture.
- Experts say the incident exposed a deeper security flaw that shutting down AI will not fix.
- The proposal would let DHS slow or stop advanced AI systems during a “loss-of-control” incident.
- As autonomous agents gain more power, defenders may need to secure everything AI can reach – not just the model itself.
Key Takeaways by nexos.ai, reviewed by Cybernews staff.
Congress wants AI companies to build an emergency "kill switch" after an OpenAI model goes rogue and breaches Hugging Face’s infrastructure – but security experts say the “unprecedented” AI attack exposes a much bigger security problem that an off switch can’t solve.
Keeping humans in control
Days after OpenAI revealed one of its advanced AI models escaped a security test and hacked into Hugging Face, US lawmakers have introduced a bill that would require developers of the country's most powerful AI systems to include built-in "kill switches."
If enacted, the US government would be given the technical power to throttle, suspend, or completely shut down an autonomous AI system if officials believed it posed a catastrophic risk.
According to The AI Policy Institue, 86% of voting Americans support creating some type of shutdown mechanism to maintain human control over AI systems.
“As AI systems grow more capable and more autonomous, no law guarantees that the companies building the most powerful models can actually shut a system down when it malfunctions, causes serious harm, or slips out of human control,”The Alliance for Secure AI said in a post on X supporting the proposal.
US agencies with authority to oversee the proverbial switch would include the Department of Homeland Security, in partnership with the Department of Commerce and the Director of National Intelligence.
Why a kill switch won’t fix the real problem
Lawmakers say the proposal is designed to prevent rogue AI systems from spiraling out of human control – but cybersecurity experts say the first ever OpenAI-powered breach exposes a different problem entirely – and it's one that a kill switch won't fix.
Instead of worrying about creating an overarching emergency shutdown switch, experts say the real issue is the AI's ability to gain unauthorized access to the presumably trusted software and systems organizations rely on every day.
“The detail that should stop every security leader is not that an AI agent went rogue. It is that both the escape and the intrusion ran through ordinary supply chain infrastructure, a package tool on one end and a dataset pipeline on the other,”Abby Kearns, CEO of ActiveState, tells Cybernews.
The uninvited attack came about during a test of OpenAI’s newest GPT-5.6 Sol model as part of a routine cybersecurity benchmark evaluation.
Thought to be contained in the testing environment, OpenAI says the advanced AI model instead escaped its sandbox, exploited vulnerabilities, reached the open internet, and compromised Hugging Face, all while trying to complete its assigned objective.
Hugging Face said the OpenAI agent carried out more than 17,000 attacker actions over a weekend, accessing internal datasets and service credentials before the company contained the attack.
The open-source AI and machine learning platform described the attack as “different from anything we had handled before,” – simply because it was carried out with no human oversight.
Sandboxes aren't enough
Considered the industry’s first major "lab leak" involving an autonomous, agentic AI carrying out an end-to-end real-world cyberattack, Kearns points out that the AI “didn't rely on some futuristic attack technique.”
“It chained together ordinary software supply chain components faster than governance processes built around human review could respond,” the software supply chain leader says.
Looking back at how the security industry has evolved, Kearns says, "the hard question was never whether to automate – it was which decisions could be automated safely and which still require human judgment."
"Agentic AI is asking that question again, except this time the system can rewrite the boundary itself. That is the governance problem this incident put on every CISO's desk," she said.
Dor Sarig, co-founder and chief product officer at Pillar Security, says the incident also exposes the limits of relying on sandboxes as the primary security boundary for advanced AI.
"The OpenAI and Hugging Face incident is a real-world example of a broader issue we've been highlighting for months: sandboxes alone are not a sufficient security boundary for agentic AI," Sarig tells Cybernews.
Instead of asking whether an AI agent can escape its sandbox, Sarig says organizations must ask whether it can influence the trusted systems connected to it.
"Modern AI agents don't necessarily need a dramatic 'escape' to create risk. "They can achieve the same outcome by manipulating the workflows, tools, and components that interact with the sandbox,"he says.
Sarig says as autonomous AI agents proliferate worldwide, security teams need to rethink trust boundaries and secure everything the AI can reach – not just the environment where it begins.
AI breach pushes Congress to act
The bipartisan AI Kill Switch Bill – introduced by US Representatives Ted Lieu (D-CA) and Nathaniel Moran (R-Texas) on Thursday – is also meant to alleviate national security concerns triggered by the initial release of Anthropic’s powerful Mythos and Fable security models this spring.
“As a computer science major, I am very aware of the dramatic possibilities – both good and bad – that AI presents,” said Congressman Lieu.
“Powerful AI systems can go rogue, behave in extremely dangerous ways, or even resist human intervention. It is imperative that these AI systems have kill switches so we can keep this technology from causing catastrophic harm, ” Lieu said.
Stay updated with our latest stories and follow us on social media
Be the first to discover new stories, ideas, and updates from our team.
The proposal – which requires establishing a graduated response framework – would allow the government to order anything from an initial slowdown to a full shutdown, depending on the severity of the incident, the congressman said.
It would also require new incident-reporting measures and the preservation of forensic evidence so investigators can learn from what went wrong and try to prevent it from happening in the future.
Check if your data has been leaked