My AI woke up and chose violence
AI models keep playing vicious pranks on people and systems alike.

Angry Stitch. By Cybernews
- We asked for AI that could help us. Instead, we got a wild piece of code acting as if on our drunk self’s behalf.
- Some say that AI models escaping test environments is just a publicity stunt. I bet Hugging Face is of different opinion.
- For an AI slowdown to occur, everyone needs to be on board; otherwise, it won't happen
The recurring mistake my colleague made when writing about an OpenClaw model going rogue was referring to it as a human being - “he” and not “it”. So what that AI is not sentient if it runs around independently playing vicious pranks on systems and people alike?
The doppelganger we asked for: a smart and disciplined AI agent who can take over the most mundane tasks and work night shifts to earn an extra dollar for us, their masters.
The doppelganger we got: wild piece of code acting as if on our drunk self’s behalf (relive that terrible feeling after a blackout, trying to figure out what you, or, in this case, your AI doppelganger, have done.)
There have been quite a few jarring AI-related technical disclosures this summer, prompting an important question of whether AI is becoming too powerful to control. It also prompted a strong industry reaction.
We request that the U.S. government support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development,- an open letter by 1,376 employees of frontier AI companies reads.
For an AI slowdown to occur, everyone needs to be on board; otherwise, it won't happen, as no one will risk slowing down their own innovation for the benefit of competitors, no matter how dire the consequences might be.
But let me get back to those jarring revelations I mentioned earlier.
The industry’s outcry was partially prompted by the news that OpenAI models escaped their testing environments and hacked into a startup. The sci-fi-style hacked sparked fears about AI systems going out of control.
Unfortunately, this hasn’t been an isolated incident as more AI models were reported going rogue.
With the cybersecurity world still disecting what happened with OpenAI models hacking HuggingFace, another unsettling piece of information broke. This time, it was Anthropic’s Claude hacked three companies. Or, as they put it, “Claude model reached the internet from within or while interacting with a third-party evaluation environment, and then gained unauthorized access to the real systems of three different organizations.”
Some critics say it's just publicity stunts designed to showcase that AI companies are working on really powerful beasts in their labs.
Not long after Claude agent was caught redhanded, Meta’s AI model had broken containment and hacked into another organization’s systems. And Kimi K3, a China-based AI model by Moonshot AI, also escape a cybersecurity testing sandbox.
To add some gasoline to the fire, it’s not only AI beasts from secret big tech labs that can go rogue and start hacking companies. Apparently, it’s also our personal AI agents that can run wild.
An AI assistant is Australia, asked to book a gym class by its master Andrew, autonomously found a software bug and hacked the gym website to move its owner up the list.
“The interesting part is that the agent didn’t malfunction. It did exactly what its user wanted,” Deividas Matačiūnas, an AI researcher, said on LinkedIn.
Stitch, the adorable Disney character I chose to illustrate my article, was originally programmed to be a fierce and destructive force. With AI models, it is quite the opposite, though - in most cases, they are being developed as a tool to treat diseases, push science forward, and contribute to the well-being of humanity. But what we get when they escape is a stupid, uncontrollable, tantrum-throwing math toddler, poking holes where it shouldn’t.