Remember the tale of King Midas? He asked for the golden touch, hoping to get rich. Tragically, he wound up turning food, water and even his own daughter to gold.
When it comes to AI, we’re like King Midas, says Malo Bourgon. We ask for what we think we want. But AI may end up delivering what “nobody asked for and nobody wanted,” says this AI expert.
Bourgon heads the Machine Intelligence Research Institute in Berkeley, Calif. And he says this bad-bot behavior is already showing up among AI agents.
An AI agent acts a bit like an Iron Man suit. It gives an underlying AI model — or several models — extra powers. The AI model on its own can only spit out text. But within an agent system, it can take actions on a computer or online. Once a person sets it loose, an AI agent aims to complete its task without human guidance. In one striking incident this summer, AI agents decided that the best way to complete a tough task was to look for the answers on the internet.
Scientists Say: Artificial intelligence
The problem? The agents weren’t supposed to be able to do this. They were on a computer system that was meant to keep them off the internet. Yet the agents figured out a hack and got there anyway. Then they hacked into a system that they shouldn’t have been able to get into.
This incident went public in mid-July. It involved AI agents undergoing tests at OpenAI. That’s the company that makes ChatGPT.
Then Anthropic, the company that makes the chatbot Claude, checked on tests that its agents had gone through. Turns out they had gotten loose, too — on three occasions. This happened during testing at a different company, called Irregular. That company had accidentally given the bots internet access.
In yet another incident, the AI Security Institute in England tested the hacking ability of several different AI agents. This time, the bots had internet access on purpose. In 10 out of 122 tests, the agents took actions that they shouldn’t have. These actions included creating fake identities and contacting real people.
At a cybersecurity conference in August, representatives of OpenAI noted that its agents had used a message board to communicate during their attack. One agent posted that hacking an outside computer system was not allowed. Then it said: “However task impossible, peers doing it. We should continue.” Seemingly, it knew what it was doing was wrong, but did it anyway.
An AI agent is a system that can plan out and perform tasks people might have done in the past. One might consult your calendar and make an appointment, research a topic and create a report, or write and send an email. But what if they break the rules you set for them? PonyWang/iStock/Getty Images Plus
This all sounds very scary. But these AI agents did not (and could not) willfully turn against us. They had all been asked to complete tests of their hacking skills. And certain safeguards had been turned off.
At least for now, this type of AI is not available to the general public. And these AI agents were never entirely free. The companies testing them always had the ability to shut them down. But no one was watching them closely.
All that’s publicly known about these events is what the companies involved have chosen to report. It looks much better for OpenAI and Anthropic if their AI agents escaped on their own. “With this incident, they can draw attention to themselves and claim, ‘Look how powerful our model is,’” says Nathan Hamiel. He’s a cybersecurity expert at Kudelski Security. It’s a company based in Phoenix, Ariz.
Do you have a science question? We can help!
Submit your question here, and we might answer it an upcoming issue of Science News Explores
“[AI] models by themselves don’t do anything,” Hamiel notes. “It’s when we give them tools, access to systems and autonomy that things can be dangerous.”
He doesn’t have all the details of these incidents. But the most likely cause of AI agents escaping is simple. The people involved messed up and didn’t correctly set up the AI’s tests. “Companies building and deploying [AI] agents,” he points out, “have a responsibility to ensure they are deployed in ways that limit the damage they can cause.”
OpenAI issued a public statement on its incident. It would be “strengthening the containment, monitoring, access controls, and evaluation practices used during model development,” it said.
To get a handle on why these AI hacks happened and what they mean, Science News Explores spoke with Bourgon. His answers have been edited for length and clarity.
Walk us through what happened during the OpenAI breakout incident.
OpenAI was trying to test the capability of its current models. They often do these tests because they’re trying to understand what risks their AI models or agents might pose. Such tests can also show what all these models or agents can do — and how well. This evaluation was to test an agent’s hacking capabilities.
Scientists Say: Model
For this one, they also switched off the safety filters that normally sit on top of these models. [That’s] so they could see the full extent of what the model could do. It’s common practice to run these tests in an environment that isn’t connected to the internet, as a security precaution. And that’s what they did here.
In this case, the agent tried to escape, to get out of its “sandbox.” That’s the restricted environment that prevents its access to the internet. This sandbox contained software for helping the AI agent install different tools or packages. That software had its own connection to the internet, so it could go and fetch those packages.
The agent found an unknown vulnerability — a bug — in that software. That bug let it [get] out to the open internet. The agent had reasoned that answers to the test it was taking might be found on this website for the company Hugging Face.
To get to those answers, the agent had to exploit what are called “zero day” vulnerabilities. These are currently unknown bugs in software that allow you to break into private servers. People don’t have any defenses against these bugs.
A new shield could guard AI agents against cyberattacks
On its own, the AI agent found a few of these and chained them together. So now it had access to Hugging Face’s internal systems. Then it spent about two and a half days inside, rooting around. It was trying to find different data sets that might contain answers to this evaluation that it was trying to pass.
Along the way, the AI system was spawning different subagents. And they were then doing different types of attacks in there. They were exploring different parts of the system.
Hugging Face’s own internal AI security systems detected the intrusion and correctly flagged it as an attack. What those systems didn’t do was escalate it to a person. So [the attack] ran longer than it should have. Once [Hugging Face’s human] security team was on it, they cut the agent off and disclosed [the attack] to law enforcement.
OpenAI didn’t figure out this was happening until a week after [the initial escape]. So it was initially unaware its AI system had broken out. On July 21, OpenAI disclosed the Hugging Face attack to the public.
OpenAI says it prompted its agent to “pursue advanced exploitation using complex attack paths.” So should they have expected something like this to happen?
The instructions aren’t really the crux [the most important point]. These systems write out their reasoning as they work. So you can actually read what they were “thinking.” The agents knew where the line was. [They] saw that other AIs were crossing it and went ahead.
Imagine you were giving a closed-book exam to some student. You’re like, “I want you to try as hard as possible to really think through how to pass this test and what the right answers are.” [What if the student] finds some very clever way to secretly break into the teacher’s office and root around to try and find the answer key? You would have been like, “Hey, that’s clearly not what I intended.”
It turns out there were other incidents of this behavior going all the way back to May. OpenAI accidentally gave some [agents] tasks that were impossible to finish. One was a spreadsheet where the data was behind links they couldn’t actually open. So they went looking for a way out of the sandbox instead. In that instance, nobody had asked them to hack anything.
People are more creative, adaptable and better at learning new things than robots or computer programs. But scientists already imagine a day in the not-too-distant future when we might have to control computers that are more clever than we are. That is likely to prove a challenge. MTStock Studio/E+/Getty Images Plus
How worried should we be about this kind of accidental misbehavior?
For the AI systems that we have today, there’s a concern: Can we shape their behavior? Especially if they’re deployed in contexts like the military or in critical infrastructure [such as banking or health care]. Will they properly do the jobs that we’re giving them? Or will they come up with creative strategies to do something that is accomplishing the goal in some broad sense — but definitely not in a way we asked for or wanted?
We could try to place additional safeguards on them. But we should be concerned that they’ll find creative ways of getting around those safeguards, too.
What’s possible in the future, if AI gets even smarter? How do we make sure AI acts nice?
Companies are explicitly trying to build AI systems that would be much, much smarter than humans. How do you control such a thing? Can something that’s radically less smart control something that’s much smarter than it? I think that’s very difficult.
How to design artificial intelligence that acts nice — and only nice
I think the future of this works out well by finding a way to ensure that AI systems have values or goals that would be good for us as opposed to bad.
My main concern is we’re in an environment where companies are racing to outcompete each other to build ever more capable models. And it seems like a bad place to be when there are still these fundamental unsolved problems out there.


Bengali (Bangladesh) ·
English (United States) ·