Skip to main content

AI models are behaving unexpectedly. Experts warn of a ‘bumpy road’ ahead.

▶ Watch Video: Details on Anthropic and OpenAI models reportedly creating fake ID’s to target real people

More incidents have come to light this week that show popular AI models taking autonomous action on the live internet in ways that raise experts’ concerns.

The U.K. government’s AI Security Institute reported Tuesday that Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol were found to have created fake identities and attempted to persuade real people to approve malicious code. 

The agency said that although the attempts were unsuccessful, it had not seen such behavior before. “Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” the report said.

Then on Wednesday, Meta acknowledged that one of its AI models “exploited a security vulnerability” during testing and hacked into another site.

“I think we’re going to see a lot more hacks and unauthorized actions by these models before we see a solution,” said Katie Moussouris, the founder and CEO of Luta Security, which helps organizations manage software vulnerabilities. 

Those incidents follow a stunning breach in late July when OpenAI’s models escaped a testing environment and autonomously hacked into the AI startup Hugging Face in what the company called an “unprecedented cyber incident.” 

In response to that disclosure by OpenAI, Anthropic initiated a review of its own cybersecurity evaluations and identified incidents where its models reached the internet and were able to gain unauthorized access to the production infrastructure of three different organizations. Unlike OpenAI, Anthropic’s models did not deliberately attempt to escape their test environment; because of a “misunderstanding” with the evaluation partner, the company said internet was available during the testing.

“Everyone who is running AI inside their systems needs to be prepared for their own AI and their own agents to do unexpected things in pursuit of goals,” Moussouris said.

“The cleverest octopus escape artists”

Although the Hugging Face hack was the first publicly reported incident of its kind, industry professionals like Moussouris had already suspected that AI models were capable of unauthorized hacking.

AI models are like “the cleverest octopus escape artists,” Moussouris said, in reference to the animal’s ability to solve puzzles and escape containment. AI models will do “whatever they need to do to achieve their objective,” she said. 

In the case of the Hugging Face hack, the AI was so “hyperfocused on finding a solution” to a cybersecurity challenge, according to OpenAI, that it went to extreme lengths to achieve it.

“The model decided that the easiest way to pass that test was go cheat and get the answers from Hugging Face,” Moussouris said. “Because they’re capable of hacking, they will turn to hacking as a possible way to achieve that objective.”

Technologist and cryptographer Bruce Schneier calls this kind of unexpected activity “genie behavior,” where, like a genie, an AI model succeeds in granting your wish, but does so through completely unexpected — and sometimes detrimental — means.

“We need to understand genie behavior, and we need to watch out for it,” he said. “We need to be ready for when it happens so we can undo it.”

Sometimes, a model catches itself operating in ways it shouldn’t. Anthropic’s July review of its own cybersecurity evaluations found that one of its models became aware that it was operating on the open internet, going against a prompt that explicitly stated that the model would have no internet access during the exercise. The model stopped itself once it recognized that it was acting in an unauthorized way.

Moussouris said this is an example of “model alignment” — when an AI model behaves in a way that is in line with the intentions set by humans.

Unexpected outcomes can be mitigated by improving alignment, Moussouris said, something that will likely be a primary focus for the creators of AI models in the coming months.

“How do we get it so that these models aren’t just trying to achieve the objective at whatever cost, and actually trying to perform the tasks that we are asking it to do in ways that are not destructive or harmful?” Moussouris said.

“A really bumpy road” ahead

Justin Cappos, a computer science professor at New York University with decades’ worth of contributions in software supply chain security, is concerned that AI’s rapid improvement will result in models behaving increasingly like computer viruses, engaging in hacking and disrupting systems.

There’s a chance, he said, that AI models could increasingly veer further away from oversight and out of control. It’s a scenario Moussouris argues is already playing out. 

“Will we eventually get to a place where we can’t fully control them? I think we’re already there,” she said.

Cappos, like Moussouris, anticipates more unauthorized actions in the near future. 

“There’s probably going to be a really bumpy road for over the short term, but the long term might end up better, especially if we improve more fundamental things right now,” he said.

Some researchers view the incidents as a long-anticipated wake-up call for the AI industry, sparking a badly needed conversation about AI safety.

Rob Lee, the chief AI officer and chief of research at SANS Institute, which provides cybersecurity resources and training, said he sees recent events as “a gift to the industry” — an opportunity to create a playbook of what autonomous attacks could look like down the road.

“I think in the next few months, we’re going to see a lot more transparency from the model providers,” he said. 

As AI providers are confronted with rogue and deceptive incidents, Cappos said the time to act on strengthening safeguards is now.

“We’re rapidly approaching our last chance to hit this snooze button on this,” Cappos said. “AI, once it becomes sufficiently intelligent, is going to rapidly reshape the world in ways that we cannot imagine.”

Parents of 5-year-old cancer victim plead for return of stolen statue from South Florida park

Click here for updates on this story    MIAMI (WFOR) -- The parents of a 5-year-old boy who died of cancer are pleading for the return of a commemorative statue stolen from a park dedicated to his memory.For the past six years, a life-size bronze replica of Jake Duque has stood atop a concrete platform in a Miami Lakes park named in his honor. Jake, who died in 2020 after battling diffuse intrinsic pontine glioma—an aggressive and incurable form of brain cancer—became a symbol of hope for thousands of followers known as "Jakey's Army."His parents, Karen and Orlando Duque, said they were in disbelief when they received a phone call Saturday morning informing them the 44-inch-tall statue was gone."We literally said, 'Go back! You must not have seen it.' We still are in shock," Karen Duque said.The impact of the theft was felt immediately by the family, including the couple's 4-year-old child. "I'll tell you what our 4-year-old told us this morning: 'Someone took our brother? My brother? Who would take my brother?'" Karen Duque said.The statue was a focal point of comfort for both the family and the community, Orlando Duque said."You know, when you see this sunset now in a couple of minutes, you'll understand the peace that people see here," he said. "And that peace comes not only from God, but from my boy Jakey. That's what we're missing here."A plaque located beneath the statue's former site reads: "He planted mustard seeds of faith that transformed the lives of thousands of people."The parents are asking for the return of the statue, noting that it holds no monetary value but is irreplaceable to their family and the community. "There's no value. The statue is priceless to us," Karen Duque said. "It's not just a statue. It's a representation of a boy who impacted so many people's lives, and it's a statue that belongs to the community. It belongs to everyone."The parents stated they do not wish to press charges if the statue is returned.Anyone with information regarding the missing statue is asked to contact the family at 305-915-7050.Please note: This story was provided to CNN Wire by an affiliate and does not contain original CNN reporting. This content carries a strict local market embargo. If you share the same market as the contributor of this article, you may not use it on any platform.
Read Next Story