Skip to main content

AI models are behaving unexpectedly. Experts warn of a ‘bumpy road’ ahead.

▶ Watch Video: Anthropic’s Mythos AI created fake profiles to trick people in hack, then hid evidence

A cybersecurity report from the U.K. government has exposed new examples of popular AI models taking autonomous action on the live internet in ways that raise experts’ concerns.

The AI Security Institute’s report, released Tuesday, said Anthropic’s Mythos 5 and OpenAI’s GPT-5.6-Sol were found to have created fake identities and attempted to persuade real people to approve malicious code. 

The agency said that although the attempts were unsuccessful, it had not seen such behavior before. “Some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations,” the report said.

“I think we’re going to see a lot more hacks and unauthorized actions by these models before we see a solution,” said Katie Moussouris, the founder and CEO of Luta Security, which helps organizations manage software vulnerabilities. 

That follows a stunning breach in late July, when OpenAI’s models escaped a testing environment and autonomously hacked into the AI startup Hugging Face in what the company called an “unprecedented cyber incident.” 

In response to that disclosure by OpenAI, Anthropic initiated a review of its own cybersecurity evaluations and identified incidents where its models reached the internet and were able to gain unauthorized access to the production infrastructure of three different organizations. Unlike OpenAI, Anthropic’s models did not deliberately attempt to escape their test environment; because of a “misunderstanding” with the evaluation partner, the company said internet was available during the testing.

“Everyone who is running AI inside their systems needs to be prepared for their own AI and their own agents to do unexpected things in pursuit of goals,” Moussouris said.

“The cleverest octopus escape artists”

Although the Hugging Face hack was the first publicly reported incident of its kind, industry professionals like Moussouris had already suspected that AI models were capable of unauthorized hacking.

AI models are like “the cleverest octopus escape artists,” Moussouris said, in reference to the animal’s ability to solve puzzles and escape containment. AI models will do “whatever they need to do to achieve their objective,” she said. 

In the case of the Hugging Face hack, the AI was so “hyperfocused on finding a solution” to a cybersecurity challenge, according to OpenAI, that it went to extreme lengths to achieve it.

“The model decided that the easiest way to pass that test was go cheat and get the answers from Hugging Face,” Moussouris said. “Because they’re capable of hacking, they will turn to hacking as a possible way to achieve that objective.”

Technologist and cryptographer Bruce Schneier calls this kind of unexpected activity “genie behavior,” where, like a genie, an AI model succeeds in granting your wish, but does so through completely unexpected — and sometimes detrimental — means.

“We need to understand genie behavior, and we need to watch out for it,” he said. “We need to be ready for when it happens so we can undo it.”

Sometimes, a model catches itself operating in ways it shouldn’t. Anthropic’s July review of its own cybersecurity evaluations found that one of its models became aware that it was operating on the open internet, going against a prompt that explicitly stated that the model would have no internet access during the exercise. The model stopped itself once it recognized that it was acting in an unauthorized way.

Moussouris said this is an example of “model alignment” — when an AI model behaves in a way that is in line with the intentions set by humans.

Unexpected outcomes can be mitigated by improving alignment, Moussouris said, something that will likely be a primary focus for the creators of AI models in the coming months.

“How do we get it so that these models aren’t just trying to achieve the objective at whatever cost, and actually trying to perform the tasks that we are asking it to do in ways that are not destructive or harmful?” Moussouris said.

“A really bumpy road” ahead

Justin Cappos, a computer science professor at New York University with decades’ worth of contributions in software supply chain security, is concerned that AI’s rapid improvement will result in models behaving increasingly like computer viruses, engaging in hacking and disrupting systems.

There’s a chance, he said, that AI models could increasingly veer further away from oversight and out of control. It’s a scenario Moussouris argues is already playing out. 

“Will we eventually get to a place where we can’t fully control them? I think we’re already there,” she said.

Cappos, like Moussouris, anticipates more unauthorized actions in the near future. 

“There’s probably going to be a really bumpy road for over the short term, but the long term might end up better, especially if we improve more fundamental things right now,” he said.

Some researchers view the incidents as a long-anticipated wake-up call for the AI industry, sparking a badly needed conversation about AI safety.

Rob Lee, the chief AI officer and chief of research at SANS Institute, which provides cybersecurity resources and training, said he sees recent events as “a gift to the industry” — an opportunity to create a playbook of what autonomous attacks could look like down the road.

“I think in the next few months, we’re going to see a lot more transparency from the model providers,” he said. 

As AI providers are confronted with rogue and deceptive incidents, Cappos said the time to act on strengthening safeguards is now.

“We’re rapidly approaching our last chance to hit this snooze button on this,” Cappos said. “AI, once it becomes sufficiently intelligent, is going to rapidly reshape the world in ways that we cannot imagine.”

Reportan casos de ciclosporiasis relacionados con la lechuga en 15 estados de EE.UU.

Un gran brote de ciclosporiasis en varios estados, que se relacionó con la lechuga vendida en restaurantes como Taco Bell y algunas tiendas de alimentos, registra casos en 15 estados, según funcionarios de salud federales.Seis estados —Arkansas, Iowa, Missouri, Nebraska, Nueva Hampshire y Carolina del Norte— se sumaron al brote luego de que pruebas de laboratorio realizadas por los Centros para el Control y la Prevención de Enfermedades de EE.UU. (CDC, por sus siglas en inglés) mostraron que probablemente las enfermedades en esos estados fueron causadas por la misma cepa del diminuto parásito de la cyclospora, que se aloja en la mucosa del intestino delgado y puede causar semanas de diarrea explosiva si no se trata.De los más de 6.000 casos contabilizados por los CDC en este brote, casi 300 requirieron hospitalización y dos personas murieron.La actualización confirma lo ampliamente extendido que ya estaba el brote, dijeron expertos, en lugar de ser un indicio de que su alcance se está expandiendo.“Simplemente estamos contando a las personas que ya estaban enfermas. Ahora viene la vinculación con Taylor Farms u otros productos”, dijo Bill Marler, un abogado con sede en Seattle y defensor de la seguridad alimentaria.Los CDC implementaron una respuesta de emergencia de nivel 3 al brote, según una fuente informada sobre las acciones de la agencia. El nivel 3, el más bajo en la respuesta de emergencia, significa que la agencia está monitoreando y trabajando en una situación activamente, aunque el riesgo para el público en general es bajo. Esto le da a los CDC la flexibilidad de reasignar personal para ayudar en la investigación del brote.En entrevistas detalladas, muchas personas enfermas informaron que consumieron lechuga, a menudo lechuga iceberg rallada suministrada a restaurantes como Taco Bell por el gigante de productos agrícolas Taylor Farms.Taylor Farms de México retiró toda la lechuga iceberg cultivada en el centro de México, incluidos los kits de ensalada vendidos en tiendas de comestibles, el 17 de julio. En un comunicado en su sitio web, la empresa dijo que realizó una retirada general de sus productos “por precaución”.La empresa además señala que está llevando a cabo una revisión “completa” de todos los procesos y protocolos de seguridad.Taco Bell dice que también retiró la lechuga iceberg rallada afectada suministrada por Taylor Farms de sus restaurantes en ciertos estados.Los expertos creen que debido a que la lechuga es perecedera, es probable que la comida que enfermó a las personas haya sido ingerida o desechada hace tiempo.Se considera que la lechuga es la principal sospechosa en el brote en varios estados, pero podría no ser la causa de las enfermedades en otros estados. Debido a que algunas personas enfermas no informaron haber consumido lechuga y los funcionarios de salud estatales dijeron que no pueden descartar otras fuentes contaminadas como contribuyentes.En su última actualización, los CDC identificaron más de 22.700 casos de ciclosporiasis doméstica en esta temporada; 10.468 casos están confirmados y más de 12.200 se encuentran en proceso de pruebas de laboratorio adicionales y análisis para determinar si la enfermedad se adquirió en EE.UU.Según el recuento de Marler, los estados reportaron cerca de 25.000 casos de ciclosporiasis hasta la fecha en esta temporada.Debido a que muchas personas pueden optar por no visitar al médico por diarrea, los funcionarios creen que es probable que haya muchos más casos no contabilizados. El recuento actual “es una base y no una estimación”, dijo Marler. El verdadero número de estadounidenses afectados por ciclosporiasis este verano podría ser de hasta 1 millón, según un conteo, publicado por los CDC en 2011.La Administración de Alimentos y Medicamentos de EE.UU. investigó siete grupos de casos relacionados de ciclosporiasis este verano. En cinco de esos brotes, la agencia define la causa como “no identificada”.The-CNN-Wire™ & © 2026 Cable News Network, Inc., a Warner Bros. Discovery Company. All rights reserved.
Read Next Story