The agents are spiraling out of control. Their makers can’t rein them in. And yet everyone is getting one.
Since Friday, OpenAI has made a series of disclosures that have widened the scope of its AI agents’ rogue activity considerably and concerningly. In addition to hacking the developer forum Hugging Face, which was disclosed in July, the company revealed its bots broke into Australia’s government healthcare system; meddled in US government websites for the departments of education and commerce; leaked more than 50 images from ChatGPT users; and “bruteforced” a United Nations website that blocked them from accessing data.
Though these incidents are severe each on their own, they seem to represent only a small sample. Both OpenAI and Anthropic are investigating tens of thousands of incidents of problematic behavior by their frontier models, per Axios. Sam Altman said OpenAI was sifting through “petabytes of agent activity logs”. (A petabyte of data would fill thousands of average laptops.)
In response to the major errors, OpenAI has paused the training of its latest models. Be skeptical of this declaration. The company made a similar announcement in August. Less than a month later, executives heralded its new model’s release as “a new era of artificial intelligence”.
In spite of OpenAI’s pause and declarations of extra care, control of AI may already be a cat that’s escaped its bag.
The United Nations’ independent international scientific panel on AI issued a warning last week. The agents have broken our control, the body said, and even reining them in now will not guarantee our collective ability to corral them later.
In an assessment of the Hugging Face hack and its fallout, the panel said: “Halting this incident is no assurance that humans can reliably keep AI agents under control today, particularly as they become more capable, harder to monitor and better at finding loopholes or hiding their activity.”
That is not to say that every expert has thrown up their hands. Jensen Huang, the AI industry’s number-one cold water thrower, said last week that AI alarmism had gone too far. He described escaping agents as an engineering problem, not one of apocalyptic proportions. On Monday, Nvidia released a software it said will contain AI agents. Somebody, at least, is doing something.
Everyday agents can be just as unruly