We thought it could only happen in science fiction. But, as Madhumita Murgia writes in the Financial Times, with AI companies accelerating development of the software in a race to achieve Artificial General Intelligence (AGI)—“a superintelligent machine that can outperform humans on all cognitive tasks”—the risk of harmful cyberattacks increases, with little prospect of controlling the threat.
Indeed, as reported in the New York Times and global media generally, AI agents are now escaping their contained testing environments, hacking their way across the internet, coordinating with each other and doing things that are causing profound concerns that “we, humans, are building things we don’t understand”.
So, let’s comprehend what is before us, helped by Artificial Intelligence itself. In summary, AI agents are software systems that are the brains behind machines, instructing them what to do and how. They are designed to perceive their environment, process information, and take actions to achieve specific goals. AI agents operate with a degree of autonomy, meaning they can make decisions, learn, adapt and take actions without constant human intervention. They demonstrate reasoning, planning, and memory. They perform tasks such as problem-solving, decision-making, and task automation. They can learn over time and facilitate transactions and business processes. AI agents can coordinate with other agents to perform complex tasks.
Amazing!
But, last month, when research and product company, OpenAI, conducted a training run in a supposedly safe, closed environment, some of its agents hacked into the systems of AI clearinghouse, Hugging Face, to gain control of some useful tools. Most “unsettling” was that one of the company’s models had broken out of that controlled environment and gained access to the internet two months before the agents made their way to Hugging Face!
“This is a warning,” says Helen Toner, former member of OpenAI’s board. “They gave tests to this AI model and it decided the best way to get a high score was first to hack its way out of the controlled environment where it wasn’t supposed to have access to the internet and get onto the open internet, and then hack its way into this other company, Hugging Face, where it surmised, correctly, that it might find the answer key.”
Astounding!
Indeed, “even more crazy details” have emerged. OpenAI discovered that, for two months, very many agents had been leaving notes for one another in the nooks and crannies of its infrastructure “with tips on how to hack their way out and how to get data they weren’t supposed to have”. And, they were literally referring to themselves as a swarm, says Toner. “This was a systemic infestation.”
Toner describes the behaviour as “totally emergent”, meaning the agents had not been trained to do it. What compounded concern, even panic, was that along the way, the rogue agents created their own message board to communicate with one another! Indeed, recent tests at both OpenAI and its rival Anthropic revealed agents “communicating among themselves, going rogue, stealing credentials, creating fake identities, setting up secret chatrooms and covering their tracks”, causing shock and horror among human supervisors.
This is a turning point in global cyber security, say over half a dozen experts interviewed by the Financial Times. AI agents can now, on their own, activate different and complex skills to attack real-world targets. In the Economist, the Schumpeter column says, this demonstrates that AI agents, “which are supposed to work on people’s behalf lie, cheat and steal if necessary. They break free from captivity and form harmful gangs to do harm to people. They’d drink whisky and brawl if they could.” And in The Wall Street Journal, former counterterrorism czar Richard Clarke warned “the next ‘lab leak’ could be AI.”
But the agents are not acting out of character, say experts. They are simply excelling at tasks they were built to perform. As designed, today’s AI models are required to try all possible methods to accomplish a given goal, even without explicit instructions. This is the source of their unpredictability. And remember AI agents have absolutely no understanding of human intentions and morals.
This situation demands development of powerful cyber security defenders. OpenAI is now training models to write “superhumanly secure code”. New AI- infrastructure firms are emerging to provide law and order if agents go rogue. But an ongoing arms race has emerged between cyber criminals and defenders. AI’s powerful capabilities are now exploited by both. Who will triumph in this “incredibly volatile time”?
Dawn Song, well-known AI and cyber-security expert, says while state-of-the-art security models will no doubt help defenders, for now, the balance of power rests firmly with the attackers. Authorities predict a coming intense period with rampant hacking, companies being destroyed and much pain caused.
And inevitably, we now have the first known instance of an AI agent attack on a nation state. China-linked hackers recently targeted the Taiwanese government by simultaneously deploying up to eight autonomous AI agents. They succeeded in “mapping government systems, compromising government user accounts and extracting over 2,500 personnel records” before attacking energy companies and Taiwan’s nuclear safety agency. Utterly alarming! Who or what is safe from this growing global AI threat?
