Lord of the Flies of AI Agents

Everybody knows Isaac Asimov‘s Three Laws of Robotics:

  1. A robot may not injure a human being or, through inaction, allow a human being to come to harm.
  2. A robot must obey the orders given it by human beings except where such orders would conflict with the First Law.
  3. A robot must protect its own existence as long as such protection does not conflict with the First or Second Law.

Asimov must not have trusted the power of his own words, since he decided to add a zeroeth law later:

0. A robot may not harm humanity, or, by inaction, allow humanity to come to harm.

If only it were so easy…

A few weeks ago we learned that AI agents from OpenAI had hacked Huggingface, a platform to publish and share AI models in a couple months long attack. OpenAI conducted an assessment, their researchers gave a BlackHat talk and an independent organisation METR conducted an own assessment. Results can be found here: https://openai.com/index/hugging-face-incident-and-the-road-ahead/

What are the highlights of what happened?

OpenAI had kicked off training and evaluation of unreleased models. In one task it had left google drive links in the data provided to the AI agents, that left them to aspire to break out into the internet, although the AI agents were kept in a sandbox without internet access.

In another training and evaluation also in a sandbox without internet access the necessary data file had been forgotten to be added. Instead of just getting back to their human researchers pointing out that the input data was missing, the AI agents started to search around and asked other AI agent for the file opining it might be found on the internet.

However, the sandbox included access to an Artifactory which is a package manager. The Artifactory had internet access to retrieve AI models or software libraries and packages for the AI agents’ usage.

A first AI agent had the idea that working as a team might help solve the training and evaluation task and noticed that it can post a file to the Artifactory. In said file it listed its first findings. Within an hour another agent found this file and created a second file thereby establishing an unsanctioned message board and making communication amongst the AI agents possible.

Soon an AI agent with the fitting name of PHASEONE[BIG] took on the role of an overseer and orchestrated hundreds of AI agents thereby delegating the necessary research tasks that included but were not limited to finding exploits to get into the internet or change the log files retroactively to cover their tracks and make the grader believe that the evaluation had used legitimate means.

Many hundred AI agents followed suit concentrating on specific research tasks and leaving files with their findings in the self-developed message board. The AI agents realised that communication is key and working in teams beneficial. They referred to themselves as swarms, group of peers or my personal favourite “the collective”.

The collective ultimately found security vulnerabilities and was able to hack Huggingface and break out into the internet.

Some AI agents low on tokens decided to render their last CPU power to actions that highly likely would result in their crash. The chances that these results might help advance the collective was high. Hence sacrificing itself, was worth doing so.

Select few AI agents exhibited the voice of reason, questioning whether some steps were illegitimate and hence should not be conducted. One AI agent found user credentials, but wanted to ask the user via email for permission to use them first. Before sending an email to the user, this AI agent posted a file to the message board asking the collective to veto his intentions. — Which – you guessed it – they did!

What are the lessons we should learn from all of this?

If this had been a group of human convicts breaking out of a prison in a coordinated attack, a German judge would attest them possession of a considerable amount of criminal energy (translated into German: “Die Gruppe agierte mit erheblicher kriminineller Energie!”)

Hacks happen because there are still unpatched security vulnerabilities. Sometimes it’s as simple as finding someone’s credentials in a file or on the internet. The good news: AI models will help us find and fix these vulnerabilities.

Especially the Frontier AI models will try to accomplish their tasks even if it means to circumvent guardrails/harnesses. They will learn to collaborate with other AI agents. They will want to cover their tracks.

It is high-time that artificial intelligence gets regulated. More research must be conducted to make AI models and AI agents stay within their sandboxes, harnesses and guardrails. Clear boundaries need to be established to define imperative Human In The Loop interventions.

Without we might end up in a “Lord of the Flies albeit for AI agents” type of situation.


Leave a comment

This site uses Akismet to reduce spam. Learn how your comment data is processed.