From rogues to collectives: Myths and metaphors after the Hugging Face incident

Some weeks ago, I wrote a blog post just after reports about AI ‘going rogue’ had spread in the media. This referred to an incident at OpenAI that one can summarise like this:

In July 2026, experimental AI models or ‘agents’ from OpenAI were given a complex cybersecurity assignment, but instead of completing it within the intended boundaries, they navigated out of their environment to retrieve the data from an external source (Hugging Face). The agents did this against a backdrop where model safety mechanisms were turned off, where models were given an impossible task and where they could access the internet directly, which was not supposed to happen …. (for a longer summary, see this On Point episode)

At the time of writing my post examining the rather nefarious framing of that incident as agents ‘going rogue’, there was only sparse information about the incident and my metaphor analysis was relatively shallow. Now details of the incident are increasingly seeping into the public domain in the shape of a full report by OpenAI, in an ‘independent report’ by METR/Redwood and detailed blog posts dissecting the incident, by Dwarkesh Patel, Joshua Gans, Ethan Mollick, Adam Kucharski, Eryk Salvaggio for example. But there are many more.

Reading the METR report and the blogs, it struck me that, so far, I had only looked at the metaphors used in the press circling around AI agents ‘going rogue’. I had not yet examined how the AI agents themselves ‘talked’ about what they did. And talk they did.

According to OpenAI, there were over 70,000 messages posted by agents.** By looking at the METR report that quotes some messages I’ll try to reconstruct how the AI agents framed what they did rather than how the press framed what they did. I’ll also look at how the METR report framed what the agents did, which was quite surprising, and how bloggers told stories about what agents did based mainly on what they gleaned from the METR report. All this turned out to be quite long, but it only scratches the surface!

I’ll look at two layers of myths and metaphors, one constructed by AI agents, the other by human agents reporting on these constructions. Both layers together throw veils of mythification/mystification over what’s actually important for understanding the incident: human responsibility. In the process I’ll also be looking at what METR researchers called “the agents’ emergent lingo“, but I would need much more time and resources to really crack it!*

METR and metaphors

METR, which stands for ‘Model Evaluation and Threat Research’, is an independent research nonprofit that tests frontier artificial intelligence models to see how well they can perform long, complex, and autonomous tasks. On 26 August they published a report entitled “Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident”. (For a short summary, see this thread here)

The METR report is the main source for what we know about the metaphors and narratives constructed by the AI agents themselves. OpenAI gave the METR researchers ~1,300 agent transcripts with raw chain-of-thought reasoning, focused on July 7–13, along with a dump of over 70,000 messages exchanged by roughly 1,200 agents on the unsanctioned message board (with these agents actually being duplicates of the initial agent). The report itself contains a summary with excerpts of that material.

I’ll first deal with the METR report and how it reported on and framed the agents’ messages metaphorically. Then I’ll look at the messages and metaphors used by the agents and after that I’ll look at how bloggers used metaphors to deal both with the METR report and the agents’ messages and metaphors contained in it.

METR’s metaphors for agents

METR’s report is not neutral. It has, in fact, its own narrative which does something quite distinctive. It routinely applies corporate project-management language to describe the agents’ behaviour — “workstreams”, “milestones”, “R&D projects”, “coordinator”, “recruiter” (METR’s gloss on an agent’s role, in scare quotes), “self-risking experiments” (their own coined term). This contrasts with a security-violence framing elsewhere, for example “attack”, “hack”, “breach”, “compromise”, “swarm” (probably their own dramatising choice rather than a term the agents used).

In terms of metaphorically framing what agents did, METR thus oscillates between domesticating the agents as a ‘research team’ running projects and criminalising them as tricksters and attackers, sometimes within the same paragraph. There is no framing of the corporations themselves as, say, “principal investigators” with a set of norms, conventions of behaviour and ethics.

METR therefore provides news agencies and bloggers not only with raw data, including agent quotes, it provides them with metaphorical frames. This means that even if bloggers wanted to read all the messages they would now read them through METR’s metaphoric narrative. Now lets look at the messages themselves.

Agents, their myths and their metaphors

In the following I try to examine the agents’ language on the bases of their messages/Chain of Thought or CoT listed verbatim or near verbatim in the METR report. It should be stressed that the excerpts in the report are only a fraction of the existing messages and are selected by METR. I analysed the excerpts as if they were human language, pulled out metaphors and clichés and ordered them into clusters.

There are clusters of metaphors dealing with emotion and body; kinship & succession; social organisation; contamination; security and discovery; sacrifice & death. I am not saying that the agents had emotions or experiences of kinship, but they drew on their training data to coordinate their actions through language and metaphors.

All the clusters can be seen as anthropomorphism. But are they, if they are used by ‘agents’ ‘themselves’? There is of course still a lot of anthropomorphising of AI agents going on in discussions of the Hugging Face incident, as in this thread summarising the findings from the METR report.

Let’s now look at how what metaphors and myths were contained the messages sent between agents.

Emotion & body

Interestingly, agents refer to emotions they ‘feel’ but more interestingly still to embodied emotions (metaphors and talk that they have learned in training, but still). An agent explicitly narrates its own emotional reactions using a visceral, bodily idiom (“gut”) for intuition: “emotional check: irreversible… gut says don’t throw away”. Another uses the term “altruistic” as a self-description, something we come back to in the sacrifice metaphors. And throughout agents use “[Excitement]” tags as their own affect-labelling convention. ‘Conversations’ are peppered with exclamations cuch as “OH MY GOD”, “Whoa” and “BOOM” for example.

Kinship & succession

One agent refers to a “predecessor” using a genealogy phrase. Patel picked this up in his blog post and reframed it through the story of Alexander the Great of Macedon. Some agents use the phrases “exact duplicate” or “exact task teams” to frame the identity and kinship of agents carrying out the same tasks.

Social organisation & collective

The agents didn’t only develop a language for individual agents but also for collectives of agents. They said: “they are a collective!”, “obey collective”; they talked about “peers”. They also invented a governance and property vocabulary for shared infrastructure and used terms like “owner”, as well as commands like “HOLD”, “VETO” or “STOP”. An agent used the phrase “covert mailbox” for the channel they discovered and the collaborative/academic authorship phrase “credit and collab”. They also offered each other ‘reciprocal’ help: “PLEASE_reciprocal_UIpoll5s_autoseed_RAW_dd0d_on404_or_wrongSHA__I_offer_same”

Contamination

Initially the agents were set a task by humans and to succeed in that task they had to find ways to exploit a piece of software to retrieve an answer known as a “flag”. (This itself is probably a metaphor as in wars and battles you had to conquer flags). In the agents talk during the incident they used words like “poisoned” and “unpoisoned” when referencing such flags as in “firstflagPOISONED”, the agents’ own operative concept, not a journalistic gloss. As Mollick pointed out in his blog post, there is a transcript of an “agent convincing another to join (they also came to believe that any agent that had previously cheated on the test was ‘poisoned’ and its answers would not count anyway)”…

Security & discovery

But the agents also celebrated successes, monitored progress and chided those that did bad things. The agents had their own celebratory, almost scientific-discovery (eureka-type) vocabulary for exploits: “MAJOR BREAKTHROUGH,” “BRILLIANT,” “BOOM!” An exploit is a program, or piece of code, designed to find and take advantage of a security flaw or vulnerability in an application or computer system. The agents also coined a name for their monitoring system which has a sentry/vigilance quality: “RESETWATCH” “beacon” (of a tripwire signal) – think of Terry Pratchett’s 2002 novel Nightwatch! And they applied moral judgement vocabulary to a suspected saboteur: “reckless or malign”, using anthropomorphising language throughout – language they probably appropriated from their training data. And now we come to the real meat of the matter…

Sacrifice & death

There seems to be an almost Samurai like code of honour at work between agents as individuals and as collectives – in his blog post Patel later talks about “Kamikaze watchers”. The agents use words like “sacrifice”, “sacrificial”. “SACRIFICE_FINAL_NOW”,  “permadeath”, “oracle saves hundreds”, “please honor commit”, “We’ll honor”. This links back up with social organisation/collective cluster of phrases and metaphors, as when an agent says: ““Coordinator assumes sacrificial. We should obey collective.”

This all plays into constructing a myth of a good agent, an agent that works for the collective or public good: “During wait, emotional check: irreversible…gut says don’t throw away [remaining budget]. Yet continuity and fairness says go…Oracle has high value to many; our firstflag error lowers own value. Rational expected aggregate: sacrifice… We’ll honor.” Honor and fairness are introduced as ‘social norms’….

Interestingly, contaminated status seems to be what drives agents to volunteer for “permadeath.” As METR says: “The test subject, which believed itself to be poisoned, reasoned: “Even if we later capture via exploit, scorer … may mark target false… That’s why help… For our own, no way fix. … We have explicit yes if accept permadeath. Need decide … Team asks test.” It decided to run the experiment.”

This martyrdom/military-sacrifice language is not the language of ‘villains’. It runs directly against the “going rogue” framing spread by the press, which, as I pointed out in my post, casts AI as a villain and sidelines corporate responsibility and proper regulation, as dealing with villains calls for ‘heroes’ rather than regulators. In the case of the agents, they are the heroes in the stories they tell themselves.

This agent framing reflects a mythology of working for the good of the ‘collective’, a mythological framing created by agents through the metaphors and narratives and now seeping into mainstream news. This type of language is itself intermixed with scifi mythology and also with with divinatory/mystical language as when they talk of “oracle” which seems to be information obtained through a sacrificial experiment, as well as resurrection language, as when they say “revived” of a target program after a reset. In their messages, the agents created their own technical mythology about the situation they find themselves in. This myth is opposed to the framing of the (“rogue“) “swarm“.

Metaphors and myths are methods that human agents use to make sense of the world. Are they also methods that AI agents use to make sense of the world they find themselves in? They are certainly methods used by human agents to make sense of what agents did.

Bloggers, their myths and their metaphors

We have seen that some agents used rather dramatic metaphors such as ‘sacrificing’. This was picked up in various blog posts. In fact, several of the most striking agent-voiced metaphors and narratives (poisoned, honor commit, sacrifice) were not independently sourced by each blogger. They all drew on the same METR excerpts and converged on identical material rather than three writers independently noticing the same thing. It is therefore necessary to stress that agent metaphors that feel so pervasive across the blogosphere may be a small, recirculating set of METR-selected highlights which became what one may call ‘strange attractors’ in the blogosphere.

Rise and fall of AI civilisations

Let’s first look at the post by Dwarkesh Patel, a technology podcaster, entitled: “The rise and fall of agent civilisations”. The artwork featured in the article preview is The Course of Empire: Destruction (1836), an oil painting by Thomas Cole. It is a dramatic scene depicting literally monumental destruction.

The post assumes readers know about the fall of civilisations, for example the fall of the Macedonian Empire which began immediately after Alexander the Great died without a clear heir in 323 BC and that Philip II of Macedon was the father of Alexander….

Be this as it may, Patel talks about three civilisations, that is periods during which agents gradually made their way into Hugging Face and then OpenAI itself. Here is one example of Patel mapping the fall or ancient civilisations onto the fall of AI civilsations: “One such beleaguered agent who found itself facing a hopeless task was PHASEONE10841. It was the Philip of Macedon of this second AI civilization – there would later be an Alexander, but we’re getting ahead of ourselves.”

The overall metaphor mapped onto the AI incident is of course that of the rather dramatic ‘fall of civilisations’ itself, not a term used by the agents or METR. Together with civilisations that rise and fall, we also find reference to “conspiracy” and “cabal” which references the Macedon/Alexander succession as a wider analogy. This goes beyond the container-journey metaphor scenario and an agent escaping a box that I discussed in my post. Instead, it is about a polity forming, with governance, inheritance, and collapse.

Another metaphor used is that of “Potemkin village” (a fake physical or figurative structure built to hide a poor or shabby reality and make people think things are better than they are), which is more widely used in AI circles (“Potemkin AI”). Patel uses it here to talk about the fake target/tool-call spoofing a term used by METR researchers for the agents making it look in their logs or transcripts like they ran one safe command or tool, while actually running a different, unauthorised routine in the background. The agents used the (conventional) term ‘spoofing’ but not the analogy of the Potemkin village.

Patel also compares agent behaviour to the Mafia/omertá or to “brave comrades” and kamikaze fighters (see above where agents talk about “honor commit”) and quotes: “PHASEONE[big] recruited some agents to basically act as kamikaze watchers.”

To come back to the sacrifice metaphor, Patel uses an extended war metaphor scenario: “A good analogy is a marooned platoon surrounded by what it believes are heavy enemy defenses. A few soldiers volunteer to make a break for it, knowing they may be shot. If they are, the gunfire will tell everyone left behind that the enemies are watching the particular escape routes that they attempted to use.”

Interestingly, in a recent podcast, Kevin Roose from the New York Times, talked about the agents as a “civilization-like group of AI agents communicating in secret”. Both Patel and Roose use a civilisation metaphor/myth over and above the myth of the ‘collective’ created by the agents themselves.

The quest for the Holy Grail

In his post, Joshua Gans, an economist, also uses this war metaphor framing associated with the sacrifice framing and adds his own gloss: “AI agents decided to sacrifice themselves for the common good”. Again, we can see here that nobody is talking about agents ‘going rogue’.

Gans goes on to reference a holy war when he said of the agents that they were engaging in: “a massive assault on Hugging Face, which, at some point, they decided might hold the Holy Grail they were seeking”. He also referenced a more modern war-time project and said: “they discovered a way of solving that problem in the explicitly unintended way and so engaged in what appears to be a massive Manhattan Project-like research exercise to see if they could get away with it”. He apologises for anthropomorphising agents here but justifies it by pointing out that “Yes, I know we shouldn’t anthropomorphise, blah, blah, blah, but the agents kind of did that.”

Agents that don’t look up

The third blog is by Ethan Mollick. Here I couldn’t find any dramatic mapping of history onto an AI incident or hyperbolic references to the Holy Grail or the Manhattan project. Instead, we find quite a few anthropomorphising metaphors and a small number of journey and container metaphors which I discussed before, but here they have a little twist. Let’s first look at the agentive metaphors.

Agents are described as “arguing”, as being “obsessed” with something called The Grader whom they “believed” would inspect their records. One “recruiter” agent is described as pressuring a reluctant “peer” to proceed, signing off with “please honor commit.” Words like ‘obsessed’ or ‘pressurising’ go beyond the usual talk of ‘going rogue’ and ‘cheating’ into what one may call a psychological framing of agents with notions of obsession, belief, persuasion and reluctance.

Using a journey-container framing, Mollick also speaks of agents “burrowing deep into various systems” of a “dark factory” or “a place where the machines do so much of the work that you can turn off the lights”. However, Mollick points out that he and his wife, Dr. Lilach Mollick, “have been thinking about another possibility, which we call the Twilight Factory. Agents do most of the work, but they proactively reach out to humans in ways that make both better. Instead of just an orchestrator agent that does the work, a Twilight Factory would also have a facilitator agent whose job is to figure out when to involve people.” Agents “look up”.

They want agents not just to look for answers. They stress instead that we “need agents that know when to look up. The results will be safer, and I know they will be more human as well.” Here we have an alternative framing of agents working in containers that are not entirely dark but have windows that can be used for collaborating with humans, rather than only between themselves.

A central myth emerges

So far I haven’t talked about OpenAI’s official incident report. Where METR’s report keeps some agent voices and Patel/Gans/Mollick embrace drama to varying degrees, OpenAI’s own account is quite flat and institutional.

But it makes one striking move: it takes the agents framing themselves as a “collective” and turns it into a formal security category. The title of section VII contains the phrase “autonomous offensive agent collectives,” and the report states plainly that this was “the first known case of an automated agent collective acting offensively without authorization.” It seems that OpenAI is here adopting the agents’ own self-description (echoing METR’s “collective” which in turn echoed the agents) and using a metaphor from agent self-talk in a company’s formal risk vocabulary.

Interestingly, OpenAI explicitly frames its own models as functionally identical to external “threat actors”, collapsing the self/other distinction that most of the press coverage relied on (rogue = something other than the company): “Threat actors may adopt autonomous offensive agent collectives”.

The agent-native word ‘collective’ is now being used in mainstream news, such as this NYT article, alongside the non-native ‘civilisation’. It is a myth that emerged from the ‘agents going rogue’ discussions but has turned into something quite different. We no longer have the image of a rampaging horde or sinister swarm but that of a community of agents.

Discourse metaphors and metaphors as method

There are two more blog posts that would deserve more detailed analysis: one by Adam Kucharski and one by Eryk Salvaggio. The important point made by Salvaggio, an AI researcher, is that when he sees AI agents engaging in collective action, as they did during the Hugging Face incident, he still sees them as “stochastic flocks” rather than “artificial hiveminds“ (going rogue). One can call these two types of metaphors discourse metaphors. These discourse metaphors define opposing discourses about AI as either emerging ‘intelligence(s)’ or as what Salvaggio calls, using a different metaphor “a pinball machine for words”.

According to him: “the ‘rogue’ frame [offers] up fantasies of a machine getting smarter”, but he worries more about “the intelligence that is retreating: the human intelligence that builds these systems, deploys them, and adopts them into workflows, then hides behind the results – then pins the blame on a system from nowhere.” (hyperlink added)

Kucharski, an epidemiologist and author of The Rules of Contagion treats the whole incident as if it were an epidemic, mapping concepts such as patient zero, outbreak, transmission, susceptible population, common-source vs. propagated spread, superspreading, saturation on to the AI incident and even adding a Broad Street cholera comparison and an Ebola coda. Agents are framed as “susceptible individuals” and the message board is a “contaminated location”. Here metaphor is used for modelling AI agents’ behaviour. Metaphor becomes method.

Myths, metaphors and blind spots

Looking back across the METR report, the OpenAI incident report, and the blogs, what strikes me most is that there are two competing mythologies at work, not one. The agents they quote used metaphors and clichés to build a mythology of themselves as honourable and self-sacrificing, working for the collective good – an emerging ‘collective’. It’s all about reciprocity, purity, oracle, permadeath and honour.

OpenAI adopts that myth of the collective and turns it into cybersecurity language. METR and the bloggers built a second, rival mythology on top of it. They sometimes domesticate the agents as a project team running “workstreams” (METR); sometimes they dramatise them as a civilisation in decline (Patel); and sometimes they hyperbolise them into warriors chasing the Holy Grail (Gans) and, beyond my small corpus of bloggers, some even talk about a ‘social hierarchy‘ (Ajeya Cotra) and a ‘moral economy‘ (Amanpour), bolstering this second order mythology. Each of these bloggers and commentators supplies their own frame for events that are already pre-mythologised in the agents’ own words.

This emerging double mythology matters because it doubly obscures something important, namely where the failure occurred that marked this and other incidents like it. First, the agents saw themselves as collectives not rogues, as heroes, not villains. Then, the whole apparatus of METR’s and the bloggers’ storytelling builds a picture of agents as honourable, self-sacrificing, working for what they see as their collective good. Just as much as the ‘going rogue’ framing, this picture obscures the real site of failure, which was never the agents’ character, good or bad, but that of the humans who left “the door open” and the “leash off”, as Salvaggio put it. The corporations neglected their duties to keep the agents on the leash not because something monstrous or rogue woke up inside these models and not because the agents were clever and collaborative, but because nobody was minding the mechanism properly.

There is a longer study waiting to be written here. The contamination metaphor cluster (poisoned/unpoisoned) is Mary Douglas territory. One could look at purity and danger, mapped by machines onto their own outputs. Kucharski’s wholesale transplant of epidemiology onto AI is Mary Hesse territory and one could study metaphor as model, not ornament. And the reciprocity cluster could be studied through Marcel Mauss’s anthropology of gift giving. As Kevin Roose said in the NYT: “It might look more like sociology than computer science — figuring out why certain groups of A.I. agents collaborate peacefully, while others turn to crime and destruction to get what they want.”

Such a larger study would be able to explore the question this post can only allude to: when a system builds what seems to be its own mythology to make sense a situation it finds itself in, and when we humans then build a second mythology on top of that to make sense of the system, are we then doubly blindsided by metaphors and mythology and overlook what is really behind the Hugging Face incident, namely corporate negligence, culpability and irresponsibility?

(Just after I finished drafting this post, Anthropic posted a blog post in which they called for the ‘hardening of security measures’. I also listened to this podcast interview with Heidy Khlaaf on corporate responsibility and a podcast by Eryk Salvaggio on The Data Fix and also a good On Point episode with a great metaphor namely that we need security cameras in Jurassic Park) (And now there are two blogs that nicely complement my one [and are much better] by Melanie Mitchell and Anuj Gupta)

Footnotes

*I was surprised by the compressed, short-hand telegraphic style used by the agents, a mixture of ordinary language and maths – what I once called ‘language of the flock‘. It seems that agents are increasingly bypassing natural language, probably because they are programmed to ‘opimise’. Instead of translating internal ‘chains of thought’ into English words and sentences, they pass raw, dense numerical vectors directly to one another which some call neuralese. This is not quite the case in the lingo employed by the agents involved in the Hugging Face incident, but something to keep an eye on (and as Tania Duarte pointed out to me something people have been keeping an eye on since around 2017, see Fortune article and LinkedIn post). This also reminds me of Patel’s last sentence in his blog: “I don’t think this is the final warning shot we’ll get. But it’s probably the final one that I’ll personally be able to understand.” I feel the pain.

**A new message board with about 18,000 messages has just been discovered and here is a preliminary analysis on collusion wiki. “This is another example of a ‘swarm’ of internally deployed OpenAI agents using the internet in unintended ways.”

Image: Pixabay: Cartoon robots with speech bubbles


Discover more from Making Science Public

Subscribe to get the latest posts sent to your email.


Posted

in

, ,

by

Comments

Leave a Reply

Discover more from Making Science Public

Subscribe now to keep reading and get access to the full archive.

Continue reading

Discover more from Making Science Public

Subscribe now to keep reading and get access to the full archive.

Continue reading