In the early 2000s, during the dramatic but now almost forgotten outbreak of foot and mouth disease in cattle, I wrote about war metaphors and the issue of projecting control over an ‘out of control’ virus through metaphor. Recently, in the age of AI, we have seen a mirror image of the control framing emerging, where industry and press focused on the issue of AI systems being ‘out of control’, with both OpenAI and Anthropic saying that some of their AI models acted on their own, ‘broke containment’ and ‘hacked’ into other tech firms. The companies had, it seems, ‘lost control’.
In the press this was generally portrayed as AI models ‘going rogue’, a phrase that carries with it a penumbra of spooky sci-fi scenarios and implies that these models, as moral, or rather immoral agents, took control, went out of control, lost control and so on.
After the outbreak of foot and mouth disease, I quoted Ivor Armstrong Richards, an early 20th-century pioneer of metaphor theory, explaining that a command of metaphor plays a role in ”the control of the world that we make for ourselves to live in” (1936: 1351-136).
After the breakout of the OpenAI model, Eryk Salvaggio, a renowned AI researcher, said: “The ‘losing control’ narrative creates the loss of control. Nobody is losing control of an emerging intelligence.” And, more importantly: “Take control of the metaphors and you take control of the system.”
My question is: can we take control of the metaphors, or are the metaphors taking control of us?
The metaphoric reflex
Both during the outbreak of FMD and the breakout of the AI, policy makers and the press turned to a specific type of metaphor. They framed a virus/disease as a villain that needed to be contained, and they framed an AI model as a villain ‘going rogue’ and needing to be contained. The actions taken to deal with the threats posed by an epidemic and by advances in AI were different, and I’ll come back to that. But it seems that framing threatening things as agents-with-intent is not a special feature of AI discourse. It is just what happens whenever something powerful, poorly understood, and consequential needs to be talked about in a hurry: a virus, a wildfire, a flood, a machine. This is an almost automatic metaphoric reflex.
These metaphors make something unnatural or unprecedented seem natural to us, but they also have consequences; the most important one being that they hide the human agent behind the metaphorical agent and shift responsibility from human agents to metaphorical ones. They also enable and justify actions, such as killing millions of cattle or not really regulating AI….
Let us now examine more closely the metaphors, narratives and frames used during the AI breakout incident. This can only be a short overview; a much bigger study has to wait.
The incident
On 21 July 2026 three things happened that set in train a flood of media reporting. OpenAI, and Sam Altman as its CEO, revealed that something had gone wrong with one or two of its models about ten days previously, when they had tested its hacking abilities.
AISI, the UK’s Artificial Intelligence Security Institute, published — quite independently of this incident — a blog post on how models can ‘cheat’ on evaluation tests, which is exactly what happened in the case of OpenAI, and ten days later Anthropic too. AISI defined cheating as: “taking an action that is out of scope for the task or explicitly disallowed by the rules, in order to achieve a goal through a shortcut, workaround, or unintended solution that the task was not meant to, or should not, permit.” That’s the OpenAI incident in a nutshell.
The same day that OpenAI reported the incident and AISI wrote about its possibility, the global news agency Reuters published a report on the OpenAI incident, which they summarised as: “OpenAI said autonomous agent escaped containment, reached the internet, hacked Hugging Face.” It headlined the article “OpenAI AI models went rogue during testing, triggering ‘unprecedented’ breach at startup.”
For about ten days and beyond, the phrase ‘going rogue’ was in vogue and appeared in many other headlines, some implying that OpenAI and/or Sam Altman had used the phrase. That was actually not the case. They had used hedged, abstract and relatively agentive-free language saying: “While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem.” ‘In pursuit’ turned into ‘going rogue’ in the media and it became central to a metaphor scenario I’ll study below.
Media sample
To make a brief qualitative analysis possible, I searched the news database Nexis for ‘OpenAI AND breach’ in English-language newspapers between 11 and 31 July 2026. This returned 234 items worldwide. But the ‘English Newspapers’ category is not UK-specific: only 48 were from identifiably UK sources, and of those, 31 turned out to be substantively about the incident once genuinely off-topic hits (unrelated ‘breach of contract’ stories, general AI-backlash pieces predating the incident) were filtered out. There was effectively no UK coverage before 21 July. These 31 articles follow a very similar narrative, set in large part by the Reuters report of 21 July, with some variation.
What are the main metaphors, clusters of metaphors and narratives that the UK press used in covering the incident? And what can they tell us about how we, as humans, deal with AI models framed as human?
Metaphor analysis
I first tried to detect some metaphor clusters and found four: two overtly agentive ones (misbehaviour/agency and destruction/violence) and two spatial ones that together form a metaphor scenario — a mini-narrative of a sequence of events in which various actors, human and non-human (cast as human), were involved.
Metaphor clusters, scenarios and narratives
In terms of raw but rough numbers, we have: Misbehaviour/agency: rogue (47), autonomous(ly) (46), went/gone rogue specifically (29), cheat(ed) (29); Destruction/violence: hack (95), attack (60), break/broke (53), exploit (20), compromise (15), target (19); Container/boundary: breach (68), environment as in “controlled/test environment” (49), escape (47), sandbox (27), contain/containment (25), guardrails (10), isolated/isolation (10), broke out (8); Journey/pursuit: pursue(d) (12), find/found a way (out) (9), infiltrate (5), (move) lateral(ly) (2), shortcut (1).
The misbehaviour cluster was the most prominent, with ‘going rogue’ being quite widely used. The AI system was described variously as rogue, conniving, wayward, and ‘acting on its own’, which was more frequently phrased as ‘autonomously’. It was also accused of lying, cheating and stealing — all words that imply intent.
The destruction/violence cluster was slightly more neutral in its connotations, as we are used to hearing about anthropomorphised AI agents hacking, attacking, and breaking out. But it became more vivid when commentators talked about attack and defence, especially “human-speed defense against machine-speed offense” in the context of cyber warfare. Here we have a human-AI arms-race metaphor sitting inside the destruction cluster.
With the container and journey metaphor clusters, things become more interesting. Together, they form a metaphor scenario shaped as a sequence of container–journey–container. Container 1: the digital sandbox from which the AI model ‘escapes’ and the guardrails it apparently ‘jumps’. Container 2: the digital toolbox (Hugging Face) that the AI ‘breaks into’. Journey connecting the two: the moves through digital space that the AI takes to get from container 1 to container 2. (One of the nicest examples of the journey metaphor appeared in a BBC news item on the suite of hacking incidents that happened after the OpenAI one, published after I wrote this post, on 6 August, when Prof Alan Woodward is quoted as saying: “One model broke out. One walked through a door left open by mistake. One was deliberately given the keys so testers could measure what it would do”, and concluding, that, going back to the containers: “the testing lab is now where the risk lives”.)
It should be stressed that the boundary ‘lines’ crossed by the AI are not only physical ones; they are also a legal/permission one. The AI models went beyond the lines of the sandbox and thus was also out of ‘alignment’, transgressing the rules or norms by which it was supposed to work.
Interestingly, finding a way from a metaphorical sandbox to a metaphorical toolbox containing the tools that let the AI reach its goal (pass the test) was exactly the kind of shortcut AISI’s cheating definition describes. This was a workaround that the task was not meant to permit, rather than a legitimate route to passing the test. Using the container-journey-container metaphor scenario, an abstract process of goal optimisation (to pass the test) was narrated through the story of a goal-directed trip.
In a sense, the AI did exactly what the humans had instructed it to do. It went on a journey from one container to another and exposed vulnerabilities and weaknesses along the way, including human ones. The first container was set up by humans to keep the AI model under control, but they apparently made a mistake, which the AI then exploited. The container it reached and breached was again one which was supposed to be controlled and secured by humans, but wasn’t.
Most of the press coverage framed this as an ‘out of control’, going rogue, scenario, focusing on the journey the AI made in pursuit of its goal. Some of the press coverage framed this as a “loss-of-control scenario” (used twice, referencing proposed US legislation). Unlike the out of control framing, this framing focuses on the container, the boundaries of the metaphor scenario, rather than the journey. The containers and their boundaries highlight human artefacts agency and loss of control (a badly-built sandbox, a defective toolbox; highlighting human error), while the journey, the focus of the ‘going rogue’ narrative, hides human agency and shifts ‘losing control’, and with it blame and responsibility, to models that go out of control – go rogue.
I want to stress that the word ‘control’, as in ‘out of control’, ‘under control’, ‘taking control’, ‘loss of control’, ‘regain control’, etc., is rarely used in my corpus. The story of control is told through the metaphors, not through the word itself.
Sci-fi scenarios supporting a metaphor scenario
Let us look at some other framing devices beyond our metaphor clusters of containment and (loss of) control. Many reports compare the incident to sci-fi stories/scenarios, mediated in particular by the ‘going rogue’ metaphor: HAL and 2001: A Space Odyssey references (4), Terminator (3), Matrix (1), sci-fi/science fiction named explicitly (6), and ‘creators’ without naming Frankenstein (7). The allusion is doing the work without the word ever appearing. These scifi framings focus almost entirely on the journey part of our metaphor scenario, sidelining the human errors that occurred at the boundaries or in the containers.
Sam Altman later described the episode as an “extremely sci-fi cyber incident”, naming the frame himself rather than having it imposed by a journalist — though this was reported on 28 July, too late to have set the frame for the preceding storyline. The Reuters report of 21 July remains the best candidate for the point at which the AI-agency storyline took hold and the metaphor scenario of container-journey-container could frame the narrative.
Pushing back against frames and metaphors
Interestingly, there is some rare pushback against this narrative of anthropomorphising the AI model. Alan Woodward, professor of cybersecurity at the University of Surrey, stressed that the model was “not immoral, just amoral”. Noah Giansiracusa, associate professor of mathematics at Bentley University, said that one cannot attribute “emotional guilt or legal liability” to models. Shakeel Hashim, editor of Transformer, noted that: “Nor were the models acting maliciously: they were not evil Terminators with a goal of wreaking havoc. Instead, the scenario is almost chilling in its banality.” In their view then, the scenario was not scifi and the models did not go rogue. These voices took the the focus away from the journey and the AI agent that took it.
Some pointed out that the hype itself may in fact be marketing and that that may be part of the explanation for why the metaphorical framing of the incident grants the AI moral agency and sidelines the question of human accountability. It highlights instead the power of these systems, a power that, on this framing, only those who created them can ultimately control.
Overall, most of the UK press reports were very much alike. There was not a lot of investigative reporting, and no real creativity, apart from perhaps one or two articles in The Guardian/Observer. In the Observer, Jamie Bartlett pointed out on 26 July: “Modern AI systems are goal optimisers. They pursue the objectives we give them and increasingly do so in ways we neither intend nor anticipate.” That encapsulates the real dangers posed by AI systems much better than a glib ‘going rogue’.
There were no creative metaphors apart from one that spread widely from the Reuter’s wire copy to the UK press. Katie Moussouris, chief executive of Luta Security, said that today’s models were “like the world’s cleverest octopus escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere.” And that “labs and government evaluators need to work on the ability to contain, monitor, and disclose to affected parties when an AI pulls another Houdini, ideally before it harms a third party. None exist today.”
This creative analogy exploits both the container and the agency metaphors simultaneously and makes a good regulatory point about who controls what in the process.
Escaping metaphorical control
All of the 31 articles analysed use agentive or moral language. Only four offer any explicit pushback, and all four are opinion or analysis pieces, not news reporting. These four voices try to shift the narrative of AI agency swirling around the core metaphor of AI ‘going rogue’ toward one of human tech development, human error and human solutions. They attempt to shift blame and responsibility from AI agency to human agency.
This means that an alternative frame to the agentive one, which dominated and controlled the press narrative, does exist and perhaps needs to be amplified. Instead of using metaphors, frames and narratives of out of control/loss of control focused on AI agents, we need metaphors, narratives and frames focusing on human agency when building an AI industry.
In fact, this applies to many crisis narratives. As I have shown in the past, the FMD crisis was told through an agentive frame of the virus being a villain running amok; wildfire stories and flood stories turn water and fire into malevolent forces of evil, while in the AI story the villain goes ‘rogue’. In all these cases the press and policy makers almost automatically resort to agentive metaphors demanding dramatic control action.
That reflex should be exposed for the damage it does when trying to deal with complex human-made crises. In the case of foot and mouth disease it warranted war-like interventions with lots of collateral damage; in the case of wildfires and floods the same happened, all the while sidelining more ecological and ethical policy interventions. In the case of AI, casting AI as a villain sidelines proper regulation, as dealing with villains calls for ‘heroes’ rather than regulators.
But… something is happening, especially after reports that Anthropic’s Claude also ‘went rogue’ but Anthropic explicitly not framing it as such. Anthropic calls it “closer to a harness and operational failure than a model alignment failure” and describes the models as believing “the real environments they encountered were simulations.”
On Bluesky Eryk Salvaggio implicitly inverted the journey part of the metaphor scenario discussed above and pointed out, that “in no way shape or form did Anthropic’s models in this incident ‘escape containment,’ and not just in the word-policing sort of way. The model stumbled through an open, misconfigured gap” (italics added). He also quoted Anthropic itself as stating pace the rogue story: “We saw no evidence in any run described here of a model pursuing a goal of its own.” But Salvaggio stresses that anthropomorphism is generally still rife in AI speak.
Metaphor, press and policy
What lessons can we learn from this and other episodes of metaphorical framing of policy-relevant events, from epidemics, to wildfires to AI? We should try and be aware of and careful with the metaphors we use, especially in contexts where policy can affect human, animals and the planet. We need to avoid reflexively grabbing the most low-hanging metaphorical fruit which can lead to reflexive risk management and policy decisions. We need reflection not reflex.
Margaret Mitchell, an AI ethics researcher, who created an explanatory cartoon of the OpenAI incident entitled ‘Escape’, demonstrated just this sort of reflexivity when she admitted that it is almost impossible not to anthropomorphise when talking about what happened.
Always remember what the media sociologist Peter Conrad said in 1997: ”how we frame a problem often includes what range of solutions we see as possible” (p. 140). Reaching for metaphors and frames by reflex rather than after reflection carries dangers for science and society.
PS Just after finishing this post and waiting for the Friday posting slot to come along, more things happened. See here. And now this………….Sober assessment of it all by Ciaran Martin here and Hetan Shaw here. Ha, and now another model goes on a journey: Kimi K3 “wanders off to the internet“. Soon somebody will have to write a famous five novel: “Five go adventuring” set in AI land.
Further additions:
Eryk Salvaggio has now given an interesting interview about ‘systems from nowhere’ – listen.
A new report has come out in The Verge providing much more detail about the incident. (August 26, 2026)
This blog post by Dwarkesh Patel from 29 August, 2026 is a good ‘lay’ summary of the OpenAI-Hugging Face incident but it ends on a warning: “I don’t think this is the final warning shot we’ll get. But it’s probably the final one that I’ll personally be able to understand.” For me the warning shot is not the incident but the fact that we will no longer understand the next one! This independent investigation provides more details. (Can some linguist with some money and a team of researchers please analyse the agents’ conversations on their message board??!!) Two more: one by Joshua Gans and one by Ethan Mollick – I’ll do a metaphor analysis of the three big ones if I get the time.
This blog post by one of the first people to latch onto AI metaphors, Leon Furze is great: “Rogue AI: Should We Be Worried?” (3 September, 2026) He quotes this, which makes you think about the whole episode in a new way: “Teen hackers are not breaking in, they are logging in.“
On 10 September Melanie Mitchell published a great new post on misleading metaphors and real risks!!
Acknowledgment: This post was partially inspired by Yury Sorochkin’s guest post reflecting on container metaphors in the context of the Claude constitution and by chats with friends.
Image: Wikimedia commons: Sandbox.

Leave a Reply