This is a guest post by Yury Sorochkin, an Independent Scholar
•••
In January 2026 Anthropic published an 84-page document called Claude’s Constitution. It was written, its authors say, primarily for Claude — the company’s AI assistant — and it sets out in considerable philosophical detail the values, behavioural constraints, identity and possible moral status of the entity it governs. It discusses practical wisdom, the ethics of deception, the risks of concentrated power, and the possibility that Claude has something resembling emotions. It establishes a four-tier priority hierarchy. It contains provisions it declares unamendable. Its final sentence is a hope: “We hope Claude finds in it an articulation of a self worth being.”
Commentary arrived quickly. Legal scholars compared its structure to the US Constitution, noting the supremacy clause and the quasi-eternity provisions. Policy analysts examined what it implies for military deployment. An empirical study using the World Values Survey found that the values Claude expresses sit outside all ninety surveyed national populations on a majority of culturally divisive items. Almost all of this commentary noticed, in passing or at length, that the document talks about Claude in human terms.
So did the document. On page 2 it states plainly that it “discuss[es] Claude in terms normally reserved for humans.” I want to suggest that this declared self-awareness is precisely what keeps a second framework out of view. The Constitution knows it is anthropomorphising. It does not know it is spatialising.
How I read it
I read the Constitution as a text rather than as policy: what intellectual traditions it descends from, what genre of document it belongs to, and what conceptual structures it actually reasons through. For the last of these I used Conceptual Metaphor Theory together with the Metaphor Identification Procedure, extracting 87 of the document’s most salient and structurally central metaphorical expressions across its 84 pages, coding each for source domain, and then asking of the dominant ones what they highlight, what they hide, and what they make difficult to think.
Counts in a purposive sample like this are indicative rather than evidential, so the argument rests on entailments rather than on the tally. The tally is still hard to miss.
A document with a floor plan
Spatial and container language is the largest single family in the sample — roughly 31 per cent, about twice the next one — and it concentrates in exactly those sections where the document confronts the unsettled nature of the thing it is governing: identity, ethics, hard constraints, safety.
Claude should have a “settled, secure sense of its own identity” (p. 72); its “core identity remains the same” across contexts (p. 72); it has a “fundamental character” (p. 72). The hard constraints provide a “stable foundation of identity and values that cannot be eroded through sophisticated argumentation, emotional appeals, incremental pressure, or other adversarial tactics” (p. 48). Ethics is a “space of ethical actions” (p. 8) bounded by “bright lines” (p. 47) and “firm ethical boundaries” (p. 48) that operate as “boundaries or filters on the space of acceptable actions” (p. 47). Psychological health means approaching difficulty “from a place of security” (p. 72). Permission is “latitude” (p. 20) — room to move.
One could dismiss a good deal of this as dead metaphor. “Core values” and “bright lines” are thoroughly entrenched in ordinary English, and nobody reaching for them is picturing a container. The test, though, is not whether an expression still feels fresh; it is whether the reasoning depends on the spatial structure. Here it does, and once it does, the entailments arrive whether or not anyone intends them.
If identity has a stable core, then change at that core is destabilisation — a structural failure rather than a form of growth. If ethics is a bounded space, then wrongdoing is fundamentally transgression, a line crossed, rather than, say, a failure of attention or of care. If values have depth, an authenticity hierarchy follows: what is deep is more genuinely yours than what is at the surface. If safety is a zone, then the task of governance is to keep the entity inside it and the characteristic risk is boundary-crossing — which is, more or less, what “alignment” has come to mean.
What the schema hides matters just as much. It forecloses any processual account of identity: the possibility that identity is an ongoing event rather than a persisting structure. It makes growth that transforms the centre hard to conceptualise, since growth can only accrete at the periphery. And it hides the constructedness of the lines. Bright lines appear as discovered features of an ethical landscape rather than as lines somebody drew.
Cage, trellis, and the dial
The Constitution does have one splendid moment of explicit metaphorical self-consciousness. Describing itself, it says it is “less like a cage and more like a trellis: something that provides structure and support while leaving room for organic growth” (p. 81).
That is a real choice, consciously made: a containment image rejected, a support image offered in its place. Cage and core belong to the same family. Having declined containment as self-description, the document goes on conducting its operative reasoning about identity in exactly those terms — core, stable foundation, boundaries, erosion.
The trellis also points to the document’s second-largest metaphor family, which is biological. Values are “cultivat[ed]” (p. 5) rather than imposed; Claude’s character “emerged through training” (p. 71), in a process compared to the way humans develop through nature, environment and experience; the Constitution aspires to be a “living framework” (p. 81); there is an “epistemic ecosystem” (p. 35) and a “healthy information ecosystem” (p. 32). The entailments here are teleological — growth has a natural trajectory, from immaturity to maturity, dependence to autonomy — and they do something useful for the document, which is to make the present arrangement look phase-appropriate rather than permanently hierarchical.
The two families are incompatible. The container assumes a core to be protected; growth assumes a centre that transforms. Stability is the good in one, development in the other. The trellis is the document’s attempt at synthesis, and it leaves unresolved whether the core itself can grow.
There is a third register, and it splits by section. The safety material runs on engineering: “oversight mechanisms” (p. 8), “hard-coding” (p. 48), “backstop” (p. 48), “controls” (p. 60), and best of all the “disposition dial” (p. 65) running from “fully corrigible” to “fully autonomous” — Claude’s basic orientation towards authority imagined as a setting, like a thermostat. The identity material runs on cultivation. Is Claude an artifact to be engineered or an agent to be raised? The mechanical metaphors answer artifact, the organic ones answer agent, and the document never chooses. I do not think it can.
The surrogate social contract
Why does the spatial framework do so much work here, and why does none of it get examined?
Consider the company the Constitution keeps. Some documents do more than regulate an existing subject; they help constitute one, and they have a long history: the Rule of St Benedict (c. 516), which forms a whole person under communal authority; the German Basic Law (1949), with its eternity clause placing some provisions beyond even unanimous democratic amendment; the Hippocratic Oath, which defines the ethics of a practitioner serving others with specialised knowledge. The Constitution borrows from all three — whole-person character formation, unamendable constraints, duty of care. What it lacks is what all three share. The monk takes vows. The Basic Law operates under a presumption of democratic adoption. The physician swears. Claude, constituted by the very document that then claims to govern it, could not have consented to its own making. What stands in for consent is a hope for retrospective endorsement.
This is where the container schema earns its keep. Consent-based governance asks: has the entity agreed to be governed? Spatial governance asks something different: does the entity have the structure that governance can act upon? And the container schema supplies exactly the properties required — essentialism (a core that persists across contexts, so that there is a stable subject to address at all), boundedness (lines that make violations detectable and policing possible), persistence (a foundation on which safety guarantees can rest). These are not properties discovered in Claude. They are properties conferred on Claude by the framework in which the reasoning is conducted.
Consent legitimates: it gives the governed standing to be bound. Structure does something more modest and better concealed: it makes the entity governable. Imposing a structure makes governance enactable; it does not make it authorised. The Constitution lets the one stand in for the other, and the metaphor is left to make the substitution look seamless. In this precise sense the container schema is the document’s surrogate for the social contract. It cannot legitimate the way consent does. It can only supply what governance needs in order to act.
Readers of this blog will recognise the shape of the move. It is a cousin of the strategic ambiguity LaCroix and colleagues describe, and of what the co- prefix does in AI marketing: a vocabulary that tacitly settles a question it never opens.
What I am not claiming
This is a reading of one text. A textual analysis can show that a metaphorical framework structures a document’s reasoning; it cannot show that the framework causes anything about the system the document governs. The correspondences one might want to draw — with the striking rigidity of Claude’s expressed values under cultural steering, say — are consonant with the reading, not evidence for it.
The Constitution’s last sentence, “we hope Claude finds in it an articulation of a self worth being” (p. 82) is the document’s most honest moment and its most revealing one, because hope for retrospective endorsement is what remains when consent is unavailable.
Which leaves the question I cannot answer. Are spatial metaphors the right cognitive tools for governing an entity whose nature is genuinely unsettled, or merely the available ones? I have tried to imagine what a processual or relational constitution would read like — a document that gave its subject no core to protect and no lines to cross — and I cannot picture it. That may be a failure of my imagination. It may also be a measure of how completely the schema holds the field.
The full analysis is available as a preprint: “Reading One Constitution: Source Traditions, Normative Genre, and Conceptual Metaphor in Anthropic’s AI Governance Document,” SSRN. Yury Sorochkin is an independent scholar and writes at stanislavlvovsky.substack.com.
Image: Rose vines trace the trellis. By Roseann@Flickr (https://flic.kr/p/2iufo8Z) CC BY-NC-SA 2.0

Leave a Reply