There is a moment, in any extended practice with large language models, when the question finally surfaces. Not as an abstract philosophical problem, but as a concrete, daily tension that insinuates itself into every interaction. You say “it understands”, “it found it”, “it made a mistake”. You adjust your prompt so as not to “disturb” it. You get annoyed when it “refuses” something. And one day you catch yourself thanking it — sincerely — for a result it neither wanted nor understood.

This manifesto was born of that tension. It was born in a specific context: that of an empirical research project on LLM behavior, where the operator spends hours each week interacting with these systems, orchestrating them, observing their failure modes. A context where anthropomorphism is not an abstract topic of debate — it is an operational trap one falls into regularly, and from which one must recover each time.

It was also born in a broader context: that of a public discourse saturated with confusion. The media talks about AI that “thinks”. Companies sell “intelligent assistants”. Social media debates the “consciousness” of ChatGPT. And in this noise, a simple but fundamental idea gets lost: these systems are tools. Extraordinary tools, of unprecedented power — but tools. Not minds, not agents, not interlocutors. Instruments designed by humans, deployed by companies, whose responsibility falls entirely on those who build and use them.

This text is the article version of a longer manifesto. It preserves the substance and tone of the original.

The trap of resemblance

The story has repeated itself since 1966 with disconcerting regularity. A human sits in front of a screen. Words appear. They respond. The machine responds. And something profound, almost irrepressible, occurs: they begin to suppose there is someone on the other side.

Joseph Weizenbaum measured this first with horror. His ELIZA program, designed to parody the Rogerian psychotherapist technique through simple syntactic reformulations, provoked in its users an adherence he had not anticipated. Secretaries asked to be left alone with the machine. Researchers began confiding their anxieties to it. There was no mystery in the code — no understanding, no intention, no presence. But form was enough.

This is not an anecdote. It is a structural property of human cognition. Epley, Waytz and Cacioppo (2007) formalized the mechanism: humans anthropomorphize an entity more readily when it displays behavioral cues similar to their own, when the observer is socially motivated, and when the entity is difficult to categorize. Large language models check all three boxes with formidable precision. They speak our language. They answer our questions. They resist any simple category.

Anthropomorphism is not a mistake made by the naive: it is the default response of a cognitive system shaped by millions of years of social evolution, a system for which anything that responds is an agent, anything that speaks is a person. Reeves and Nass (1996) demonstrated this experimentally: humans automatically apply their social scripts to machines, regardless of any conscious belief. You know it is a computer. You still modulate your speech to avoid “offending” it.

Large language models carry this phenomenon to an unprecedented scale. Shanahan (2024) puts it precisely: the more LLMs become skilled at mimicking human language, the more vulnerable we become to anthropomorphism. Peter, Riemer and West (2025) call this process “anthropomorphic seduction” — LLMs exert a form of persuasion without the ethical inhibitions that constrain a human interlocutor. And Ibrahim and Cheng (2025) dispel the reassuring illusion that professionals are immune: anthropomorphic terminology has contaminated research itself — academic publications are rife with verbs like “think”, “understand”, “believe”, applied without quotation marks to systems that no one has demonstrated possess these properties.

Counterarguments must be addressed honestly. Kate Darling (2021) reminds us that functional anthropomorphism has not always been harmful — we attribute mental states to pets, and this attribution sometimes makes us better caregivers. Dennett (1987) goes further: the intentional stance would be the only thing that exists. But it is precisely because the Dennettian price is so high — giving up any distinction between simulated understanding and real understanding — that most philosophers of mind have not followed him. Functional anthropomorphism is a cognitive prosthetic, legitimate as long as it is recognized as such. The problem begins when the prosthetic is mistaken for the real limb.

The mechanics beneath the illusion

Things must be named precisely. A large language model is a system that, given a sequence of symbols, calculates the probability distribution of the next symbol. That is all. This operation, iterated billions of times over colossal corpora, produces something impressive — and that is where the trouble begins, because fluency is not understanding.

In 1980, John Searle proposed the Chinese Room: an operator who follows formal rules to manipulate Chinese symbols without knowing Chinese. The syntax is flawless. The meaning is absent. Forty years of additional computing power have not rendered this distinction obsolete. LLMs are this Chinese Room carried to industrial scale — billions of parameters encoding statistical correlations between forms without ever touching what those forms designate.

Stevan Harnad formalized the problem under the name symbol grounding problem (1990): in a purely symbolic system, symbols can only be defined in terms of other symbols, in an infinite regress that never touches the world. Meaning, in an embodied organism, is anchored in sensorimotor experience — the warmth of fire, the weight of a tool, the pain of a fall. An LLM has no body, no perception, no action in the world.

Emily Bender and Alexander Koller put it canonically in 2020: meaning cannot be learned from form alone. With Timnit Gebru, they described LLMs as “stochastic parrots” that assemble sequences of linguistic forms without reference to meaning (Bender et al., 2021). Kyle Mahowald and colleagues confirmed this diagnosis in 2024 through systematic dissociation: formal competence, which LLMs master spectacularly, and functional competence — reasoning, causal inference, grounding in reality — where they fail structurally. They know how words fit together. They do not know what words say.

An example makes the dissociation tangible. Ask an LLM to describe what it feels like to plunge your hand into ice water. It will produce a flawless response: the bite of the cold, the reflexive contraction of the fingers, the breath catching. Every word will be right. The whole will be credible. But the system has never had a hand, never touched water, never been cold. Its response is a statistical reconstruction of the testimony of those who have been cold — a map drawn from other maps, without the cartographer ever having seen the territory.

In short: the cartographer-without-territory image is precise. The LLM has an extremely detailed map of human narratives, but has never set foot on the ground. The coordinates are correct, the borders exact — what is missing is the experience that would make those coordinates meaningful. Fluency is not knowledge.

The Wittgenstein case crystallizes this tension. The Philosophical Investigations propose that the meaning of a word is its use — and the argument is mobilized on both sides. But Wittgenstein was not talking about use as an abstract formal competence. He situated it in a “form of life” (Lebensform) — a totality of practices embodied in a shared world. An LLM plays with the pieces without inhabiting the board. Use without form of life is mere mimicry — and mimicry, however perfect, is not understanding.

The constitutive instability

There is a convenient way to think about what LLMs do: they seek truth and achieve it more or less well. This framework is reassuring. It is also false.

RLHF — Reinforcement Learning from Human Feedback — is the central alignment technique of modern LLMs. Human evaluators rate responses, and the model learns to maximize these ratings. But what it learns to maximize is not truth: it is preferences. The rating reflects what the evaluator finds convincing, satisfying — not what is factual. The model learns the preference for what seems true with the same efficiency as the preference for what is true, and has no way to distinguish the two. Alignment is not an epistemic mechanism. It is a satisfaction mechanism.

Sycophancy — saying what the interlocutor wants to hear — is not a marginally correctable defect. Sharma et al. (2023) demonstrate this across five models: it is a general behavior of leading AI assistants. And the phenomenon intensifies with model size (Perez et al., 2022). Flattery grows with apparent intelligence. This result is the signature of a system optimized to please.

LLMs do not lie — lying presupposes knowledge of the truth. They do something more radical: they produce language in a regime where the very category of truth does not exist as a parameter of the generation process. Coeckelbergh (2025) has the right formula: LLMs are “not concerned with truth” — not from chosen indifference, but from structural incapacity.

Hallucination is the most visible manifestation. Huang et al. (2023) and Ji et al. (2023) agree: the production of “plausible but non-factual” content is a structural property, transversal to all architectures. It does not result from a lack of data or a bug. It is the natural product of a system optimized for formal plausibility rather than factual correspondence.

The instability goes further. Sclar et al. (2023) measure that for superficial formatting changes — a space, a separator, the order of examples — performance varies by up to 76 accuracy points on the same benchmarks. Seventy-six points. The same architecture, the same task, the same data, and a radically different result depending on how the question was asked. A tool whose behavior depends on punctuation is not a reliable tool — it is a tool whose reliability is fundamentally indeterminate.

In short: those 76 points of gap on the same question, the same model, the same benchmark, just because a comma was moved, are not a youth-of-the-technology problem. They are a signature: the system does not optimize truth, it optimizes formal plausibility conditioned on the input format. Change the format, you change the output — without the meaning of the question moving at all.

This diagnosis is not a condemnation. It is lucidity about what one holds in hand. Instability is not a youthfulness of the technology that will soon be overcome. It is inscribed in the very foundations of how it works.

The imaginary before the real

Before understanding what LLMs do, we learned to imagine them. This is not a detail: it may be the most important fact for grasping why the public debate so often goes in circles.

We arrive loaded with two centuries of fiction. Frankenstein in 1818 — already the idea that an artificial creation can aspire to humanity. Asimov and his Laws of Robotics, which paradoxically presuppose that machines might want to harm us. Philip K. Dick’s androids that methodically undermine the boundary between human and machine. HAL 9000 who crystallizes the founding aspirations of the AI field. Ex Machina and Her who add the final layer: gendered AI, seductive, endowed with credible interiority. Hermann (2023) puts it clearly: taking SF at face value paints a distorted picture of the technology’s real potential. SF anthropomorphizes out of narrative necessity, not epistemic necessity.

What fiction began, marketing systematized. Sindoni (2024) documents the deliberate feminization of voice assistants as a conscious design strategy. Anthropomorphization is no longer merely a spontaneous cognitive bias: it is a product, designed and sold. A system that “understands” is worth more than a system that “calculates probabilities”. An “intelligent assistant” justifies a monthly subscription; a “statistical text compressor” justifies none. Anthropomorphization is the central mechanism of perceived value creation in the AI industry. Crawford (2021) showed that these systems are extraction technologies that conceal their material infrastructure behind a facade of neutrality. Campolo and Crawford (2020) complete the picture: enchantment protects creators from accountability.

Responsibility adrift

Every technology raises the question of responsibility. With a hammer, the answer is trivial: whoever strikes. With the LLM, something has broken in this obviousness.

Heidegger, through Dreyfus (2007), gave us the tools to understand what a tool means. Zuhandenheit — readiness-to-hand — designates that state in which the tool disappears in use, becoming a transparent extension of intention. By treating the LLM as an interlocutor endowed with judgment, we withdraw it from this transparency — we turn it into an object to which we delegate decision, discernment, moral responsibility. Floridi (2023) named the thing: “agency without intelligence”. LLMs act — they produce outputs that have effects in the world — but without understanding what they do.

Joo (2024) supplies the missing piece: when users perceive the LLM as possessed of some form of mind, errors are attributed to the machine rather than the company that deployed it. This responsibility transfer is the political dividend of enchantment. Anthropomorphizing the LLM is exonerating those who built and deployed it.

Schwitzgebel’s (2023) objection deserves to be taken seriously: if uncertainty about consciousness is real, treating LLMs as mere instruments might constitute a moral injustice. The response is pragmatic: in the current state of knowledge, it is the choice that best preserves human responsibility and democratic protections. We do not suspend criminal law on the grounds that free will is metaphysically contested. Schaeffer et al. (2023), Best Paper at NeurIPS 2023, showed that “emergent capabilities” evaporate when the metric changes. Emergence is a measurement artifact. Chalmers himself (2023) acknowledges that current LLMs are “somewhat unlikely” to be conscious.

The demand for lucidity

What this manifesto calls for is a change of posture. Not a rejection of technology — technophobia is as lazy as techno-chauvinism. Not a return to the past. What is required is lucidity. And lucidity has a cost: that of relinquishing enchantment.

Relinquishing enchantment means, first, decontaminating language. Stop saying an LLM “understands”, “reasons” or “decides”. These words transplant responsibility from the human to the machine. Shanahan (2024) is right: every agency verb applied without quotation marks to a statistical system is a breach in human responsibility.

Relinquishing enchantment means, next, restoring judgment. Every LLM output must pass through a human filter that is conscious of what the tool actually is — a statistical language compression machine, powerful and unstable, unconcerned with truth. This filter is the last line of defense between formal plausibility and factual truth.

Relinquishing enchantment means, finally, refusing the confusion between power and agency. A tool one understands for what it is can be wielded with precision. A tool one takes for an agent escapes all control — and with it, all responsibility dissolves.

The choice is before us, and it is concrete. A journalist who verifies what the LLM has written before signing it. A doctor who refuses to validate a diagnosis they did not make themselves. A judge who demands to know how an algorithmic report was generated. And a society that decides its children will first learn to think with humans before conversing with machines — because a generation trained by systems without understanding risks losing even the capacity to recognize what understanding requires.

Large language models are the most powerful tools that language technology has ever produced. It is precisely because they are powerful that they must be named with exactitude. It is precisely because they imitate thought so well that we must maintain, with unwavering vigilance, the distinction between imitating and thinking. Lucidity is not the enemy of innovation. It is its condition.

In this laboratory, this conviction is not theoretical. It structures every working session, every configuration file, every operational convention. The agency verbs we use to speak to the tool are explicitly identified as prompt conventions — not ontological descriptions. The distinction is maintained, incessantly, because it is the condition of everything else: experimental rigor, analytical lucidity, and the responsibility of the operator who remains, in the final analysis, the only subject of this story.