A Mind That Has Never Touched Anything
Our most capable text machines have no hands. The phenomenologists would have called that a category problem, not a limitation.
Ask a competent model to explain how to ride a bicycle and you will get a good answer: keep your speed up, look where you want to go, steer slightly into the fall to bring the wheels back under the centre of mass. Every clause is correct. The description is better organised than the one your father gave you in a car park thirty years ago. It is also not the thing that happened in the car park, and everybody who has learned to ride a bicycle knows this without needing the philosophy. What you acquired that afternoon was not a set of propositions. Something in your hips learned to answer a question your head could not formulate.
Maurice Merleau-Ponty built the Phenomenology of Perception in 1945 around the refusal to treat that as a curiosity at the edge of cognition. His claim is that it is the centre. The body, in his account, is the subject of perception: the thing that has a world at all. Treating it as an object the mind steers reverses the relation. We do not perceive a field of sense data and then infer a room; we perceive a room, because we are the kind of thing that can move through one. The doorway looks passable, the cup looks graspable, the stairs look climbable-when-tired, and for a body that could do those things, that is what the visual scene is made of. Nothing has been added to it by inference.
He is best known for the case that makes it vivid. A patient with an amputated arm continues to experience the limb, reaches with it, protects it. Merleau-Ponty’s reading is that the phantom is not a hallucination of a missing object but the persistence of a world — a set of possibilities for action that the body still holds open. The limb is missing from the anatomy and present in the grammar.
Whatever else language models are, they are the first serious candidates for mind in our history that have no such grammar. No proprioception, no fatigue, no reach, no grip, nothing that is tiring to hold. They learned entirely from the traces left by creatures that had all of it.
The paradox that has been sitting there since 1988
The empirical shape of the problem was noticed early and named after Hans Moravec. The things we find hardest — chess, integration, formal logic, competent prose — turned out to be the things machines got first. The things any toddler does — recognising a face across lighting conditions, walking over a cluttered floor, picking up an unfamiliar object without crushing it — turned out to be the last and are still not solved to a toddler’s standard. Moravec’s own explanation was evolutionary: perception and locomotion have been under optimisation for hundreds of millions of years and are therefore very good in ways that are invisible to us, while abstract reasoning is a recent, thin, poorly compressed layer that happens to be the part we can hear ourselves doing.
Hubert Dreyfus had made the philosophical version of the argument in 1972, drawing directly on Merleau-Ponty and Heidegger, and was mocked for it for two decades. His claim was that human expertise rests on a background of embodied know-how that no rule set reproduces, and that symbolic AI would keep hitting this wall. He was substantially right about symbolic AI and, in the specific form he gave it, substantially wrong about what came next — a combination that makes his book more useful now than a straightforwardly correct prediction would have been.
Rodney Brooks took the constructive route in 1991: build robots with no world model at all, let the world be its own representation, get competent behaviour out of tight coupling between sensors and actuators. Francisco Varela, Evan Thompson and Eleanor Rosch gathered the broader position in The Embodied Mind the same year, and Lawrence Barsalou’s perceptual symbol systems gave it a cognitive-science formulation in 1999: concepts as partial re-enactments of perception and action rather than amodal symbols. On all of these accounts, a system trained only on text is holding the labels and not the things.
The labels were produced by creatures with hands. That is the whole question: whether the hands came through in the text.
The paradox has partly dissolved since 1988, and the way it dissolved is instructive. Machines did get much better at perception and at manipulation, and they got better by the same route as everything else — enormous quantities of data and a great deal of compute — rather than by acquiring anything a phenomenologist would recognise as a body. A robot that grasps an unfamiliar object reliably has not developed a body schema. It has a policy learned over a very large number of grasps, most of them in simulation.
That is either the refutation of the embodiment thesis or a demonstration of how much of what we attributed to embodiment was statistics all along, and the two are hard to tell apart from outside. What has not dissolved is the narrower gap, and it shows up where the data runs out: in the situations no one has done ten million times, where a person’s answer comes from having a body rather than from having seen bodies.
Borrowed grounding
Here is where the argument gets more interesting than its usual form. The text a model learns from is the sediment of billions of hours of embodied life, written throughout by creatures with bodies — people describing weight, resistance, cold, the specific effort of carrying something awkward up stairs. The structure of embodiment is in the corpus, in the same way that the structure of three- dimensional space is recoverable from a great many two-dimensional photographs.
So the question is not whether a text-trained system has access to embodiment. It plainly has access to an enormous amount of it, second-hand. The question is what second-hand access is worth, and it is open. It might be nearly everything, if what matters about a concept is its relational structure and that structure survives the transfer. It might be much less, if the relational structure is anchored at the edges by things you can only have by having had them — which is what Merleau-Ponty and the enactivists claim.
The sharpest recent statement of the sceptical side puts it as a point about training signal rather than about bodies. Emily Bender and her co-authors argued in 2021 that a system trained on form alone — sequences of symbols, with no access to the communicative intent behind them — has access to form alone, and that fluency invites us to supply the meaning ourselves. The argument does not depend on embodiment, and it is stronger for that: it applies equally to a system with cameras bolted on, if what the cameras deliver is another stream of form.
There is a cheap version of this argument that should be avoided: the assertion that a system which has never been cold cannot use the word “cold” correctly. It plainly can, by any behavioural test we have. The serious version is narrower and harder to test: that certain judgements are made out of the residue of having done the thing, and that competence in describing them is not evidence of having the residue. How much force is too much when the object might be fragile. Which way a body is about to move a half-second before it moves. When an instruction is physically impossible in a way so obvious that no one who has ever done the job bothered to write it down.
The case against
The case against is best put as a history: the embodiment argument has an unbroken record of losing, and its defenders have never once conceded a round.
The pattern is well established. A capacity is declared to require embodied know-how; a system without a body acquires it; the argument is restated with the goalposts moved to whatever remains. Dreyfus’s original list of what computers could not do has been steadily emptied, largely by systems that violate his premises. Language itself was the flagship case — meaning was supposed to require being in a world, and text prediction was supposed to be able to produce syntax at best. That prediction failed badly enough that the sensible response is humility about the next one.
There is also a live empirical alternative to the embodiment thesis: that grounding is social more than it is sensory. Barbara Landau and Lila Gleitman’s 1985 study of a congenitally blind child found her acquiring visual vocabulary — look, see — on the ordinary schedule and using it with a structure fitted to her own way of examining things, which is not what you would expect if such words were empty tokens for her. What a blind speaker has is a place in a community of speakers who do see, and it appears to be enough to do most of the work. If that is where grounding comes from, then a system trained on the whole record of that community has been handed the relevant thing, and the missing sensorium is a distraction. And the practical trend runs the same way: the systems in question are no longer text-only, and connecting them to cameras, robots and force sensors is an engineering programme rather than a philosophical impossibility. An argument that keeps having to be reformulated after each result has stopped tracking a boundary and started defending it.
The multimodal reply deserves a closer look than it usually gets, because it settles less than its confidence suggests. Connecting a model to cameras and force sensors gives it streams of numbers correlated with the world. What Merleau-Ponty describes is a body that constitutes the world as a field of things it could do, which is a different sort of arrangement altogether. The difference is whether perception is data arriving at a system or a system’s way of having a situation. A robot with excellent sensors may well acquire the first and remain entirely without the second, and nothing in the engineering programme is aimed at the second, because nobody knows what would count as aiming at it.
That said, the sceptic has to concede the obvious: this is the kind of distinction that has repeatedly turned out to make no measurable difference, and a distinction that makes no measurable difference is under permanent suspicion of making none at all.
What survives, for me, is smaller than the thesis and harder to dismiss than the objection. The bicycle explanation was correct. It was assembled from every description of bicycles ever written, which is a very large number of descriptions written by people who had fallen off. Whether that is knowledge of bicycles or an exceptionally good report about knowledge of bicycles is not answerable from the text, and it may not be answerable at all — but it becomes urgent the moment such a system is asked, as it now routinely is, to judge whether a physical procedure is safe for the person carrying it out.
Sources
- 01Merleau-Ponty, Maurice · 1945 · Phénoménologie de la perception (Phenomenology of Perception) · Gallimard
- 02Dreyfus, Hubert L. · 1972 · What Computers Can't Do: A Critique of Artificial Reason · Harper & Row; revised as What Computers Still Can't Do, MIT Press, 1992
- 03Moravec, Hans · 1988 · Mind Children: The Future of Robot and Human Intelligence · Harvard University Press
- 04Brooks, Rodney A. · 1991 · Intelligence without representation · Artificial Intelligence 47(1–3), 139–159
- 05Barsalou, Lawrence W. · 1999 · Perceptual symbol systems · Behavioral and Brain Sciences 22(4), 577–660
- 06Varela, Francisco J.; Thompson, Evan; Rosch, Eleanor · 1991 · The Embodied Mind: Cognitive Science and Human Experience · MIT Press
- 07Bender, Emily M.; Gebru, Timnit; McMillan-Major, Angelina; Shmitchell, Shmargaret · 2021 · On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? · Proceedings of FAccT '21, 610–623
- 08Landau, Barbara; Gleitman, Lila R. · 1985 · Language and Experience: Evidence from the Blind Child · Harvard University Press