Does Deep Learning Refute Scott?
It is the custom of every intellectual generation to discover that its predecessors were naive, and the generation now writing about artificial intelligence has discovered it with unusual speed. Among its favorite inheritances is James C. Scott’s Seeing Like a State, from which it has drawn a simple and satisfying lesson: formal systems flatten the world, tacit knowledge escapes them, and any technology that runs on abstractions must therefore do violence to the local, the contextual, and the human. One finds this lesson applied to algorithms roughly once a week, generally under a title containing the words “Seeing Like a Platform.”
The lesson has only one defect, which is that its central premise has quietly become false. It is the purpose of this essay to say so plainly, and then to rescue what remains — which is, as it happens, the more important half.
What Scott actually claimed
Scott’s argument deserves to be stated carefully, since it is usually invoked rather than read. States, he observed, can only govern what they can see, and so they remake the world to be seeable: standard surnames, uniform land registers, monoculture forests, cities on grids. The instrument of this remaking is the schema — a simplified formal representation that keeps the features useful to power and discards the rest. What the schema discards Scott called metis: the practical, adaptive, largely unwritable knowledge of practitioners — how this river floods, when this soil tires, which neighbor’s promise is good.
Half a century before Scott’s cadastral maps, the same argument had appeared in economic dress. Friedrich Hayek’s 1945 essay “The Use of Knowledge in Society” argued that the knowledge that matters to an economy exists as dispersed, tacit, local fragments — “knowledge of the particular circumstances of time and place” — which no central authority can assemble, because the assembling destroys it. Scott is Hayek’s knowledge problem in anthropological costume; that the two men would have disagreed about nearly everything else is one of the small ironies that make the history of ideas worth reading. Both rest on a single load-bearing premise: tacit knowledge cannot be centralized. You may employ the man who knows the harbor, but you cannot extract the harbor from the man.
For two centuries of administrative technology, the premise held. Every instrument the centralizers built — the census, the survey, the standard form, the relational database — could store only what could first be made explicit. The schema came first; whatever fit it, survived; whatever did not, vanished. On this premise the entire critical tradition rests.
The premise fails
Now consider what a large language model is. It is not a schema that the messy world was forced into. It is a statistical compression of the mess itself — trillions of words of working notes, forum answers, code reviews, kitchen improvisations, harbor lore — written by millions of practitioners for one another, in their own unstandardized idiom. Nobody designed its categories, and nobody could: its “categories” are whatever regularities the text of human practice happens to contain.
The results are familiar to anyone who has used the technology honestly. Ask a model why a sourdough starter behaves differently in a humid kitchen, how to calm a particular kind of anxious customer, what an error message usually means as opposed to what it says — and you receive something that behaves remarkably like metis: context-sensitive, hedged, alive to the exceptions. It follows that the first technology in history has appeared which ingests tacit knowledge without first making it explicit. The extraction that Hayek declared impossible and Scott declared destructive is now performed, imperfectly but at planetary scale, by gradient descent.
To this the traditional critic replies that the model does not really understand the harbor. Perhaps not; the question is interesting and this essay will not settle it. But the reply misses the practical point. Scott’s argument never depended on the state’s comprehension — the Prussian forestry office understood nothing; the damage came from what its instruments could and could not carry. The new instruments carry the mess. A critique aimed at the limitations of the carrier must reckon with a carrier that no longer has those limitations, and the honest conclusion is uncomfortable: the naive claim that “AI destroys context” is dead, and its defenders should stop making it. An argument kept alive after its premise has died persuades no one but its owners, and discredits the sounder arguments standing next to it.
The relocation
Where, then, did the legibility problem go? It did not dissolve; it changed address — from the representation to the supply chain — and Kate Crawford’s Atlas of AI is, in effect, its new cadastral survey.
Consider what must be made legible in order that the model may remain illiterate in the old, blessed sense. The workers who label training data are managed through piecework platforms that score them in fractions of a cent, their judgments forced into taxonomies precisely as rigid as any Prussian forest register. The text of human practice is scraped under terms no practitioner negotiated, catalogued, deduplicated, filtered — an enclosure of the written commons conducted with the same serene confidence as the enclosures of land that gave Scott’s states their fields. The minerals are mined, the electricity metered, the data centers sited, each through administrative schemata of the classical kind. And when the model’s fluent output is bolted into a bureaucracy — a benefits office, a content moderation pipeline, a hiring funnel — it is wrapped in exactly the scores, thresholds, and appeal-proof categories that this site’s other essays examine. The mess is in the middle; the grid closes over both ends.
It follows that Scott requires not refutation but a change of address. The violence of simplification, which the critical tradition kept looking for inside the representation, now operates in the industrial process that produces the representation and in the institutional process that consumes it. One may put the revised thesis in a single sentence: deep learning did not abolish the legibility problem; it relocated the problem from the map to the map-making economy.
Nor has the claim shrunk in the move; it has arguably grown, for the old schemas were at least public — one could read the census categories and object to them. A supply chain is private, distributed across jurisdictions, and visible chiefly to those who own it. The cadastral map has, so to speak, gone underground, and critics who continue to inspect the surface will report, quite accurately, that they cannot find it.
What remains of the old religion
The reader may reasonably ask what practical difference the relocation makes. Three follow at once.
First, the target of reform changes. If the trouble were in the representation, the remedy would be better representations — richer categories, “context-aware” models — and the industry would be delighted to sell them to us. If the trouble is in the supply chain, the remedies are labor law for labelers, provenance and consent for training corpora, and public accounting of material costs: dull instruments, which is generally the mark of real ones.
Second, the evaluative question changes. One should ask of any AI deployment not “does the model understand the context?” — it increasingly does, or counterfeits understanding to a standard the question cannot detect — but “what had to be gridded so that this fluency could be delivered here?” The answer is an inventory of labelers, scrapes, mines, and scores, and it differs from deployment to deployment in ways that permit actual judgment.
Third, a prediction, since a thesis that risks nothing is worth nothing. If the relocation view is right, the harms that surface over the coming years will cluster not around models flattening meaning — the translation will keep getting better — but around the two gridded ends: disputes over what was taken to train, and injuries done by the institutional casings into which the fluent middle is fitted. If instead the characteristic scandals of the next decade are semantic — machines catastrophically misreading context in ways their supply chains cannot explain — then Scott’s original address was correct after all, and the relocation argued here will have been a detour: an error one could hardly regret discovering, since it would mean the maps were once again lying where everyone could inspect them.
The intelligent position, in short, is neither the triumphalism that declares the knowledge problem solved nor the nostalgia that repeats a critique whose premise has expired. The state has finally learned to read our handwriting. The urgent question is what it costs everyone downstream of the reading lesson — and that question belongs to Scott still, provided his heirs consent to follow the problem to its new address.