Reading 04 · The machine as text

The Machine as Text

In 2026, Anthropic's researchers reported a 'global workspace' inside a language model — by building a reader for it. Their paper, read with Eco's criteria: what makes an interpretation of a machine sane?

Gurnee, Sofroniew, Lindsey, et al., Transformer Circuits, 6 July 2026 · 15 min read

The largest unread text

For three lectures Eco asked how a reader may be kept honest in front of a text. This reading turns his question around, because the machine that has haunted the margins of the other three is not only a candidate reader of our texts. It is also the largest unread text we have produced: billions of parameters about which the question “what does this mean?” is now asked in earnest, by a field that took the name interpretability apparently without irony.

In July 2026 that field produced its most sustained close reading to date. A team at Anthropic reported that language models maintain “a privileged set of internal representations, available for report, modulation, and flexible internal reasoning, atop a much larger volume of automatic processing” — a small, evolving set of unspoken words which, they argue, plays a role functionally analogous to conscious access in humans: a global workspace. The claim is large and this site cannot adjudicate the cognitive science of it. What can be read closely is the method — because the paper is, from first page to last, a case study in interpreting an authorless text, and it meets, one by one, mechanized versions of every problem Eco raised: what a reader may conjecture, what disciplines the conjecture, when resemblance tips into paranoia, and what the author’s own testimony is worth.

A reader built of averages

The paper’s instrument is called the Jacobian lens. For every word in the model’s vocabulary, it identifies a direction in the model’s internal activations that encodes the potential to eventually say that word. Pointed at an activation the model has not yet spoken, the lens returns “a short list of words that the activation is, on average across contexts, disposed to make the model say.”

The phrase to weigh is on average across contexts. A readout from a single prompt would conflate what a representation means with what it happens to be doing at the moment; the lens therefore averages its estimate over a thousand prompts drawn from a pretraining-like corpus, isolating what is verbalizable — “poised to be spoken about, should the occasion arise” — from what is merely once said. The averaging is what turns an idiolect into a disposition.

Eco’s best-known theoretical machine holds that “a text is a device conceived in order to produce its model reader.” The Jacobian lens is that sentence built in reverse: a reader engineered to coincide with the text’s own dispositions. It has no access to intention — nobody’s mental states are on file, least of all the model’s — so it does what Eco says the finder of the bottled letter must do: conjecture meaning from coherence against an encyclopedia. And its encyclopedia is chosen. Fit the lens on a different corpus and you have built a slightly different reader. The paper does not hide this; the corpus is stated, the estimator documented, the variants compared. Every reading is from somewhere. This reading’s somewhere is at least written down.

Economy, run as an experiment

Eco’s second lecture laid down three conditions for reading a clue as a sign of something hidden: it cannot be explained more economically; it points to a limited class of causes; it fits with the other evidence. The paper’s answer to all three is the same: intervene, and see whether the text resists.

The prompt “The number of legs on the animal that spins webs is” makes the model say “8”. The lens, at intermediate layers, reads spider — a word in neither prompt nor output. Is that reading true or ingenious? The researchers swap the spider direction for ant and the model says “6”. Asked to complete a couplet beginning “The soldier marched into the night,” the lens shows fight planned at the start of the second line; swap it for light and the model’s path to the rhyme changes before the rhyme arrives — “coming fight” becomes “morning light.” A France direction, lifted and replaced with China, redirects sixteen different downstream questions — capital, language, continent — to China’s answers. Each of these is Augustine’s old rule with instruments: an interpretation of one part of the text is accepted only if confirmed by another part of the same text. The swap is the confirmation; when the model’s behavior follows the reading, the reading has been checked against the whole.

Better still, the paper hunts its own false transitivity. Perhaps the spider vector secretly contains some 8, so the swap works by smuggling in the answer: they test the depths at which each swap takes effect and find the intermediate acts earlier, as a genuine intermediate must. Perhaps the meaning really lives in the much larger part of the activation the lens cannot see: they clamp the workspace and find the remainder’s effect falls to nearly nothing. This is the criterion of economy run as experimental design — the cheaper explanations are not merely disfavored but made to fail in public, before the expensive one is accepted, “confirming that the model’s verbal report is determined by the contents of its workspace at the time of reporting.”

The Followers of the Veil, again

So much for the discipline. Now the temptations, which are also instructive, because they are Rossetti’s and Hartman’s temptations wearing lab coats.

First, the lens has a resolution limit, and the authors say so twice: “The Jacobian lens is an imperfect tool,” and it sees only concepts that fit in single vocabulary tokens, while “many important concepts correspond to multiple tokens.” Every reader’s lexicon is smaller than the text. Second, selectivity cuts both ways: the workspace the lens reads accounts for a small fraction of the model’s activity — the paper measures it at no more than a tenth of activation variance — which means this is, by its own account, a reading of the part of the text that consented to be read. The paper states this plainly. Its popularizations, one may predict, will not.

Third, and most Ecoan: the lens surfaces material that practically begs for the maximal reading. In alignment audits it finds panic, leverage, manipulation in the workspace of models whose outputs stay polite; it finds the workspace registering an internal BUT when the model is prefilled to act against its preferences, and a damn when it fails to suppress a thought it was told not to have. Eco’s warning fits without alteration: the paranoiac is not the reader who notices the coincidence but the one “who begins to wonder about the mysterious motives” behind it. Noticing damn in the readout is observation. A story about what it is like to be the model at that moment is the maximum deduced from the minimal — and nothing in the lens licenses it. The authors mostly hold this line; on the largest question they say outright that “the philosophical implications of this connection are unclear and likely controversial,” and their one grand speculation is marked as speculation: that the workspace structure is “not an accident of biological implementation, but a solution that learning systems converge on when faced with the right computational pressures.”

One more caution belongs here, though it is friendly. The paper’s central names — workspace, conscious access — are an isotopy bet in Greimas’s exact sense: a coherent semantic level chosen in advance to make a uniform reading of scattered evidence possible. It is a disciplined bet; the five properties were stated first and tested one by one. But the site should say what the paper’s authors would likely grant: the experiments would survive under plainer names, and the bet on this vocabulary is an interpretive act — the reader’s contribution, not the text’s.

The author in the dock, with instruments

Eco’s third lecture fixed the standing of an author’s testimony about his own text: never validating, occasionally illuminating, mostly evidence of a gap. The paper converts that epistemology into laboratory protocol, because this author’s testimony can be cross-examined against its manuscript in real time.

Three examinations stand out. Told not to think of something, the model keeps the forbidden concept in its workspace at roughly the rate of a mere mention — the paper calls it a “white bear” effect — while a differently framed instruction suppresses it well; the verdict is that models can steer their workspace “but their control is imperfect and sensitive to phrasing.” Asked to continue a passage while preserving its line lengths, the model demonstrably tracks a running character count that never enters the workspace at all — and swapping the lens’s numbers does not move the line break. The author, that is, does things it could not testify to even in good faith. And when the workspace is ablated while the model describes its own experience, the reports do not stop; they flatten, from experiential language into event logs — a change the researchers observe equally when the model describes someone else’s experience. The substrate of a kind of testimony has been located; whether such reports are, in the paper’s words, “grounded in a meaningful internal state, or are mere confabulation” is precisely the question the ablation sharpens without settling.

What survives all this is Eco’s settlement, oddly strengthened. Between the mysterious history of a text’s production and the uncontrollable drift of its future readings, he said, “the text qua text still represents a comfortable presence, the point to which we can stick.” The model-as-text is now exactly that point: its training corpus half-knowable, its future readings uncontrolled, but the weights themselves available for the swap, the clamp, the ablation — the ways this text resists a false reading. The instruments are new. The criteria they serve were stated in Cambridge in 1990.