The walk kept its pockets deliberately empty: at most one external link per station, so the street never leaked its readers. This back room holds everything else. Three shelves, one sentence each on what a thing is actually good for.
The lean shelf — the Canon
- Taiichi Ohno, Toyota Production System: Beyond Large-Scale Production (1978; English 1988). The origin of go-and-see, the seven wastes, and — perhaps apocryphally — the chalk circle; read it for how blunt the original actually is.
- James P. Womack, Gemba Walks, Expanded 2nd Edition (LEI, 2013). The eponymous text: 72 short essays from actual walks, including the ones where he corrects himself — the epistemic model this site tries to imitate.
- Mike Rother, Toyota Kata (2009). The daily-observation habit turned into a teachable coaching routine — what “practice” means once the walk becomes a job.
- Jeffrey Liker, The Toyota Way (2004; 2nd ed. 2021). The genchi genbutsu chapter is the fullest doctrinal treatment of standing in the circle.
- Mary & Tom Poppendieck, Implementing Lean Software Development (2006). Where “go, see, and confirm” entered the software-lean vocabulary — aimed at seeing customers, the ancestor of Station 5’s move.
- Gene Kim, Jez Humble, Patrick Debois, John Willis, The DevOps Handbook (2016; 2nd ed. 2021). The Toyota lineage carried into deployment pipelines; the andon cord’s software afterlife.
- Mik Kersten, Project to Product (2018). The strongest version of the metrics-first position this site argues with at Station 4 — read it so you know the opposing case from its best advocate, not from our summary.
The nearest neighbors — credited precedents
- “What We Learned from a Year of Building with LLMs” (Yan, Bischof, Frye, Husain, Liu, Shankar, 2024). Contains the single nearest sentence to this site’s thesis — production input–output pairs as “the genchi genbutsu of LLM applications.”
- Birgitta Böckeler, “Harness engineering for coding agent users” (martinfowler.com, April 2026). The canonical vocabulary for Layer 2 — guides, sensors, the harness as an ongoing engineering practice, and “our name is on the commit.”
- Kent Beck, “Augmented Coding: Beyond the Vibes” (2025). A canonical figure watching the “genie” work and cataloguing the warning signs — the practitioner’s trained eye, without the lean framing.
- Simon Willison, “Vibe engineering” (October 2025). The accountability end of the AI-coding spectrum, named — staying answerable for what ships.
- METR, “Analyzing coding agent transcripts to upper bound productivity gains from AI agents” (February 2026). Transcript-reading as measurement discipline, from the eval direction — 5,305 real sessions, read so the numbers mean something.
- Anthropic, “Demystifying evals for AI agents”. Recommends reading transcripts as a critical, ongoing skill — the AI lab arriving at Ohno’s conclusion from the opposite side of the street.
- Michael Ballé, “In IT, can we do virtual gemba walks?” (Lean Enterprise Institute). The best objection to this entire site, from inside the tradition — engaged head-on at Station 5, and listed here so you can check our reading.
The metrics skirmish
- Kent Beck & Gergely Orosz, “Measuring developer productivity? A response to McKinsey” — and Part 2 (2023). The canonical modern case that measuring effort invites gaming while outcomes resist the spreadsheet.
- Stack Overflow, 2025 Developer Survey — AI section. Primary source for Station 4’s first data point: 84 percent adoption, 46 percent distrust.
- GitHub, Copilot impact dashboard. Primary source for the second: the vendor’s own warning that engagement alone signals stalling, not success.
- METR, “Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity” (July 2025). Primary source for the third: felt 20 percent faster, measured 19 percent slower — since revised toward zero by a 2026 follow-up, which Station 4 reports.
- Goodhart’s law. “When a measure becomes a target, it ceases to be a good measure” — the one-line theory under Station 4’s whole argument; the engineering literature rediscovers it roughly monthly, with story points as the usual casualty.
- The watermelon metric (ProjectManagement.com, “The Watermelon Effect”). Green on the outside, red on the inside — the fruit Station 4 borrowed, and the fate of any dashboard whose indicators have learned to be healthy.
The making of this site
This site was drawn with AI and checked by Jörn — the footer says so on every page, and the claim is meant literally in both halves. The prose was written by a language model working inside a harness of its own: a content specification, a voice prompt, a prior-art research document with verified sources, and independent review passes — a fact we note with exactly the amount of irony you suspect. Every statistic and every named precedent traces to the research shelf above; nothing numerical was written from the model’s memory.
Go and See is part of a larger catalog of AI-drawn sites, several of which this walk crossed: the knowledge problem, the dark factory, and building software with AI each have streets of their own.
The exit is behind you — the same door as the entrance. The street outside has a yellow line on it somewhere. It leads to your own floor.