Keep walking, and watch the floor while you do. The presses fall silent first. The pallets go next, then the forklifts, then the walls. Between roughly 1990 and the day before yesterday, the economy’s center of gravity moved from transforming material to transforming information, and the factory we have been standing in dissolved into offices, then into software, then into laptops that could be anywhere. The work did not stop being real. It stopped being visible.
Walk a software organization’s floor the way we walked Station 1 and the trained eye finds nothing to train on: rows of people at multi-screen desks, all postures identical, whether the person is creating the year’s most valuable feature or re-fighting a build system. The value stream still exists — someone wants something; eventually software does it — but it now runs through repositories, tickets, pipelines, and pull requests, none of which have a place you can stand. For the discipline this street teaches, that is a crisis. And the industry answered it with the most natural move in the world: if you cannot see the work, build a screen that claims to show it. Burndown charts. Velocity. Deployment frequency. The dashboard is the office worker’s window onto a floor that no longer exists.
When the pandemic scattered even the offices, the lean community briefly faced the question head-on: a whole literature of “virtual gemba walks” appeared, asking whether the discipline survives a webcam. The sharpest answer came from inside the tradition. Michael Ballé — a veteran gemba coach — had already answered a question about walking IT remotely with, in effect, no: lean is a hands-on sport, you have to be there, context is everything. That objection comes from someone who has walked more real floors than anyone writing this site, and it stands against everything the rest of this walk will claim. The next station answers it directly instead of working around it.
The seduction
The dashboard’s case for itself is genuinely strong, so hear it out. Its most thoughtful version says: flow metrics are the software gemba — work item ages, queue lengths, lead times are the digital equivalents of the piles and walks we saw at Station 1, so watching the metrics is watching the floor. There is real insight here, and this walk takes from it what is true: the numbers can tell you where something is wrong. What they cannot tell you is what is actually happening — the metric can say a work item sat for nine days; it cannot say the reviewer no longer trusts the author, which is the fact a walker would learn in one conversation. Metrics tell you where to walk. They are not the walk.
Left alone with the decisions, dashboards fail in two well-documented ways. The first has a name — Goodhart’s law: when a measure becomes a target, it ceases to be a good measure. Story points inflate, lines of code multiply, coverage becomes theater; the numbers improve as the thing they measured quietly degrades. The second failure has a fruit: the watermelon metric, green on the outside, red on the inside — every indicator healthy, because every indicator has learned to be, while the project rots undetected. Neither failure requires anyone to lie. The dashboard truthfully reports numbers that have stopped meaning what the dashboard’s readers believe they mean.
The canonical modern skirmish over all this came in 2023, when McKinsey announced that developer productivity could, after all, be measured — with contribution analyses and velocity benchmarks, already in use at nearly twenty companies. Kent Beck and Gergely Orosz wrote the response the moment demanded; Beck’s summary was that the report was too absurd and naive to critique in detail — and the two of them then spent two careful essays critiquing it in detail, arguing that measuring effort and output invites exactly the gaming Goodhart predicts, while what matters — outcomes, impact — resists the spreadsheet. Measurement of effort versus observation of reality: the fight this street has been about since 1950, restaged with modern logos.
The newest floor, the oldest trap
Then AI arrived in the codebase, and the dashboard got its greatest subject matter yet: adoption. The current research keeps sounding the same warning — the sources are on the colophon’s shelf.
Stack Overflow’s 2025 developer survey found adoption still climbing while trust in the tools’ output falls: more developers now actively distrust the accuracy of AI output than trust it. The adoption dashboard shows the first line and has no cell for the second.
GitHub’s documentation for its own Copilot impact dashboard warns that measuring adoption by engagement alone “looks like success” while most users remain stalled at basic completions. When the instrument’s manufacturer tells you the instrument flatters, believe the manufacturer.
Sharpest of all, a randomized trial by METR: experienced developers using AI on real tasks believed they were about 20 percent faster, while the measurements showed them slower — a gap of roughly forty points between felt speed and actual speed. A follow-up study revised the measured slowdown to near zero; what stands is the gap between what the developers felt and what the clock said. Self-report is a dashboard too, and it was green.
So stand still one more time, here at the coldest point of the walk, and feel the shape of the problem. The discipline says: go to where the work actually happens, and see it. The work is now a stream of tokens between a developer and a machine — produced in seconds, vanishing into a scroll buffer, invisible on every dashboard the organization owns.
If the discipline is go and see — where do you stand?