The register
Every figure in the chapters resolves to an entry below. The register records how good each source is, and the build fails if a claim points at nothing, which is what makes that sentence a guarantee rather than an intention.
The tier is never typed into the prose. Each chip looks its tier up here, so a source can be re-graded in one place and every chip that cites it follows.
Each tier also says what it licenses you to conclude; a citation without that line invites you to treat a press release the way you would treat a paper.
- primary
- The organisation's own publication, or the paper itself. Licenses: quote it as that organisation's own account, which is not the same as independently true.
- reputable
- Established secondary reporting. Licenses: trust the shape, check the digits.
- vendor-reported
- Self-reported by the seller, with no independent measurement. Licenses: a direction at most; never repeat it as a fact.
- contested
- Disputed, or retracted by its own authors. Licenses: citing it only with its dispute attached. On this site it appears that way or not at all. Read the note first.
- workspace
- Repositories on the author's own machine, cited by path and read date. Real evidence, and not independently verifiable by you — the two facts travel together, and the pages that use it say so. Licenses: a worked example, never proof.
primary · 35 entries
| ID | Source | Date |
|---|---|---|
| src-0001 |
Software Factories, Light and Dark The light/dark frame, comprehension debt, back pressure, and the outer loop. Chapters 1, 3, 5, 6 and 8 paraphrase its argument. |
2026-07-22 |
| src-0002 |
Harness Engineering is not Enough: Why Software Factories Fail AI Engineer World's Fair talk. Source of the four-month dark-factory report and the three-to-ten-step rule. The talk date is the conference window; the research document does not record the exact day. |
2026-06-01 |
| src-0003 |
12-Factor Agents Own your control flow. The repository predates the talk; the register records the project rather than a release. |
2025-01-01 |
| src-0004 |
The Five Levels: from Spicy Autocomplete to the Dark Factory The autonomy ladder, modelled on the NHTSA/SAE driving-automation levels. Shapiro places himself at Level 4. |
2026-01-23 |
| src-0006 |
Exploring Generative AI: harness engineering Agent = Model + Harness. Guides and sensors, computational and inferential, the steering loop, three regulation categories. |
2026-04-02 |
| src-0007 |
Anatomy of an agent harness Böckeler attributes the Agent = Model + Harness formulation here. |
2026-01-01 |
| src-0008 |
AI Harness Engineering: A Runtime Substrate for Foundation-Model Software Agents Argues the software-engineering gap is a harness problem (context, tools, verification, rollback), not merely a model-capability problem. Preprint, not peer-reviewed. |
2026-03-01 |
| src-0009 |
Ralph Wiggum as a software engineer (the Ralph loop) Infinite shell loop, one prompt file per iteration, filesystem as memory, fresh context each cycle. |
2025-07-01 |
| src-0010 |
Back pressure engineering, Gas Town, and evolutionary software By January 2026 Huntley demonstrated a system that auto-healed a production bug with no human intervention. |
2026-01-01 |
| src-0014 |
Spec Kit — spec-driven development toolkit Constitution, specify, plan; 80k+ stars; supports 30+ coding agents. Star counts are a mid-2026 snapshot. |
2025-10-01 |
| src-0015 |
BMAD-METHOD — Breakthrough Method for Agile AI-Driven Development Simulates an agile team of 12–21 specialised agents producing versioned artifacts. ~37k stars. |
2025-06-01 |
| src-0017 |
OpenSpec — lightweight brownfield change management |
2025-09-01 |
| src-0018 |
Kiro — spec-driven IDE, general availability Popularised EARS requirements notation. |
2026-01-01 |
| src-0019 |
Claude Code — subagents, skills, hooks, worktrees Subagents cannot spawn sub-subagents; practitioners report 3–5 concurrent subagents and 4–8 worktrees per developer. |
2026-04-01 |
| src-0020 |
Technology Radar v33 and v34 Claude Code moved to Adopt. Warns of cognitive debt, semantic diffusion, AI-accelerated shadow IT, and complacency with AI-generated code. |
2026-04-01 |
| src-0025 |
Codex — cloud sandbox agent, CLI and desktop apps |
2026-02-01 |
| src-0026 |
Antigravity 2.0 — agent-first IDE; 76.2% on SWE-bench Verified Relaunched at Google I/O 2026. Ships desktop app, Go CLI, SDK and a Managed Agents API. |
2026-05-19 |
| src-0027 |
OpenHands — open-source autonomous platform, CodeAct paradigm Reports roughly 72–81% on SWE-bench Verified with a frontier model; ships a proprietary critic model. ~81k stars. |
2026-01-01 |
| src-0028 |
SWE-agent and mini-swe-agent mini-swe-agent is roughly 100 lines of Python and scores above 74% on SWE-bench Verified: a standing argument that scaffolding beats scaffolding complexity. |
2026-01-01 |
| src-0030 |
Minions, part 2 — over 1,300 agent-authored PRs merged per week, all human-reviewed Five-layer blueprint of deterministic and agentic nodes; isolated cloud devboxes; 3M tests and 500 MCP tools; a two-round CI limit that bails out to humans. Built on a decade of prior tooling investment, a transferability caveat the write-up makes itself. |
2026-02-20 |
| src-0032 |
35% of internally merged PRs are created by cloud agents The February 2026 computer-use launch post cited 'more than 30%'. Code review remains the human bottleneck; 'self-driving codebases' is explicitly aspirational. |
2026-04-01 |
| src-0033 |
Software Factories and the Agentic Moment The reference dark factory. Two rules: code must not be written by humans, and code must not be reviewed by humans. Founded July 2025. |
2026-02-01 |
| src-0034 |
Scenario testing as holdout sets, 'satisfaction', and the Digital Twin Universe End-to-end user stories stored outside the codebase so agents cannot overfit them; a probabilistic success metric replacing boolean 'tests pass'; behavioural clones of Okta, Jira, Slack and Google Workspace as self-contained Go binaries. |
2026-02-01 |
| src-0035 |
Token spend benchmark: at least $1,000 per day per engineer Roughly $20,000 per engineer per month. Simon Willison flagged this as the make-or-break question for whether the pattern generalises beyond well-capitalised teams. |
2026-02-01 |
| src-0036 |
AI/works — agentic delivery platform, and the 3-3-3 model Idea to production in 90 days; partnership with Mechanical Orchard for mainframe renewal. |
2026-01-20 |
| src-0037 |
'The inflection point isn't so much about technology — it's about technique' |
2026-01-20 |
| src-0039 |
Frontier agent capability on discriminative subsets falls from ~73% to ~11% Discriminative subsets are tasks where a benchmark can distinguish real capability from memorisation. SWE-bench-Live resolution on novel issues runs ~18–20%. |
2026-02-01 |
| src-0040 |
'Verification debt' — the accumulating cost of unverified AI code |
2025-11-01 |
| src-0043 |
Claude Code review tooling: ~20 minutes per PR, issues flagged in 84% of large PRs (avg 7.5 findings) Anthropic also reports code output per engineer grew 200% in a year: the generator of the review bottleneck it sells into. |
2026-01-01 |
| src-0046 |
A utility model for verification findings Findings should maximise P(correct) × C_saved − C_human_verification − P(incorrect) × C_false_alarm. |
2026-01-01 |
| src-0050 |
Why SWE-bench Verified no longer measures frontier coding capabilities |
2026-02-01 |
| src-0051 |
SWE-bench: 2,294 real GitHub issues across 12 Python repositories SWE-bench Verified is OpenAI's human-filtered subset. |
2024-05-01 |
| src-0053 |
Terminal-Bench — command-line task benchmark |
2026-04-01 |
| src-0054 |
Content Agent — every change staged as a draft 'Nothing publishes without a human hitting approve': the same gate placement as a code factory, on a content pipeline. |
2026-01-01 |
| src-0056 |
Agent Factory blog series and Agent 365 A parallel coinage: 'agent factory' is Microsoft's enterprise-agent blueprint, not the software factory this site describes. Jay Parikh framed a shift 'from software factory to agent factory'. |
2025-08-13 |
reputable · 18 entries
| ID | Source | Date |
|---|---|---|
| src-0005 |
Amplification of the Five Levels essay |
2026-01-28 |
| src-0011 |
Kubernetes for agents / the Molecular Expression of Work |
2026-01-01 |
| src-0012 |
Lights-out manufacturing at FANUC and Xiaomi FANUC has run unattended robot plants since 2001; Xiaomi opened a lights-out smartphone plant in 2024. The name 'dark factory' comes from here: dark because robots do not need to see. |
2024-01-01 |
| src-0013 |
Mass Produced Software Components (the origin of 'software factory') 'Software Factory' was trademarked by Systems Development Corporation in 1974; the AI-era usage is a repurposing, not a coinage. |
1968-10-01 |
| src-0016 |
Reported monthly API cost for heavy multi-agent SDD frameworks $800–2,000 per developer per month on frontier models. Community-reported, spread over varied usage patterns; directional. |
2026-01-01 |
| src-0021 |
Series D at $26B post-money; ARR $37M → $492M in twelve months Co-led by Lux Capital, General Catalyst and 8VC. Revenue figures are company-disclosed at fundraise. |
2026-05-27 |
| src-0022 |
Series C, $150M at $1.5B post-money Led by Khosla Ventures. A separate company from Cognition, frequently confused with it. |
2026-04-16 |
| src-0029 |
LangGraph, CrewAI, AutoGen/AG2 — orchestration framework comparisons Recurring conclusion across several 2026 comparisons: the framework is rarely the differentiator; harness and context design is. |
2026-03-01 |
| src-0031 |
Rewriting All of Spotify's Code Base, All the Time (Honk / Fleet Management) 1,500+ PRs in nine months, then ~1,000 every ten days; 60–90% time savings on migrations. Deterministic scripts still handle ~70% of migration volume. The LLM judge, which initially vetoed ~25% of sessions, was removed by March 2026 as models improved. |
2026-03-01 |
| src-0041 |
2026 State of Code Developer Survey (1,149 developers) 96% do not fully trust AI-generated code; only 48% always verify before committing; AI accounts for 42% of committed code; 38% say reviewing AI code takes more effort than reviewing human code. |
2026-01-08 |
| src-0042 |
Cost of review: agent ≈ $0.05 per PR, human ≈ $15–25 per PR No single publication is the source: the figures are collated from several 2026 cost analyses via the research document. The spread, not the exact figures, is the load-bearing part, and both ends move with token prices and salaries. |
2026-02-01 |
| src-0044 |
20–40% of AI review comments are false positives Collated from several practitioner reports rather than one publication. Located by the research document, not by a permalink. |
2026-01-01 |
| src-0045 |
AI-generated code introduced over 10,000 new security findings per month by June 2025 A tenfold increase from December 2024. |
2025-09-01 |
| src-0047 |
$70M Series B for AI code verification |
2026-01-01 |
| src-0049 |
Built by Agents, Tested by Agents, Trusted by Whom? The liability counterpoint: who answers for unreviewed security code. |
2026-02-08 |
| src-0052 |
Mid-2026 SWE-bench Verified leaderboard, roughly 80–94% One reported 93.9% result had roughly 20% of its 'solved' cases judged semantically wrong on inspection, which is the reason the top of this leaderboard is not a capability measurement. |
2026-06-01 |
| src-0055 |
State of AI Agents 2026 — the 3–7 agent sweet spot Researcher → writer → critic → publisher, with human review at write and side-effect boundaries. |
2026-01-01 |
| src-0065 |
Forward Deployed Engineering programme with ServiceNow Also partnerships with Google Cloud and WaveMaker on two-pass deterministic code generation. |
2026-05-01 |
vendor-reported · 3 entries
| ID | Source | Date |
|---|---|---|
| src-0023 |
Droid scores 58.75% on Terminal-Bench v0.1.1 State of the art at the time of the claim, on the vendor's own run. |
2025-09-01 |
| src-0024 |
31x faster feature delivery, 96% shorter migration times Company-reported marketing figures with no independent measurement. Treat as directional at most. |
2026-04-16 |
| src-0038 |
Expected efficiency gains of ~35% in development and ~50% in support EPAM projections presented at Knowledge 2026, not measured outcomes. |
2026-05-01 |
contested · 1 entries
| ID | Source | Date |
|---|---|---|
| src-0048 |
Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity 16 developers, 246 tasks: allowing AI increased completion time by 19%, against a forecast 24% speedup and a post-hoc estimate of 20% — a ~39-point perception gap. In February 2026 METR said selection effects made the follow-up unreliable and that its honest position is now 'we don't know'. Cite the perception gap, not the 19%. |
2025-07-10 |
workspace · 10 entries
| ID | Source | Date |
|---|---|---|
| src-0057 |
website-factory design note — the floor / judge / feel gate split The deciding axis between the factories is not the stack but what the checker can be. |
2026-07-11 |
| src-0058 |
dark-website-factory design note — which gates were darkened and what replaced each Two lights stay on: the brief and the publish. Everything between them is decided by agents. |
2026-07-25 |
| src-0059 |
Dark factory retro — three runs, gate-by-gate catches and escapes Across three runs the darkened gates caught 30+ real defects; confirmed content escapes to the human: zero, in the single run that reached a human review. The retro states its own n=1 caveat. |
2026-07-25 |
| src-0060 |
The baseurl sweep — a guard that existed for nine days before anyone knew it was needed Three sites shipping 56 dead internal links between them; the check that catches all of them had been in the template for nine days. One of the three had a link check the whole time that could not see the bug by construction. |
2026-07-25 |
| src-0061 |
first-factory — the spec-and-verify software factory Six gated stages, three human gates, worker agents that read a Commands table rather than hard-coding build tools. |
2026-07-16 |
| src-0062 |
demo-factory — the iterate-and-judge visual demo studio Where the screenshot judge and the capped-rounds pattern were first proved. |
2026-07-10 |
| src-0063 |
visual-novel-generator — deterministic story graph, judged prose Branch completeness and asset wiring gated hard; prose quality judged with capped rounds and a human read-through. |
2026-07-16 |
| src-0064 |
This site's own build — scaffold and loop records Lights On, Lights Off was itself produced by the lit pipeline it documents, through the stages named on the About page. |
2026-07-26 |
| src-0066 |
The five-variant benchmark — one brief, five sites, archived before unpublishing Built output and home-page captures of the five sites produced from the same brief by different model/effort pairs (Fable baseline, Opus 5 high, GPT-5.6 medium, GPT-5.6 high, Opus 5 medium). Archived at merge time because the four superseded sites were scheduled to go dark, leaving this the only record the benchmark happened. |
2026-07-26 |
| src-0067 |
Merge spec — one site from five variants The comparative verdicts (which variant won which dimension and why), the graft list, and the conflict decisions, written before the merge session and treated by it as binding. |
2026-07-26 |
What the attestation covers
A machine enforces that every claim in the chapters points at a registered entry, and that every entry is well formed: an immutable id, an ISO date, a known tier, and a locator. That is a real guarantee and the build fails without it.
One boundary the legend above cannot carry alone: the tiers describe the original source, not this site's reading of it. A primary stamp means the underlying account is the organisation's own; it does not mean the sentence next to the chip has been independently re-verified against that account. Where a figure matters to a decision you are making, go to the source.
No machine enforces that an entry is fairly characterised. Whether a source is summarised honestly is exactly the check that can be faked out by the same process being checked — which is chapter 8's last row, arriving on this site's own doorstep. So the colophon claims the first and not the second, and this paragraph exists so the difference is on the page rather than in a footnote.