The register

Every figure in the chapters resolves to an entry below. The register records how good each source is, and the build fails if a claim points at nothing, which is what makes that sentence a guarantee rather than an intention.

The tier is never typed into the prose. Each chip looks its tier up here, so a source can be re-graded in one place and every chip that cites it follows.

Each tier also says what it licenses you to conclude; a citation without that line invites you to treat a press release the way you would treat a paper.

primary
The organisation's own publication, or the paper itself. Licenses: quote it as that organisation's own account, which is not the same as independently true.
reputable
Established secondary reporting. Licenses: trust the shape, check the digits.
vendor-reported
Self-reported by the seller, with no independent measurement. Licenses: a direction at most; never repeat it as a fact.
contested
Disputed, or retracted by its own authors. Licenses: citing it only with its dispute attached. On this site it appears that way or not at all. Read the note first.
workspace
Repositories on the author's own machine, cited by path and read date. Real evidence, and not independently verifiable by you — the two facts travel together, and the pages that use it say so. Licenses: a worked example, never proof.
On locators · the tier says how good a source is; the link says where to find it, and the two are kept independent. This site was assembled from two research documents that record titles, authors and dates, not always permalinks. Where a source has a publisher, the entry carries that publisher's canonical domain and the real locator is the title, author and date. Where a figure was collated from several reports with no single publication behind it, the entry points into this repository's own research document instead. A URL that looked more precise than the evidence behind it would be the exact failure this register exists to prevent.

primary · 35 entries

IDSourceDate
src-0001 Software Factories, Light and Dark
Addy Osmani · addyosmani.substack.comThe light/dark frame, comprehension debt, back pressure, and the outer loop. Chapters 1, 3, 5, 6 and 8 paraphrase its argument.
2026-07-22
src-0002 Harness Engineering is not Enough: Why Software Factories Fail
Dex Horthy, HumanLayer · humanlayer.devAI Engineer World's Fair talk. Source of the four-month dark-factory report and the three-to-ten-step rule. The talk date is the conference window; the research document does not record the exact day.
2026-06-01
src-0003 12-Factor Agents
Dex Horthy, HumanLayer · github.comOwn your control flow. The repository predates the talk; the register records the project rather than a release.
2025-01-01
src-0004 The Five Levels: from Spicy Autocomplete to the Dark Factory
Dan Shapiro · danshapiro.comThe autonomy ladder, modelled on the NHTSA/SAE driving-automation levels. Shapiro places himself at Level 4.
2026-01-23
src-0006 Exploring Generative AI: harness engineering
Birgitta Böckeler, Thoughtworks / martinfowler.com · martinfowler.comAgent = Model + Harness. Guides and sensors, computational and inferential, the steering loop, three regulation categories.
2026-04-02
src-0007 Anatomy of an agent harness
LangChain · blog.langchain.devBöckeler attributes the Agent = Model + Harness formulation here.
2026-01-01
src-0008 AI Harness Engineering: A Runtime Substrate for Foundation-Model Software Agents
arXiv preprint · arxiv.orgArgues the software-engineering gap is a harness problem (context, tools, verification, rollback), not merely a model-capability problem. Preprint, not peer-reviewed.
2026-03-01
src-0009 Ralph Wiggum as a software engineer (the Ralph loop)
Geoffrey Huntley · ghuntley.comInfinite shell loop, one prompt file per iteration, filesystem as memory, fresh context each cycle.
2025-07-01
src-0010 Back pressure engineering, Gas Town, and evolutionary software
Geoffrey Huntley · ghuntley.comBy January 2026 Huntley demonstrated a system that auto-healed a production bug with no human intervention.
2026-01-01
src-0014 Spec Kit — spec-driven development toolkit
GitHub · github.comConstitution, specify, plan; 80k+ stars; supports 30+ coding agents. Star counts are a mid-2026 snapshot.
2025-10-01
src-0015 BMAD-METHOD — Breakthrough Method for Agile AI-Driven Development
BMAD project · github.comSimulates an agile team of 12–21 specialised agents producing versioned artifacts. ~37k stars.
2025-06-01
src-0017 OpenSpec — lightweight brownfield change management
OpenSpec project · github.com
2025-09-01
src-0018 Kiro — spec-driven IDE, general availability
AWS · kiro.devPopularised EARS requirements notation.
2026-01-01
src-0019 Claude Code — subagents, skills, hooks, worktrees
Anthropic · docs.claude.comSubagents cannot spawn sub-subagents; practitioners report 3–5 concurrent subagents and 4–8 worktrees per developer.
2026-04-01
src-0020 Technology Radar v33 and v34
Thoughtworks · thoughtworks.comClaude Code moved to Adopt. Warns of cognitive debt, semantic diffusion, AI-accelerated shadow IT, and complacency with AI-generated code.
2026-04-01
src-0025 Codex — cloud sandbox agent, CLI and desktop apps
OpenAI · openai.com
2026-02-01
src-0026 Antigravity 2.0 — agent-first IDE; 76.2% on SWE-bench Verified
Google · antigravity.googleRelaunched at Google I/O 2026. Ships desktop app, Go CLI, SDK and a Managed Agents API.
2026-05-19
src-0027 OpenHands — open-source autonomous platform, CodeAct paradigm
All Hands AI · github.comReports roughly 72–81% on SWE-bench Verified with a frontier model; ships a proprietary critic model. ~81k stars.
2026-01-01
src-0028 SWE-agent and mini-swe-agent
Princeton / Stanford · github.commini-swe-agent is roughly 100 lines of Python and scores above 74% on SWE-bench Verified: a standing argument that scaffolding beats scaffolding complexity.
2026-01-01
src-0030 Minions, part 2 — over 1,300 agent-authored PRs merged per week, all human-reviewed
Stripe Engineering · stripe.comFive-layer blueprint of deterministic and agentic nodes; isolated cloud devboxes; 3M tests and 500 MCP tools; a two-round CI limit that bails out to humans. Built on a decade of prior tooling investment, a transferability caveat the write-up makes itself.
2026-02-20
src-0032 35% of internally merged PRs are created by cloud agents
Michael Truell, Cursor · cursor.comThe February 2026 computer-use launch post cited 'more than 30%'. Code review remains the human bottleneck; 'self-driving codebases' is explicitly aspirational.
2026-04-01
src-0033 Software Factories and the Agentic Moment
StrongDM · strongdm.comThe reference dark factory. Two rules: code must not be written by humans, and code must not be reviewed by humans. Founded July 2025.
2026-02-01
src-0034 Scenario testing as holdout sets, 'satisfaction', and the Digital Twin Universe
StrongDM · strongdm.comEnd-to-end user stories stored outside the codebase so agents cannot overfit them; a probabilistic success metric replacing boolean 'tests pass'; behavioural clones of Okta, Jira, Slack and Google Workspace as self-contained Go binaries.
2026-02-01
src-0035 Token spend benchmark: at least $1,000 per day per engineer
StrongDM · strongdm.comRoughly $20,000 per engineer per month. Simon Willison flagged this as the make-or-break question for whether the pattern generalises beyond well-capitalised teams.
2026-02-01
src-0036 AI/works — agentic delivery platform, and the 3-3-3 model
Thoughtworks · thoughtworks.comIdea to production in 90 days; partnership with Mechanical Orchard for mainframe renewal.
2026-01-20
src-0037 'The inflection point isn't so much about technology — it's about technique'
Rachel Laycock, CTO, Thoughtworks · thoughtworks.com
2026-01-20
src-0039 Frontier agent capability on discriminative subsets falls from ~73% to ~11%
IBM Research · research.ibm.comDiscriminative subsets are tasks where a benchmark can distinguish real capability from memorisation. SWE-bench-Live resolution on novel issues runs ~18–20%.
2026-02-01
src-0040 'Verification debt' — the accumulating cost of unverified AI code
Werner Vogels, CTO, AWS · allthingsdistributed.com
2025-11-01
src-0043 Claude Code review tooling: ~20 minutes per PR, issues flagged in 84% of large PRs (avg 7.5 findings)
Anthropic · anthropic.comAnthropic also reports code output per engineer grew 200% in a year: the generator of the review bottleneck it sells into.
2026-01-01
src-0046 A utility model for verification findings
OpenAI · openai.comFindings should maximise P(correct) × C_saved − C_human_verification − P(incorrect) × C_false_alarm.
2026-01-01
src-0050 Why SWE-bench Verified no longer measures frontier coding capabilities
OpenAI · openai.com
2026-02-01
src-0051 SWE-bench: 2,294 real GitHub issues across 12 Python repositories
Princeton / Stanford, ICLR 2024 · swebench.comSWE-bench Verified is OpenAI's human-filtered subset.
2024-05-01
src-0053 Terminal-Bench — command-line task benchmark
ICLR 2026 · tbench.ai
2026-04-01
src-0054 Content Agent — every change staged as a draft
Sanity · sanity.io'Nothing publishes without a human hitting approve': the same gate placement as a code factory, on a content pipeline.
2026-01-01
src-0056 Agent Factory blog series and Agent 365
Yina Arenas, Microsoft · azure.microsoft.comA parallel coinage: 'agent factory' is Microsoft's enterprise-agent blueprint, not the software factory this site describes. Jay Parikh framed a shift 'from software factory to agent factory'.
2025-08-13

reputable · 18 entries

IDSourceDate
src-0005 Amplification of the Five Levels essay
Simon Willison · simonwillison.net
2026-01-28
src-0011 Kubernetes for agents / the Molecular Expression of Work
Steve Yegge · steve-yegge.medium.com
2026-01-01
src-0012 Lights-out manufacturing at FANUC and Xiaomi
Manufacturing press, various · fanuc.co.jpFANUC has run unattended robot plants since 2001; Xiaomi opened a lights-out smartphone plant in 2024. The name 'dark factory' comes from here: dark because robots do not need to see.
2024-01-01
src-0013 Mass Produced Software Components (the origin of 'software factory')
Doug McIlroy, NATO Software Engineering Conference · cs.dartmouth.edu'Software Factory' was trademarked by Systems Development Corporation in 1974; the AI-era usage is a repurposing, not a coinage.
1968-10-01
src-0016 Reported monthly API cost for heavy multi-agent SDD frameworks
Practitioner reports, various · github.com$800–2,000 per developer per month on frontier models. Community-reported, spread over varied usage patterns; directional.
2026-01-01
src-0021 Series D at $26B post-money; ARR $37M → $492M in twelve months
Cognition (Devin) · cognition.aiCo-led by Lux Capital, General Catalyst and 8VC. Revenue figures are company-disclosed at fundraise.
2026-05-27
src-0022 Series C, $150M at $1.5B post-money
Factory.ai · factory.aiLed by Khosla Ventures. A separate company from Cognition, frequently confused with it.
2026-04-16
src-0029 LangGraph, CrewAI, AutoGen/AG2 — orchestration framework comparisons
Practitioner comparisons, various · langchain-ai.github.ioRecurring conclusion across several 2026 comparisons: the framework is rarely the differentiator; harness and context design is.
2026-03-01
src-0031 Rewriting All of Spotify's Code Base, All the Time (Honk / Fleet Management)
Jo Kelly-Fenton and Aleksandar Mitic, QCon London, reported by InfoQ · infoq.com1,500+ PRs in nine months, then ~1,000 every ten days; 60–90% time savings on migrations. Deterministic scripts still handle ~70% of migration volume. The LLM judge, which initially vetoed ~25% of sessions, was removed by March 2026 as models improved.
2026-03-01
src-0041 2026 State of Code Developer Survey (1,149 developers)
Sonar · sonarsource.com96% do not fully trust AI-generated code; only 48% always verify before committing; AI accounts for 42% of committed code; 38% say reviewing AI code takes more effort than reviewing human code.
2026-01-08
src-0042 Cost of review: agent ≈ $0.05 per PR, human ≈ $15–25 per PR
Industry cost analyses collated in this site's research documents · specs/software_factories_mid_2026.md, read 2026-07-26No single publication is the source: the figures are collated from several 2026 cost analyses via the research document. The spread, not the exact figures, is the load-bearing part, and both ends move with token prices and salaries.
2026-02-01
src-0044 20–40% of AI review comments are false positives
Practitioner measurements collated in this site's research documents · specs/software_factories_mid_2026.md, read 2026-07-26Collated from several practitioner reports rather than one publication. Located by the research document, not by a permalink.
2026-01-01
src-0045 AI-generated code introduced over 10,000 new security findings per month by June 2025
Apiiro · apiiro.comA tenfold increase from December 2024.
2025-09-01
src-0047 $70M Series B for AI code verification
Qodo · qodo.ai
2026-01-01
src-0049 Built by Agents, Tested by Agents, Trusted by Whom?
Stanford Law CodeX · law.stanford.eduThe liability counterpoint: who answers for unreviewed security code.
2026-02-08
src-0052 Mid-2026 SWE-bench Verified leaderboard, roughly 80–94%
Leaderboard aggregation · swebench.comOne reported 93.9% result had roughly 20% of its 'solved' cases judged semantically wrong on inspection, which is the reason the top of this leaderboard is not a capability measurement.
2026-06-01
src-0055 State of AI Agents 2026 — the 3–7 agent sweet spot
a16z · a16z.comResearcher → writer → critic → publisher, with human review at write and side-effect boundaries.
2026-01-01
src-0065 Forward Deployed Engineering programme with ServiceNow
Accenture · accenture.comAlso partnerships with Google Cloud and WaveMaker on two-pass deterministic code generation.
2026-05-01

vendor-reported · 3 entries

IDSourceDate
src-0023 Droid scores 58.75% on Terminal-Bench v0.1.1
Factory.ai · factory.aiState of the art at the time of the claim, on the vendor's own run.
2025-09-01
src-0024 31x faster feature delivery, 96% shorter migration times
Factory.ai · factory.aiCompany-reported marketing figures with no independent measurement. Treat as directional at most.
2026-04-16
src-0038 Expected efficiency gains of ~35% in development and ~50% in support
EPAM · epam.comEPAM projections presented at Knowledge 2026, not measured outcomes.
2026-05-01

contested · 1 entries

IDSourceDate
src-0048 Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity
METR · arxiv.org16 developers, 246 tasks: allowing AI increased completion time by 19%, against a forecast 24% speedup and a post-hoc estimate of 20% — a ~39-point perception gap. In February 2026 METR said selection effects made the follow-up unreliable and that its honest position is now 'we don't know'. Cite the perception gap, not the 19%.
2025-07-10

workspace · 10 entries

IDSourceDate
src-0057 website-factory design note — the floor / judge / feel gate split
This workspace · ../factories/website-factory/specs/design.md, read 2026-07-26The deciding axis between the factories is not the stack but what the checker can be.
2026-07-11
src-0058 dark-website-factory design note — which gates were darkened and what replaced each
This workspace · ../factories/dark-website-factory/specs/design.md, read 2026-07-26Two lights stay on: the brief and the publish. Everything between them is decided by agents.
2026-07-25
src-0059 Dark factory retro — three runs, gate-by-gate catches and escapes
This workspace · ../factories/dark-website-factory/specs/retro/2026-07-25.md, read 2026-07-26Across three runs the darkened gates caught 30+ real defects; confirmed content escapes to the human: zero, in the single run that reached a human review. The retro states its own n=1 caveat.
2026-07-25
src-0060 The baseurl sweep — a guard that existed for nine days before anyone knew it was needed
This workspace · ../factories/website-factory/specs/retro/2026-07-25.md, read 2026-07-26Three sites shipping 56 dead internal links between them; the check that catches all of them had been in the template for nine days. One of the three had a link check the whole time that could not see the bug by construction.
2026-07-25
src-0061 first-factory — the spec-and-verify software factory
This workspace · ../factories/first-factory/README.md, read 2026-07-26Six gated stages, three human gates, worker agents that read a Commands table rather than hard-coding build tools.
2026-07-16
src-0062 demo-factory — the iterate-and-judge visual demo studio
This workspace · ../factories/demo-factory/README.md, read 2026-07-26Where the screenshot judge and the capped-rounds pattern were first proved.
2026-07-10
src-0063 visual-novel-generator — deterministic story graph, judged prose
This workspace · ../factories/visual-novel-factory/README.md, read 2026-07-26Branch completeness and asset wiring gated hard; prose quality judged with capped rounds and a human read-through.
2026-07-16
src-0064 This site's own build — scaffold and loop records
This workspace · specs/brief.md, read 2026-07-26Lights On, Lights Off was itself produced by the lit pipeline it documents, through the stages named on the About page.
2026-07-26
src-0066 The five-variant benchmark — one brief, five sites, archived before unpublishing
This workspace · specs/archive/, read 2026-07-26Built output and home-page captures of the five sites produced from the same brief by different model/effort pairs (Fable baseline, Opus 5 high, GPT-5.6 medium, GPT-5.6 high, Opus 5 medium). Archived at merge time because the four superseded sites were scheduled to go dark, leaving this the only record the benchmark happened.
2026-07-26
src-0067 Merge spec — one site from five variants
This workspace · specs/ai-factories-merge-spec.md, read 2026-07-26The comparative verdicts (which variant won which dimension and why), the graft list, and the conflict decisions, written before the merge session and treated by it as binding.
2026-07-26

What the attestation covers

A machine enforces that every claim in the chapters points at a registered entry, and that every entry is well formed: an immutable id, an ISO date, a known tier, and a locator. That is a real guarantee and the build fails without it.

One boundary the legend above cannot carry alone: the tiers describe the original source, not this site's reading of it. A primary stamp means the underlying account is the organisation's own; it does not mean the sentence next to the chip has been independently re-verified against that account. Where a figure matters to a decision you are making, go to the source.

No machine enforces that an entry is fairly characterised. Whether a source is summarised honestly is exactly the check that can be faked out by the same process being checked — which is chapter 8's last row, arriving on this site's own doorstep. So the colophon claims the first and not the second, and this paragraph exists so the difference is on the page rather than in a footnote.