05

The dated field

As of 24 July 2026

The landscape in July 2026

The dominant production reality is a light factory: autonomous execution, attended approval. True no-review production exists, but remains rare, expensive, and difficult to generalize.

01

Light dominates production

Primary1,300+

Stripe Minion PRs / week

Agent-written, but human-reviewed before shipping.

Stripe engineering, 20 Feb 2026
Primary35%

Cursor internal merged PRs

Created by cloud agents; human review remains the bottleneck.

Cursor CEO, Apr 2026
Reported~1,000

Spotify PRs / 10 days

Honk operates atop years of migration automation and CI.

QCon / InfoQ, Mar 2026

These are large throughput figures, not evidence that the approval gate disappeared. Autonomy sits in execution. Accountability still sits with people.

02

A recognizable stack has formed

01

Specifications

Spec Kit, OpenSpec, BMAD, Kiro

02

Harnesses

Codex, Claude Code, Devin, Droids, OpenHands

03

Orchestration

Ralph loops, graphs, state machines, agent teams

04

Verification

CI, holdouts, critics, scenario suites, human review

Framework choice is becoming less differentiating. The durable advantage lies in project context, control flow, evaluation data, and verification design.

03

StrongDM is the dark reference

StrongDM’s Software Factory is the clearest public Level-5 case. Its code is neither written nor reviewed by humans. The crucial innovations are external to the generator: held-out scenario testing, behavioral service twins, and probabilistic “satisfaction” rather than trust in a passing unit suite.

04

The honest state is unsettled

METR’s 2025 randomized study found experienced developers 19% slower with early-2025 AI, even while they believed they were faster. In February 2026, METR cautioned that follow-up selection effects made its current estimate unreliable. The right conclusion is not “AI slows developers.” It is that perception is a poor measurement system.

Likewise, saturated coding benchmarks overstate mergeability. Novel and discriminative task sets produce much lower results. Measure useful changes in your own environment: cycle time, review effort, escaped defects, and recovery cost.