01
Light dominates production
Stripe Minion PRs / week
Agent-written, but human-reviewed before shipping.
Stripe engineering, 20 Feb 2026Cursor internal merged PRs
Created by cloud agents; human review remains the bottleneck.
Cursor CEO, Apr 2026Spotify PRs / 10 days
Honk operates atop years of migration automation and CI.
QCon / InfoQ, Mar 2026These are large throughput figures, not evidence that the approval gate disappeared. Autonomy sits in execution. Accountability still sits with people.
02
A recognizable stack has formed
Specifications
Spec Kit, OpenSpec, BMAD, Kiro
Harnesses
Codex, Claude Code, Devin, Droids, OpenHands
Orchestration
Ralph loops, graphs, state machines, agent teams
Verification
CI, holdouts, critics, scenario suites, human review
Framework choice is becoming less differentiating. The durable advantage lies in project context, control flow, evaluation data, and verification design.
03
StrongDM is the dark reference
StrongDM’s Software Factory is the clearest public Level-5 case. Its code is neither written nor reviewed by humans. The crucial innovations are external to the generator: held-out scenario testing, behavioral service twins, and probabilistic “satisfaction” rather than trust in a passing unit suite.
04
The honest state is unsettled
METR’s 2025 randomized study found experienced developers 19% slower with early-2025 AI, even while they believed they were faster. In February 2026, METR cautioned that follow-up selection effects made its current estimate unreliable. The right conclusion is not “AI slows developers.” It is that perception is a poor measurement system.
Likewise, saturated coding benchmarks overstate mergeability. Novel and discriminative task sets produce much lower results. Measure useful changes in your own environment: cycle time, review effort, escaped defects, and recovery cost.