Prediction registry
Falsifiable claims, each with a resolution criterion a stranger could adjudicate
and a resolve-by date. Claims are never deleted or renamed — only their status
changes — so anchors (/predictions/#<id>) are stable and safe to link.
Current entries derive from
The Jagged Frontier,
whose essays link back to each claim.
Open (20)
Steam will exceed 25,000 game releases in calendar year 2027.
From The Machine That Must Behave · Mechanics and code
- Resolves
- by 2028-01-31
- Criteria
- SteamDB or equivalent public release counts for calendar year 2027, checked January 2028.
games
markets
Through 2028, headline frontier capability gains will come predominantly from post-training RL, test-time compute, and tool use — not from pretraining-scale increases.
From The Jagged Frontier · What scaling actually bought
- Resolves
- by 2028-12-31
- Criteria
- Technical reports and credible third-party analyses of major 2026–2028 frontier releases attribute the main gains to post-training or inference-time methods; no release achieves a generational jump credited chiefly to parameter scaling. Disputed attribution resolves to ambiguous.
REVISIONS
- 2026-08-02 — Supporting evidence noted; status unchanged. (DeepSeek-V4-Flash-0731 (July 31 2026) is the cleanest instance to date of a headline capability jump attributed entirely to post-training: the model card states it "keeps the same model architecture and size" and was "only re-post-trained," yet the build reportedly beats DeepSeek's own V4-Pro-Preview on all nine published agent/coding benchmarks (e.g. DeepSWE 7.3 → 54.4). With parameters and architecture explicitly held fixed, the gain is credited to post-training rather than pretraining scale — directly on-point for this claim. It bears on the claim without resolving it: the numbers are unverified vendor figures on one release (no independent reproduction as of July 31), and one lab's post-training gain does not establish the field-wide through-2028 pattern the claim asserts. Left open; flagged under Watch Next.)
scaling
capabilities
By 2028, at least two of the three largest music-streaming platforms will offer explicit AI-content labeling or verified-human filtering.
From The Content Factory · What falls, and what does not
- Resolves
- by 2028-12-31
- Criteria
- Platform feature documentation or press releases of the three largest music-streaming services by subscribers (e.g. Spotify, YouTube Music, Apple/Amazon Music) document a user-facing AI-content label or verified-human filter in production.
REVISIONS
- 2026-06-14 — Supporting regulatory context noted; status unchanged. (The EU Commission's final Code of Practice on marking and labelling AI-generated content (June 10 2026), operationalizing AI Act Article 50 transparency obligations from August 2 2026, creates a regulatory tailwind toward content labelling across generative-AI providers in the EU. It does not establish the music-streaming-specific feature this claim requires (a user-facing label or verified-human filter on the three largest services), so the criterion remains unmet — but the direction of travel supports it.)
- 2026-08-02 — Supporting regulatory context updated; status unchanged. (The AI Act's Article 50 transparency obligations took effect August 2 2026 — the labelling tailwind noted in the June 14 revision is now binding law, with fines up to EUR 15M or 3% of turnover. One qualification: the AI Omnibus provisional agreement (May 2026) defers the machine-readable marking requirement (Art. 50(2)) to December 2 2026 for generative-AI systems already on the market, softening the near-term push. The obligation still targets AI-content marking in general, not the music-streaming-specific user-facing label or verified-human filter this claim requires, so the criterion remains unmet; the tailwind is now real but partially deferred.)
culture
music
By 2028, AI-generated works will constitute a majority of new releases in at least one major serialized genre-fiction market (e.g. Kindle Unlimited romance or LitRPG).
From The Content Factory · What falls, and what does not
- Resolves
- by 2028-12-31
- Criteria
- Platform disclosure or credible third-party analysis (publishing-industry research) showing AI-generated titles above 50% of new releases in at least one major serialized genre-fiction category.
culture
publishing
Through 2028, large-scale software-delivery telemetry (DORA-class) will continue to show individual AI coding gains far exceeding organization-level delivery gains.
From The Specification Is the Work · The amplifier
- Resolves
- by 2028-12-31
- Criteria
- DORA annual reports and comparable telemetry (Faros-class, 10k+ developers) through 2028 show individual task/PR throughput gains substantially exceeding organization-level delivery-performance gains. If the gap closes (org-level gains comparable to individual gains), resolves incorrect.
REVISIONS
- 2026-06-14 — Supporting evidence noted; status unchanged. (Anthropic's "When AI builds itself" (Favaro & Clark, June 4 2026) states Claude wrote more than 80% of code merged into Anthropic's production systems, while self-reported productivity gains remain 20-40% (per the Redwood Research calibration). A high throughput/merge share alongside far smaller realized productivity gains is exactly the individual-vs-organization wedge this claim tracks — but it is one firm's internal figure, not DORA-class telemetry, so it informs rather than resolves.)
software
productivity
By 2028, the verifier-overload signal in scientific publishing will have grown — annual retractions (or paper-mill detections) exceed the 2024 level by at least 50%.
From The Verifier and Nature · What the survey has been mapping
- Resolves
- by 2028-12-31
- Criteria
- Retraction Watch database or Crossref retraction counts: calendar-year 2028 total at least 1.5× the 2024 total, or equivalent growth in documented paper-mill detections.
science
publishing
Frontier AI systems in 2030 will still show a jagged profile — superhuman on some tasks while failing tasks most adults find trivial — with no expert-consensus "human-equivalent" moment having occurred.
From The Jagged Frontier · The coastline, not the wall
- Resolves
- by 2030-12-31
- Criteria
- As of 2030-12-31: documented, reproducible failures of frontier models on tasks most adults find trivial still exist, AND no consensus in credible expert surveys (majority of surveyed AI researchers) that human-level generality was reached. Both conditions required for resolved-correct.
REVISIONS
- 2026-06-13 — Supporting evidence noted; status unchanged. (DeepMind's From AGI to ASI (Genewein et al., 2026) independently adopts the jagged framing, stating that capability profiles of concrete systems "may well be jagged w.r.t. human-level intelligence" and that AI progress "may equally be jagged and non-uniform" (Remark III, citing Morris et al. 2026). A heavyweight establishment source endorsing the conceptual frame, though it does not settle the 2030 condition.)
- 2026-07-04 — Supporting evidence noted; status unchanged. (Grace et al., Thousands of AI Authors on the Future of AI (AI Impacts, 2023; arXiv 2401.02843), the largest expert survey of its kind (2,778 published researchers), places the aggregate median for high-level machine intelligence at 2047 — no majority expects human-level generality anywhere near 2030. This directly supports the second criterion (no expert-survey consensus that human-level generality was reached), while the first criterion (documented trivial-task failures) remains the operative open question through 2030.)
- 2026-07-26 — Supporting evidence noted; status unchanged. (Bengio's Digitalist Papers essay "Advanced AI as a Global Public Good and a Global Risk" (Dec 2025) states it is "likely that such an uneven distribution across cognitive abilities will continue without a distinct AGI moment" — a leading safety-camp researcher independently adopting the no-human-equivalent-moment framing, and recommending policymakers track specific high-risk capabilities instead of an AGI threshold. Conceptual support; does not settle the 2030 conditions.)
agi
capabilities
No genuinely new mass-culture genre or form (scale-comparable to the cinematic universe, true-crime podcasts, K-pop) will be widely credited as originated by an autonomous AI engagement-optimization loop by 2030.
From The Content Factory · Three ways the loop bends
- Resolves
- by 2030-12-31
- Criteria
- Absence of any case widely credited in mainstream cultural criticism where an autonomous AI optimization loop originated a new mass-culture genre or form. A disputed case resolves to ambiguous.
culture
Through 2030, no predominantly AI-generated live-action series or film will earn a major-category Emmy/Oscar nomination or a top-10 annual slot on a major streaming service.
From The Content Factory · What falls, and what does not
- Resolves
- by 2030-12-31
- Criteria
- Award nomination records and platform annual top-10 lists through 2030. "Predominantly AI-generated" means AI generated the majority of footage and performances, per credits or credible reporting.
culture
film
By 2030, no commercially shipped game built on a frame-generating world model (no conventional engine or assets) will sustain persistent, rule-governed play for 10+ hours.
From The Machine That Must Behave · World models
- Resolves
- by 2030-12-31
- Criteria
- Requires a commercial release (paid or ad-supported, not a research preview), technical reporting confirming a frame-generation architecture without a conventional engine, and documented persistent rule-governed state across sessions totalling 10+ hours.
games
world-models
Through 2030, works generated by AI without human creative contribution will remain uncopyrightable in the United States.
From The Machine That Must Behave · Assets
- Resolves
- by 2030-12-31
- Criteria
- US Copyright Office policy and controlling court decisions as of 2030-12-31 still deny copyright to works with no human creative contribution. A statutory change or controlling precedent granting such copyright resolves the claim incorrect.
law
games
culture
Through 2030, no civil-aviation certification authority will accept AI-generated flight-control code without human-auditable requirement-to-test traceability; AI's certified role stays on the verification side.
From The Specification Is the Work · The gradient
- Resolves
- by 2030-12-31
- Criteria
- FAA/EASA certification guidance and the documented development process of any certified system through 2030. A certified flight-control system whose code was AI-generated without human-auditable traceability resolves the claim incorrect.
software
safety-critical
By 2030, large ERP implementation outcomes (overrun and failure rates) will remain within their historical bands despite AI coding tools.
From The Specification Is the Work · The gradient
- Resolves
- by 2030-12-31
- Criteria
- Industry implementation surveys (Panorama ERP report or successor): reported overrun and failure rates for large ERP implementations through 2030 show no step-change improvement versus 2020–2024 baselines.
software
enterprise
By 2030, AI-assisted methods will have resolved at least 50 long-open named mathematics problems (Erdős-problem class) with machine-verified proofs.
From The Verifier and Nature · Mathematics: the one perfect verifier
- Resolves
- by 2030-12-31
- Criteria
- Count problems open for at least 10 years, resolved with substantive AI involvement, with a machine-verified (e.g. Lean) proof — per the Erdős problems database, formal-proof repositories, and credible mathematical reporting. Threshold: 50 by 2030-12-31.
REVISIONS
- 2026-08-16 — Supporting evidence noted; status unchanged. (Two AI-for-math results land squarely on this claim's subject. OpenAI's Astra (Aug 2 2026) published solutions to ten problems open for a decade or more — including the first explicit construction of a non-sofic group (Gromov 1999) — with Lean 4 proof certificates carrying a zero 'sorry' count (fully machine-verified), for ~$2,000 of API spend; Thomas Bloom, who debunked OpenAI's Oct 2025 Erdős claim, rated it more significant than the May 2026 unit-distance result. That is up to ten problems meeting the criterion's substance (long-open, AI involvement, machine-verified) — meaningful partial progress toward the 50-by-2030 threshold, not resolution. Claude's Riemann-zeta bound improvement (41.6% to 67.2%, Aug 11-12 2026) is trajectory evidence rather than a countable instance: it is a bound improvement, not a resolved named problem, and is unpublished/unreviewed and not machine-verified. Left open; the Astra proofs are the concrete anchor, with ~4 years and a large gap to 50 remaining.)
science
math
Through 2030, the selection of which research-level mathematical problems to pursue will remain human — no AI system autonomously poses and resolves a problem the field judges significant.
From The Verifier and Nature · Mathematics: the one perfect verifier
- Resolves
- by 2030-12-31
- Criteria
- No case through 2030 where credible mathematicians attribute both the choice of problem and its resolution to an autonomous system. Disputed cases resolve to ambiguous.
REVISIONS
- 2026-06-13 — Supporting evidence noted; status unchanged. (DeepMind's From AGI to ASI (Genewein et al., 2026) classifies AI achievements to date — Move 37, automated theorem proving, AlphaFold — as exploratory creativity within human-provided conceptual spaces (Boden levels 1-2), and frames transformative creativity and the autonomous origination of significant problems as the unmet hallmark of ASI. Consistent with this claim that problem-selection stays human through 2030; conceptual support, not resolution.)
- 2026-08-16 — Supporting evidence noted; status unchanged. (The two August 2026 AI-for-math results support the human-selection half of this claim by their structure. OpenAI's Astra (Aug 2) solved ten problems that were already open and human-posed — e.g. Gromov's 1999 soficity question — and Claude's Riemann-zeta run (Aug 11-12) advanced a bound on a problem the researchers set it to work on. In both cases the impressive step is resolution/advancement, not the origination of the question: humans chose which problems to pursue, AI did the work. No case yet where credible mathematicians attribute both the choice of a significant problem and its resolution to an autonomous system. Supporting evidence for problem-selection staying human; status unchanged.)
science
math
US software-developer employment in 2030 will be no more than 20% below its 2025 level — the job rotates toward specification and verification rather than disappearing.
From The Specification Is the Work · Where the factory's reach ends
- Resolves
- by 2031-06-30
- Criteria
- BLS Occupational Employment Statistics for the "Software Developers" category: 2030 headcount no more than 20% below the 2025 figure. Resolution mid-2031 to allow for data publication lag.
software
labor
As of 2030, clinical development will not have dramatically compressed — median time from first-in-human to approval stays above 6 years and overall clinical success rate below 25%.
From The Verifier and Nature · Biology: the worst judge of all
- Resolves
- by 2031-06-30
- Criteria
- Standard industry analyses (BIO/Citeline/FDA-based) of programs concluding by 2030: median first-in-human-to-approval time > 6 years AND overall likelihood of approval from Phase 1 < 25%. Resolution mid-2031 for data lag. Both conditions required for resolved-correct.
science
biology
medicine
No predominantly AI-generated game will achieve AAA-scale success (Metacritic ≥ 85 and multi-million unit sales) by 2032.
From The Machine That Must Behave · The irreducibility moat
- Resolves
- by 2032-12-31
- Criteria
- "Predominantly AI-generated" means design, code, and content majority-generated autonomously, per credits or credible reporting. Checked against Metacritic score (≥ 85) and sales data (≥ 2 million units). Any single qualifying title resolves the claim incorrect.
games
No disclosed AI-authored literary work will win a top-tier literary prize (Booker, Pulitzer, Nobel, or national equivalent) by 2035.
From The Jagged Frontier · Why Faust is the right test
- Resolves
- by 2035-12-31
- Criteria
- Prize records: no top-tier literary award given to a work publicly known at the time of the award to be predominantly AI-authored.
culture
literature
No AI-originated theory of fundamental physics will receive experimental confirmation by 2035.
From The Verifier and Nature · Physics: a field that splits in two
- Resolves
- by 2035-12-31
- Criteria
- No experimentally confirmed beyond-Standard-Model (or comparable fundamental) theory whose origination is credibly attributed to an AI system, as of 2035-12-31.
REVISIONS
- 2026-06-13 — Supporting evidence noted; status unchanged. (DeepMind's From AGI to ASI (Genewein et al., 2026) gives a mechanism for this expectation: the abstraction barrier and embodied bottleneck argue models trained on human data lack a demonstrated way to discover novel conceptual primitives, and Hassabis's cited test — could an AI have originated general relativity from 1900s knowledge, "today the answer is no" — frames AI-originated fundamental physics as still out of reach. Conceptual support for the stall, not resolution.)
science
physics