Index

Prediction registry

Falsifiable claims, each with a resolution criterion a stranger could adjudicate and a resolve-by date. Claims are never deleted or renamed — only their status changes — so anchors (/predictions/#<id>) are stable and safe to link. Current entries derive from The Jagged Frontier, whose essays link back to each claim.

20 OPEN 0 RESOLVED CORRECT 0 RESOLVED INCORRECT 0 REVISED 0 AMBIGUOUS

Open (20)

jf-pretraining-plateau-02 open conf medium

Through 2028, headline frontier capability gains will come predominantly from post-training RL, test-time compute, and tool use — not from pretraining-scale increases.

From The Jagged Frontier · What scaling actually bought

Resolves
by 2028-12-31
Criteria
Technical reports and credible third-party analyses of major 2026–2028 frontier releases attribute the main gains to post-training or inference-time methods; no release achieves a generational jump credited chiefly to parameter scaling. Disputed attribution resolves to ambiguous.
REVISIONS
  • 2026-08-02 — Supporting evidence noted; status unchanged. (DeepSeek-V4-Flash-0731 (July 31 2026) is the cleanest instance to date of a headline capability jump attributed entirely to post-training: the model card states it "keeps the same model architecture and size" and was "only re-post-trained," yet the build reportedly beats DeepSeek's own V4-Pro-Preview on all nine published agent/coding benchmarks (e.g. DeepSWE 7.3 → 54.4). With parameters and architecture explicitly held fixed, the gain is credited to post-training rather than pretraining scale — directly on-point for this claim. It bears on the claim without resolving it: the numbers are unverified vendor figures on one release (no independent reproduction as of July 31), and one lab's post-training gain does not establish the field-wide through-2028 pattern the claim asserts. Left open; flagged under Watch Next.)
scaling capabilities
jf-human-premium-04 open conf medium

By 2028, at least two of the three largest music-streaming platforms will offer explicit AI-content labeling or verified-human filtering.

From The Content Factory · What falls, and what does not

Resolves
by 2028-12-31
Criteria
Platform feature documentation or press releases of the three largest music-streaming services by subscribers (e.g. Spotify, YouTube Music, Apple/Amazon Music) document a user-facing AI-content label or verified-human filter in production.
REVISIONS
  • 2026-06-14 — Supporting regulatory context noted; status unchanged. (The EU Commission's final Code of Practice on marking and labelling AI-generated content (June 10 2026), operationalizing AI Act Article 50 transparency obligations from August 2 2026, creates a regulatory tailwind toward content labelling across generative-AI providers in the EU. It does not establish the music-streaming-specific feature this claim requires (a user-facing label or verified-human filter on the three largest services), so the criterion remains unmet — but the direction of travel supports it.)
  • 2026-08-02 — Supporting regulatory context updated; status unchanged. (The AI Act's Article 50 transparency obligations took effect August 2 2026 — the labelling tailwind noted in the June 14 revision is now binding law, with fines up to EUR 15M or 3% of turnover. One qualification: the AI Omnibus provisional agreement (May 2026) defers the machine-readable marking requirement (Art. 50(2)) to December 2 2026 for generative-AI systems already on the market, softening the near-term push. The obligation still targets AI-content marking in general, not the music-streaming-specific user-facing label or verified-human filter this claim requires, so the criterion remains unmet; the tailwind is now real but partially deferred.)
culture music
jf-genre-fiction-falls-07 open conf medium

By 2028, AI-generated works will constitute a majority of new releases in at least one major serialized genre-fiction market (e.g. Kindle Unlimited romance or LitRPG).

From The Content Factory · What falls, and what does not

Resolves
by 2028-12-31
Criteria
Platform disclosure or credible third-party analysis (publishing-industry research) showing AI-generated titles above 50% of new releases in at least one major serialized genre-fiction category.
culture publishing
jf-productivity-paradox-12 open conf medium

Through 2028, large-scale software-delivery telemetry (DORA-class) will continue to show individual AI coding gains far exceeding organization-level delivery gains.

From The Specification Is the Work · The amplifier

Resolves
by 2028-12-31
Criteria
DORA annual reports and comparable telemetry (Faros-class, 10k+ developers) through 2028 show individual task/PR throughput gains substantially exceeding organization-level delivery-performance gains. If the gap closes (org-level gains comparable to individual gains), resolves incorrect.
REVISIONS
  • 2026-06-14 — Supporting evidence noted; status unchanged. (Anthropic's "When AI builds itself" (Favaro & Clark, June 4 2026) states Claude wrote more than 80% of code merged into Anthropic's production systems, while self-reported productivity gains remain 20-40% (per the Redwood Research calibration). A high throughput/merge share alongside far smaller realized productivity gains is exactly the individual-vs-organization wedge this claim tracks — but it is one firm's internal figure, not DORA-class telemetry, so it informs rather than resolves.)
software productivity
jf-replication-flood-20 open conf medium-high

By 2028, the verifier-overload signal in scientific publishing will have grown — annual retractions (or paper-mill detections) exceed the 2024 level by at least 50%.

From The Verifier and Nature · What the survey has been mapping

Resolves
by 2028-12-31
Criteria
Retraction Watch database or Crossref retraction counts: calendar-year 2028 total at least 1.5× the 2024 total, or equivalent growth in documented paper-mill detections.
science publishing
jf-jagged-persists-01 open conf high

Frontier AI systems in 2030 will still show a jagged profile — superhuman on some tasks while failing tasks most adults find trivial — with no expert-consensus "human-equivalent" moment having occurred.

From The Jagged Frontier · The coastline, not the wall

Resolves
by 2030-12-31
Criteria
As of 2030-12-31: documented, reproducible failures of frontier models on tasks most adults find trivial still exist, AND no consensus in credible expert surveys (majority of surveyed AI researchers) that human-level generality was reached. Both conditions required for resolved-correct.
REVISIONS
  • 2026-06-13 — Supporting evidence noted; status unchanged. (DeepMind's From AGI to ASI (Genewein et al., 2026) independently adopts the jagged framing, stating that capability profiles of concrete systems "may well be jagged w.r.t. human-level intelligence" and that AI progress "may equally be jagged and non-uniform" (Remark III, citing Morris et al. 2026). A heavyweight establishment source endorsing the conceptual frame, though it does not settle the 2030 condition.)
  • 2026-07-04 — Supporting evidence noted; status unchanged. (Grace et al., Thousands of AI Authors on the Future of AI (AI Impacts, 2023; arXiv 2401.02843), the largest expert survey of its kind (2,778 published researchers), places the aggregate median for high-level machine intelligence at 2047 — no majority expects human-level generality anywhere near 2030. This directly supports the second criterion (no expert-survey consensus that human-level generality was reached), while the first criterion (documented trivial-task failures) remains the operative open question through 2030.)
  • 2026-07-26 — Supporting evidence noted; status unchanged. (Bengio's Digitalist Papers essay "Advanced AI as a Global Public Good and a Global Risk" (Dec 2025) states it is "likely that such an uneven distribution across cognitive abilities will continue without a distinct AGI moment" — a leading safety-camp researcher independently adopting the no-human-equivalent-moment framing, and recommending policymakers track specific high-risk capabilities instead of an AGI threshold. Conceptual support; does not settle the 2030 conditions.)
agi capabilities
jf-strip-miner-05 open conf medium

No genuinely new mass-culture genre or form (scale-comparable to the cinematic universe, true-crime podcasts, K-pop) will be widely credited as originated by an autonomous AI engagement-optimization loop by 2030.

From The Content Factory · Three ways the loop bends

Resolves
by 2030-12-31
Criteria
Absence of any case widely credited in mainstream cultural criticism where an autonomous AI optimization loop originated a new mass-culture genre or form. A disputed case resolves to ambiguous.
culture
jf-prestige-resists-06 open conf medium-high

Through 2030, no predominantly AI-generated live-action series or film will earn a major-category Emmy/Oscar nomination or a top-10 annual slot on a major streaming service.

From The Content Factory · What falls, and what does not

Resolves
by 2030-12-31
Criteria
Award nomination records and platform annual top-10 lists through 2030. "Predominantly AI-generated" means AI generated the majority of footage and performances, per credits or credible reporting.
culture film
jf-world-model-games-09 open conf medium-high

By 2030, no commercially shipped game built on a frame-generating world model (no conventional engine or assets) will sustain persistent, rule-governed play for 10+ hours.

From The Machine That Must Behave · World models

Resolves
by 2030-12-31
Criteria
Requires a commercial release (paid or ad-supported, not a research preview), technical reporting confirming a frame-generation architecture without a conventional engine, and documented persistent rule-governed state across sessions totalling 10+ hours.
games world-models
jf-ai-uncopyrightable-11 open conf medium-high

Through 2030, works generated by AI without human creative contribution will remain uncopyrightable in the United States.

From The Machine That Must Behave · Assets

Resolves
by 2030-12-31
Criteria
US Copyright Office policy and controlling court decisions as of 2030-12-31 still deny copyright to works with no human creative contribution. A statutory change or controlling precedent granting such copyright resolves the claim incorrect.
law games culture
jf-avionics-verifier-14 open conf high

Through 2030, no civil-aviation certification authority will accept AI-generated flight-control code without human-auditable requirement-to-test traceability; AI's certified role stays on the verification side.

From The Specification Is the Work · The gradient

Resolves
by 2030-12-31
Criteria
FAA/EASA certification guidance and the documented development process of any certified system through 2030. A certified flight-control system whose code was AI-generated without human-auditable traceability resolves the claim incorrect.
software safety-critical
jf-erp-unmoved-15 open conf medium

By 2030, large ERP implementation outcomes (overrun and failure rates) will remain within their historical bands despite AI coding tools.

From The Specification Is the Work · The gradient

Resolves
by 2030-12-31
Criteria
Industry implementation surveys (Panorama ERP report or successor): reported overrun and failure rates for large ERP implementations through 2030 show no step-change improvement versus 2020–2024 baselines.
software enterprise
jf-math-acceleration-16 open conf medium-high

By 2030, AI-assisted methods will have resolved at least 50 long-open named mathematics problems (Erdős-problem class) with machine-verified proofs.

From The Verifier and Nature · Mathematics: the one perfect verifier

Resolves
by 2030-12-31
Criteria
Count problems open for at least 10 years, resolved with substantive AI involvement, with a machine-verified (e.g. Lean) proof — per the Erdős problems database, formal-proof repositories, and credible mathematical reporting. Threshold: 50 by 2030-12-31.
REVISIONS
  • 2026-08-16 — Supporting evidence noted; status unchanged. (Two AI-for-math results land squarely on this claim's subject. OpenAI's Astra (Aug 2 2026) published solutions to ten problems open for a decade or more — including the first explicit construction of a non-sofic group (Gromov 1999) — with Lean 4 proof certificates carrying a zero 'sorry' count (fully machine-verified), for ~$2,000 of API spend; Thomas Bloom, who debunked OpenAI's Oct 2025 Erdős claim, rated it more significant than the May 2026 unit-distance result. That is up to ten problems meeting the criterion's substance (long-open, AI involvement, machine-verified) — meaningful partial progress toward the 50-by-2030 threshold, not resolution. Claude's Riemann-zeta bound improvement (41.6% to 67.2%, Aug 11-12 2026) is trajectory evidence rather than a countable instance: it is a bound improvement, not a resolved named problem, and is unpublished/unreviewed and not machine-verified. Left open; the Astra proofs are the concrete anchor, with ~4 years and a large gap to 50 remaining.)
science math
jf-math-direction-17 open conf medium

Through 2030, the selection of which research-level mathematical problems to pursue will remain human — no AI system autonomously poses and resolves a problem the field judges significant.

From The Verifier and Nature · Mathematics: the one perfect verifier

Resolves
by 2030-12-31
Criteria
No case through 2030 where credible mathematicians attribute both the choice of problem and its resolution to an autonomous system. Disputed cases resolve to ambiguous.
REVISIONS
  • 2026-06-13 — Supporting evidence noted; status unchanged. (DeepMind's From AGI to ASI (Genewein et al., 2026) classifies AI achievements to date — Move 37, automated theorem proving, AlphaFold — as exploratory creativity within human-provided conceptual spaces (Boden levels 1-2), and frames transformative creativity and the autonomous origination of significant problems as the unmet hallmark of ASI. Consistent with this claim that problem-selection stays human through 2030; conceptual support, not resolution.)
  • 2026-08-16 — Supporting evidence noted; status unchanged. (The two August 2026 AI-for-math results support the human-selection half of this claim by their structure. OpenAI's Astra (Aug 2) solved ten problems that were already open and human-posed — e.g. Gromov's 1999 soficity question — and Claude's Riemann-zeta run (Aug 11-12) advanced a bound on a problem the researchers set it to work on. In both cases the impressive step is resolution/advancement, not the origination of the question: humans chose which problems to pursue, AI did the work. No case yet where credible mathematicians attribute both the choice of a significant problem and its resolution to an autonomous system. Supporting evidence for problem-selection staying human; status unchanged.)
science math
jf-dev-jobs-rotate-13 open conf medium-high

US software-developer employment in 2030 will be no more than 20% below its 2025 level — the job rotates toward specification and verification rather than disappearing.

From The Specification Is the Work · Where the factory's reach ends

Resolves
by 2031-06-30
Criteria
BLS Occupational Employment Statistics for the "Software Developers" category: 2030 headcount no more than 20% below the 2025 figure. Resolution mid-2031 to allow for data publication lag.
software labor
jf-bio-calendar-19 open conf high

As of 2030, clinical development will not have dramatically compressed — median time from first-in-human to approval stays above 6 years and overall clinical success rate below 25%.

From The Verifier and Nature · Biology: the worst judge of all

Resolves
by 2031-06-30
Criteria
Standard industry analyses (BIO/Citeline/FDA-based) of programs concluding by 2030: median first-in-human-to-approval time > 6 years AND overall likelihood of approval from Phase 1 < 25%. Resolution mid-2031 for data lag. Both conditions required for resolved-correct.
science biology medicine
jf-aaa-moat-08 open conf high

No predominantly AI-generated game will achieve AAA-scale success (Metacritic ≥ 85 and multi-million unit sales) by 2032.

From The Machine That Must Behave · The irreducibility moat

Resolves
by 2032-12-31
Criteria
"Predominantly AI-generated" means design, code, and content majority-generated autonomously, per credits or credible reporting. Checked against Metacritic score (≥ 85) and sales data (≥ 2 million units). Any single qualifying title resolves the claim incorrect.
games
jf-faust-canon-03 open conf high

No disclosed AI-authored literary work will win a top-tier literary prize (Booker, Pulitzer, Nobel, or national equivalent) by 2035.

From The Jagged Frontier · Why Faust is the right test

Resolves
by 2035-12-31
Criteria
Prize records: no top-tier literary award given to a work publicly known at the time of the award to be predominantly AI-authored.
culture literature
jf-physics-stall-18 open conf high

No AI-originated theory of fundamental physics will receive experimental confirmation by 2035.

From The Verifier and Nature · Physics: a field that splits in two

Resolves
by 2035-12-31
Criteria
No experimentally confirmed beyond-Standard-Model (or comparable fundamental) theory whose origination is credibly attributed to an AI system, as of 2035-12-31.
REVISIONS
  • 2026-06-13 — Supporting evidence noted; status unchanged. (DeepMind's From AGI to ASI (Genewein et al., 2026) gives a mechanism for this expectation: the abstraction barrier and embodied bottleneck argue models trained on human data lack a demonstrated way to discover novel conceptual primitives, and Hassabis's cited test — could an AI have originated general relativity from 1900s knowledge, "today the answer is no" — frames AI-originated fundamental physics as still out of reach. Conceptual support for the stall, not resolution.)
science physics