AI predictions, scored against reality
AI leaders make claims that move markets, shape regulation and influence how people plan their careers. This ledger preserves those claims, separates forecasts from warnings, and checks the measurable ones when their deadlines arrive.
Version 1.3 · 100 source-verified claims · Last evidence review: September 14, 2026
The first lesson is already clear: the loudest forecasts are often the hardest to score. “AGI,” “smarter than humans,” “most code” and “join the workforce” can each shift meaning after the deadline. A useful prediction states the outcome, scope, date and measurement rule before reality arrives. For the operational evidence behind the agent claims, see Kingy’s State of AI Agents 2026 and the Kingy Score v2 methodology.
The ledger
A 10% chance that AI causes human extinction within the next decade.
This is a probabilistic forecast, not a claim that extinction will occur. It cannot be marked wrong merely because the outcome does not happen; calibration requires a reference class of comparable forecasts.
Source and incentive context
Coxon made the estimate while publicly resigning and said he left before his Anthropic equity vested. That reduces one obvious financial incentive to promote Anthropic, but does not validate the forecast. Sources: Associated Press; Axios interview.
More than a 10% chance that AI kills all humans within the next decade.
The estimate is unusually explicit but still lacks a published model showing how the number was derived. Keep it as a forecast, not a fact about current systems.
Source and incentive context
Hubinger works for a frontier lab whose public case for safety rules can affect regulation and competition. His technical work on deceptive model behavior is relevant evidence, but employment alone proves neither sincerity nor strategy. Sources: Washington Post; Sleeper Agents paper.
“Powerful AI” matching or exceeding Nobel-level experts across most disciplines could arrive in late 2026 or early 2027.
The definition is demanding and more useful than a bare AGI label. A fair resolution still needs agreed tests for expert-level performance, autonomous work and reliability across disciplines.
Source and incentive context
Anthropic sells access to frontier models and also argues for stronger frontier-model oversight. Both interests belong in the record. Sources: Machines of Loving Grace; Anthropic submission summary.
AI could eliminate half of entry-level white-collar jobs within one to five years.
This needs a fixed occupation list and baseline job count. Job openings, employment levels and task automation are different measures and should not be swapped after the deadline.
Source and incentive context
The forecast came from a direct interview. Anthropic benefits when employers believe its systems can perform valuable work, while stronger displacement expectations may also support regulation and workforce programs. Source: Axios.
AI-driven displacement could push US unemployment to 10–20% within one to five years.
This is measurable, but causation matters. A recession that raises unemployment without demonstrated AI displacement would not satisfy the forecast.
Source and incentive context
Record alongside the entry-level-jobs claim, but score separately because the labor-force consequence can fail even if task automation rises. Source: Axios.
AI would write 90% of code within three to six months and nearly all code within a year.
Coding-agent use rose quickly, but no credible industry-wide measurement approached 90% by September 2025. Amodei later said it happened “at least at some places,” which narrows the scope after the fact.
Evidence
Open-source research found more than 320,000 agent-attributed commits per month by April 2026, substantial but not evidence that AI wrote 90% of all code. Sources: claim and follow-up review; repository census.
Powerful AI could compress 50–100 years of biological progress into 5–10 years and help cure or prevent most major diseases.
The clock begins only after Amodei’s “powerful AI” threshold, so the medical claim cannot honestly be assigned a fixed deadline until that trigger is resolved.
The first AI agents could “join the workforce” in 2025 and materially change company output.
Working agents arrived and enterprise adoption followed. The undefined words “join” and “materially” prevent a clean win, and OpenAI’s strongest quantitative evidence comes from its own users and staff.
Evidence
By mid-2026 OpenAI reported that every internal department used Codex as its primary AI work tool, while external adoption remained uneven. Sources: original prediction; OpenAI usage study; research paper.
Systems that can produce novel insights would likely arrive in 2026.
Resolution requires a disclosed, independently validated discovery, not a model generating a plausible new sentence or helping a human researcher work faster.
Source
Robots capable of doing tasks in the real world may arrive in 2027.
Robots already perform constrained real-world tasks. The claim needs a stronger pre-registered bar, such as generality, autonomy and reliability across unfamiliar environments.
Source
AGI with human-level cognitive versatility is five to ten years away.
Hassabis supplied a capability concept and a window, but no public test suite. The ledger will not let a company announcement settle the question on its own.
Source and incentive context
Google DeepMind is pursuing AGI and stands to benefit from confidence in the field’s progress. Hassabis has generally offered a wider window than Amodei. Sources: TIME interview; TIME video.
A 50% chance of “minimal AGI” by 2028.
This updates a forecast Legg made in 2011. Scoring requires the operational definition discussed in the source, not whatever system receives an AGI label in 2028.
Human-level AI would take several years, if not a decade, and would not arrive in the next year or two.
The near-term exclusion can be evaluated first. LeCun also predicted that scaling current language models alone would not produce human-level intelligence, a separate architectural claim.
AI would probably become smarter than any single human by the end of 2025.
No agreed scalar measure of “smarter” exists. Frontier systems exceeded every human on some tests and remained brittle on basic tasks. The deadline passed, but the wording prevents a defensible verdict.
Source
Tesla would have more than one million robotaxis on the road in 2020.
Tesla had no public robotaxi service in 2020. Its September 2026 Texas fleet numbered about 420 registered robotaxis, roughly 99.96% below the promised million and almost six years late.
Evidence
Sources: Autonomy Day transcript; September 2026 fleet report; Tesla 2025 annual filing.
Hospitals should stop training radiologists because deep learning would outperform them within five years.
AI became valuable in image analysis, but radiologists were not made obsolete. Demand remained high, and Hinton later acknowledged that his forecast was wrong.
Evidence
The American College of Radiology still describes a workforce shortage and AI as a tool to reduce workload. Sources: Hinton retrospective; ACR workforce update.
A 10–20% chance that AI causes human extinction within 30 years.
As with the Coxon and Hubinger forecasts, one binary outcome cannot establish individual calibration. The value comes from preserving the probability, horizon and subsequent revisions.
Source
Within five years, AI could pass every test a human takes, including specialized professional exams.
The claim concerns test performance, not general intelligence or reliable professional practice. Resolution should use a frozen test set and human comparison rule.
Source and incentive context
Nvidia’s revenue is closely tied to demand for AI compute, so forecasts of rapid capability growth have clear commercial relevance. Source: Stanford SIEPR.
An AI coding agent with roughly mid-level-engineer capability would become possible in 2025.
Coding agents became capable of substantial bounded work, but “mid-level engineer” bundles judgment, ownership and reliability that benchmark scores do not establish. Meta later shifted part of the timeline into 2026.
Evidence
Sources: Meta Q1 2025 earnings transcript; 2026 coding-agent task study.
AI could make a two-day workweek possible within a decade as machines handle most routine work.
Technical productivity does not automatically shorten paid work. The resolution rule should track the median standard workweek across major economies, not isolated four-day-work pilots or individual choice.
Source
No predictions match those filters.
Download the ledger and inspect calibration
Use the exports for your own analysis. CSV and JSON contain the same 100 records, source URLs and current editorial status. The calibration chart below shows scored claims only; review-queue records remain visible but do not count as hits or misses.
Generated in your browser from the records above.
Each bar is an editorial score from −100 to +100. A dashed track means the forecaster has no independently scored claim yet. Rankings remain withheld until at least five resolvable claims exist.
Machine-readable endpoints: JSON API v1 · CSV API v1 · Revision API v1.
Revision history
Every evidence review gets a dated entry. Records are never silently overwritten: a changed deadline, score or source stays visible in the revision trail.
| Version | Date | Change |
|---|---|---|
| 1.0 | September 14, 2026 | Initial source-first ledger with 20 detailed records. |
| 1.1 | September 14, 2026 | Expanded to 100 source-verified claims, including 80 atomic review-queue records. |
| 1.2 | September 14, 2026 | Added public CSV/JSON downloads and per-forecaster calibration chart. |
| 1.3 | September 14, 2026 | Added stable versioned APIs and this auditable revision history. |
The monthly evidence review will append future versions when a source, status, score, deadline or record count changes. Unchanged reviews are intentionally not represented as fake revisions.
How scoring works
An accuracy score runs from −100 to +100 only after a claim can be evaluated. The score combines outcome, timing, magnitude, scope and mechanism. It is an editorial judgment with an evidence trail, not a scientific probability.
Leaderboards are withheld until a forecaster has at least five independently resolvable claims. Ranking people on one cherry-picked hit or miss would manufacture certainty.
What counts as evidence of influence
People selling AI have an incentive to emphasize capability and adoption. Safety-focused labs can benefit from rules that raise the cost of entry. Investors benefit from growth narratives. Critics gain attention when forecasts fail. These incentives justify scrutiny, not automatic dismissal.
A regulatory-capture claim needs documented lobbying, proposed rules, competitive effects and a plausible causal link. A “psy-op” claim needs evidence of coordinated deception. Similar public statements, even when they appear suddenly, do not prove either one. The ledger will preserve policy positions and financial interests so readers can test those hypotheses against records rather than vibes.
Corrections and submissions
Submit the original source, exact date, full surrounding context, stated deadline and a proposed resolution rule. Corrections should identify the specific field at issue and provide stronger evidence. Archived copies are preferred for posts that can be edited or deleted.
Editorial rule: a forecast can be revised, but its earlier version stays in the history. Quietly moving a deadline is itself useful data.
Core sources
- Associated Press. “New warnings about the risks of AI to humanity revive a long-running debate.” September 14, 2026.
- Dario Amodei. “We Must Pace the Frontier.” September 2026.
- Dario Amodei. “Machines of Loving Grace.” October 2024.
- Anthropic. Recommendations to OSTP for the US AI Action Plan. March 6, 2025.
- Sam Altman. “The Gentle Singularity.” June 10, 2025.
- OpenAI. “How agents are transforming work.” June 25, 2026.
- Google DeepMind. “From AGI to ASI.” June 12, 2026.
- Rodney Brooks. “Predictions Scorecard, 2025 January 01.”
- Grace et al. “Thousands of AI Authors on the Future of AI.” 2024.
Disclosure: The featured image is an AI-generated editorial illustration, not documentary evidence. “Source-verified” means the claim is tied to a public primary record; it does not mean the forecast has come true. Forecast evaluation is an editorial judgment based on documented sources and may be revised as evidence changes. This ledger is a growing dataset, not a claim that every public AI forecast has already been captured.
