AI News

The AI Prediction Ledger: Who Was Right, Who Was Wrong, and What Comes Next

Kingy.ai accountability project

AI predictions, scored against reality

AI leaders make claims that move markets, shape regulation and influence how people plan their careers. This ledger preserves those claims, separates forecasts from warnings, and checks the measurable ones when their deadlines arrive.

Version 1.3 · 100 source-verified claims · Last evidence review: September 14, 2026

100Claims tracked
6Scored claims
2Partly supported
3False
1Unresolvable
80New review queue
The current alarm is now on the record. Jacob Coxon assigned a 10% probability to AI-caused human extinction within a decade after leaving Anthropic. Dario Amodei then called for slower frontier capability gains, and Sam Altman publicly agreed. Those are consequential statements. They are not evidence, by themselves, of a coordinated “psy-op” or regulatory capture. The ledger records commercial and policy incentives, but it does not infer hidden motives without documentation.

The first lesson is already clear: the loudest forecasts are often the hardest to score. “AGI,” “smarter than humans,” “most code” and “join the workforce” can each shift meaning after the deadline. A useful prediction states the outcome, scope, date and measurement rule before reality arrives. For the operational evidence behind the agent claims, see Kingy’s State of AI Agents 2026 and the Kingy Score v2 methodology.

The ledger



Jacob Coxon
Former pretraining researcher, Anthropic and OpenAI · September 8, 2026
Pending

A 10% chance that AI causes human extinction within the next decade.

PendingRiskDeadline: September 2036Probability supplied

This is a probabilistic forecast, not a claim that extinction will occur. It cannot be marked wrong merely because the outcome does not happen; calibration requires a reference class of comparable forecasts.

Source and incentive context

Coxon made the estimate while publicly resigning and said he left before his Anthropic equity vested. That reduces one obvious financial incentive to promote Anthropic, but does not validate the forecast. Sources: Associated Press; Axios interview.

Evan Hubinger
Alignment science lead, Anthropic · September 2026
Pending

More than a 10% chance that AI kills all humans within the next decade.

PendingRiskDeadline: September 2036Probability supplied

The estimate is unusually explicit but still lacks a published model showing how the number was derived. Keep it as a forecast, not a fact about current systems.

Source and incentive context

Hubinger works for a frontier lab whose public case for safety rules can affect regulation and competition. His technical work on deceptive model behavior is relevant evidence, but employment alone proves neither sincerity nor strategy. Sources: Washington Post; Sleeper Agents paper.

Dario Amodei
CEO, Anthropic · October 2024 / March 2025
Pending

“Powerful AI” matching or exceeding Nobel-level experts across most disciplines could arrive in late 2026 or early 2027.

PendingCapabilitiesDeadline: early 2027

The definition is demanding and more useful than a bare AGI label. A fair resolution still needs agreed tests for expert-level performance, autonomous work and reliability across disciplines.

Source and incentive context

Anthropic sells access to frontier models and also argues for stronger frontier-model oversight. Both interests belong in the record. Sources: Machines of Loving Grace; Anthropic submission summary.

Dario Amodei
CEO, Anthropic · May 28, 2025
Pending

AI could eliminate half of entry-level white-collar jobs within one to five years.

PendingJobsDeadline: May 2030

This needs a fixed occupation list and baseline job count. Job openings, employment levels and task automation are different measures and should not be swapped after the deadline.

Source and incentive context

The forecast came from a direct interview. Anthropic benefits when employers believe its systems can perform valuable work, while stronger displacement expectations may also support regulation and workforce programs. Source: Axios.

Dario Amodei
CEO, Anthropic · May 28, 2025
Pending

AI-driven displacement could push US unemployment to 10–20% within one to five years.

PendingJobsDeadline: May 2030

This is measurable, but causation matters. A recession that raises unemployment without demonstrated AI displacement would not satisfy the forecast.

Source and incentive context

Record alongside the entry-level-jobs claim, but score separately because the labor-force consequence can fail even if task automation rises. Source: Axios.

Dario Amodei
CEO, Anthropic · March 10, 2025
−65

AI would write 90% of code within three to six months and nearly all code within a year.

FalseCodingDeadline passedScope error

Coding-agent use rose quickly, but no credible industry-wide measurement approached 90% by September 2025. Amodei later said it happened “at least at some places,” which narrows the scope after the fact.

Evidence

Open-source research found more than 320,000 agent-attributed commits per month by April 2026, substantial but not evidence that AI wrote 90% of all code. Sources: claim and follow-up review; repository census.

Dario Amodei
CEO, Anthropic · October 2024, reiterated September 2026
Pending

Powerful AI could compress 50–100 years of biological progress into 5–10 years and help cure or prevent most major diseases.

PendingScienceConditional horizon

The clock begins only after Amodei’s “powerful AI” threshold, so the medical claim cannot honestly be assigned a fixed deadline until that trigger is resolved.

Source

Machines of Loving Grace; We Must Pace the Frontier.

Sam Altman
CEO, OpenAI · January 6, 2025
+45

The first AI agents could “join the workforce” in 2025 and materially change company output.

Partly supportedJobsDeadline passed

Working agents arrived and enterprise adoption followed. The undefined words “join” and “materially” prevent a clean win, and OpenAI’s strongest quantitative evidence comes from its own users and staff.

Evidence

By mid-2026 OpenAI reported that every internal department used Codex as its primary AI work tool, while external adoption remained uneven. Sources: original prediction; OpenAI usage study; research paper.

Sam Altman
CEO, OpenAI · June 10, 2025
Pending

Systems that can produce novel insights would likely arrive in 2026.

PendingScienceDeadline: December 2026

Resolution requires a disclosed, independently validated discovery, not a model generating a plausible new sentence or helping a human researcher work faster.

Source

The Gentle Singularity.

Sam Altman
CEO, OpenAI · June 10, 2025
Pending

Robots capable of doing tasks in the real world may arrive in 2027.

PendingRoboticsDeadline: December 2027

Robots already perform constrained real-world tasks. The claim needs a stronger pre-registered bar, such as generality, autonomy and reliability across unfamiliar environments.

Source

The Gentle Singularity.

Demis Hassabis
CEO, Google DeepMind · April 2025
Pending

AGI with human-level cognitive versatility is five to ten years away.

PendingCapabilitiesWindow: 2030–2035

Hassabis supplied a capability concept and a window, but no public test suite. The ledger will not let a company announcement settle the question on its own.

Source and incentive context

Google DeepMind is pursuing AGI and stands to benefit from confidence in the field’s progress. Hassabis has generally offered a wider window than Amodei. Sources: TIME interview; TIME video.

Shane Legg
Co-founder and chief AGI scientist, Google DeepMind · 2023
Pending

A 50% chance of “minimal AGI” by 2028.

PendingCapabilitiesDeadline: 2028Probability supplied

This updates a forecast Legg made in 2011. Scoring requires the operational definition discussed in the source, not whatever system receives an AGI label in 2028.

Source

Dwarkesh Podcast transcript; 2011 forecast record.

Yann LeCun
Then chief AI scientist, Meta · October 2024
Pending

Human-level AI would take several years, if not a decade, and would not arrive in the next year or two.

PendingCapabilitiesWindow: roughly 2027–2034

The near-term exclusion can be evaluated first. LeCun also predicted that scaling current language models alone would not produce human-level intelligence, a separate architectural claim.

Source

TechCrunch interview coverage; Lex Fridman transcript.

Elon Musk
CEO, Tesla and xAI · April 2024
No score

AI would probably become smarter than any single human by the end of 2025.

UnresolvableCapabilitiesDeadline passed

No agreed scalar measure of “smarter” exists. Frontier systems exceeded every human on some tests and remained brittle on basic tasks. The deadline passed, but the wording prevents a defensible verdict.

Source

Axios.

Elon Musk
CEO, Tesla · April 22, 2019
−100

Tesla would have more than one million robotaxis on the road in 2020.

FalseRoboticsDeadline passedMagnitude and timing error

Tesla had no public robotaxi service in 2020. Its September 2026 Texas fleet numbered about 420 registered robotaxis, roughly 99.96% below the promised million and almost six years late.

Evidence

Sources: Autonomy Day transcript; September 2026 fleet report; Tesla 2025 annual filing.

Geoffrey Hinton
AI researcher · 2016
−70

Hospitals should stop training radiologists because deep learning would outperform them within five years.

FalseJobsDeadline: 2021Mechanism error

AI became valuable in image analysis, but radiologists were not made obsolete. Demand remained high, and Hinton later acknowledged that his forecast was wrong.

Evidence

The American College of Radiology still describes a workforce shortage and AI as a tool to reduce workload. Sources: Hinton retrospective; ACR workforce update.

Geoffrey Hinton
AI researcher · December 2024
Pending

A 10–20% chance that AI causes human extinction within 30 years.

PendingRiskDeadline: 2054Probability supplied

As with the Coxon and Hubinger forecasts, one binary outcome cannot establish individual calibration. The value comes from preserving the probability, horizon and subsequent revisions.

Source

BBC interview coverage.

Jensen Huang
CEO, Nvidia · March 2024
Pending

Within five years, AI could pass every test a human takes, including specialized professional exams.

PendingCapabilitiesDeadline: March 2029

The claim concerns test performance, not general intelligence or reliable professional practice. Resolution should use a frozen test set and human comparison rule.

Source and incentive context

Nvidia’s revenue is closely tied to demand for AI compute, so forecasts of rapid capability growth have clear commercial relevance. Source: Stanford SIEPR.

Mark Zuckerberg
CEO, Meta · January / April 2025
+35

An AI coding agent with roughly mid-level-engineer capability would become possible in 2025.

Partly supportedCodingDeadline passed

Coding agents became capable of substantial bounded work, but “mid-level engineer” bundles judgment, ownership and reliability that benchmark scores do not establish. Meta later shifted part of the timeline into 2026.

Evidence

Sources: Meta Q1 2025 earnings transcript; 2026 coding-agent task study.

Bill Gates
Microsoft co-founder · March 2025
Pending

AI could make a two-day workweek possible within a decade as machines handle most routine work.

PendingJobsDeadline: 2035

Technical productivity does not automatically shorten paid work. The resolution rule should track the median standard workweek across major economies, not isolated four-day-work pilots or individual choice.

Source

Fortune interview coverage.

No predictions match those filters.

Download the ledger and inspect calibration

Use the exports for your own analysis. CSV and JSON contain the same 100 records, source URLs and current editorial status. The calibration chart below shows scored claims only; review-queue records remain visible but do not count as hits or misses.

Public data:


Generated in your browser from the records above.
Positive scorePartly supportedNegative score

Each bar is an editorial score from −100 to +100. A dashed track means the forecaster has no independently scored claim yet. Rankings remain withheld until at least five resolvable claims exist.

Machine-readable endpoints: JSON API v1 · CSV API v1 · Revision API v1.

Revision history

Every evidence review gets a dated entry. Records are never silently overwritten: a changed deadline, score or source stays visible in the revision trail.

Version Date Change
1.0 September 14, 2026 Initial source-first ledger with 20 detailed records.
1.1 September 14, 2026 Expanded to 100 source-verified claims, including 80 atomic review-queue records.
1.2 September 14, 2026 Added public CSV/JSON downloads and per-forecaster calibration chart.
1.3 September 14, 2026 Added stable versioned APIs and this auditable revision history.

The monthly evidence review will append future versions when a source, status, score, deadline or record count changes. Unchanged reviews are intentionally not represented as fake revisions.

How scoring works

An accuracy score runs from −100 to +100 only after a claim can be evaluated. The score combines outcome, timing, magnitude, scope and mechanism. It is an editorial judgment with an evidence trail, not a scientific probability.

OutcomeDid the specified event happen?
TimingDid it happen inside the stated window?
MagnitudeWas the size of the effect close?
ScopeDid it apply to the promised population or only a narrow example?

Leaderboards are withheld until a forecaster has at least five independently resolvable claims. Ranking people on one cherry-picked hit or miss would manufacture certainty.

What counts as evidence of influence

People selling AI have an incentive to emphasize capability and adoption. Safety-focused labs can benefit from rules that raise the cost of entry. Investors benefit from growth narratives. Critics gain attention when forecasts fail. These incentives justify scrutiny, not automatic dismissal.

A regulatory-capture claim needs documented lobbying, proposed rules, competitive effects and a plausible causal link. A “psy-op” claim needs evidence of coordinated deception. Similar public statements, even when they appear suddenly, do not prove either one. The ledger will preserve policy positions and financial interests so readers can test those hypotheses against records rather than vibes.

Corrections and submissions

Submit the original source, exact date, full surrounding context, stated deadline and a proposed resolution rule. Corrections should identify the specific field at issue and provide stronger evidence. Archived copies are preferred for posts that can be edited or deleted.

Editorial rule: a forecast can be revised, but its earlier version stays in the history. Quietly moving a deadline is itself useful data.

Core sources

  1. Associated Press. “New warnings about the risks of AI to humanity revive a long-running debate.” September 14, 2026.
  2. Dario Amodei. “We Must Pace the Frontier.” September 2026.
  3. Dario Amodei. “Machines of Loving Grace.” October 2024.
  4. Anthropic. Recommendations to OSTP for the US AI Action Plan. March 6, 2025.
  5. Sam Altman. “The Gentle Singularity.” June 10, 2025.
  6. OpenAI. “How agents are transforming work.” June 25, 2026.
  7. Google DeepMind. “From AGI to ASI.” June 12, 2026.
  8. Rodney Brooks. “Predictions Scorecard, 2025 January 01.”
  9. Grace et al. “Thousands of AI Authors on the Future of AI.” 2024.

Disclosure: The featured image is an AI-generated editorial illustration, not documentary evidence. “Source-verified” means the claim is tied to a public primary record; it does not mean the forecast has come true. Forecast evaluation is an editorial judgment based on documented sources and may be revised as evidence changes. This ledger is a growing dataset, not a claim that every public AI forecast has already been captured.