AI News

Most broken AI promises were never promises

We set out to build a ledger of broken AI promises. The first three we investigated — the most notorious ones we could think of — were not broken promises at all. Two were never promised. One was delivered on time and simply did not work.

The answer in one sentence: The AI industry is widely believed to over-promise and under-deliver, but when you demand that a promise be a specific thing, by a stated time, from the vendor’s own mouth, most of the famous failures dissolve — what remains is a much smaller, much more interesting list, and the vendors on it are not always the ones you would guess.

What we were trying to build

The plan was ordinary: a public ledger of things AI companies said they would do, set against what actually happened, with an archived copy of every original claim so readers could check our reasoning rather than take it on trust. Sortable, filterable, one row per promise.

We applied one test to every candidate row. A promise had to be:

  1. A specific thing — a named product, model, or capability.
  2. By a stated time — a date, a quarter, or a phrase like “in about six months” that can be checked against a calendar.
  3. In the vendor’s own words — a statement we can quote and archive, not a characterisation in someone’s coverage.

That test turned out to be the whole story.

The three that dissolved

Kimi K3 was never promised under a Modified MIT licence

The widely-reported version: Moonshot AI said K3 would ship under the same Modified MIT licence as K2, then shipped a restrictive bespoke licence instead. We went looking for the statement. There isn’t one. The K3 repository carries a bespoke licence from its initial commit, with no licence-replacement commit in its history, and Moonshot’s own GitHub README says both code and weights are released under the Kimi K3 License. No Moonshot statement promising Modified MIT exists.

What happened is that outlets extrapolated from the previous generation. Including us. Our own K3 deep dive asserted Modified MIT four times and framed it as a vendor commitment — “committed to”, “the promised Modified MIT terms”. It now carries a dated correction. The phrase that made it look like a broken promise was written on this site, not by Moonshot.

Meta never committed to releasing Llama 4 Behemoth

Behemoth is the model everyone remembers as promised and undelivered. Meta announced it in April 2025, it never shipped, and by 2026 the company’s frontier work had moved to closed weights. The Hugging Face organisation still lists only Scout and Maverick.

But Meta’s own announcement says: “While we’re not yet releasing Llama 4 Behemoth as it is still training, we’re excited to share more technical details about our approach.” That is a status report. There is no release date, and no commitment to release at all. The expectation was real; the promise was not.

Rabbit shipped its Large Action Model on the date it named

This is the sharpest case. The r1 is the emblem of AI over-promising: a device that was going to “replace apps and control all services with a single sentence”, and then did not.

That CES 2024 pitch carried no date, so it fails the test. The one dated commitment Rabbit made — that the web-based Large Action Model agent would arrive on 1 October 2024 — was met. LAM Playground shipped that month. What failed was quality: it struggled with CAPTCHAs and unintended behaviour, and by September 2024 roughly 5,000 of about 100,000 buyers were using the device at any given moment.

Recording that as a broken promise would mean scoring the vendor on whether the product was good, which is a different question and one a promise ledger is not built to answer.

What survived

Four rows cleared the test. This is the evidence table — every claim is dated and quoted, and every outcome was checked against a primary surface on 8 August 2026.

Vendor The dated promise What happened Verdict Evidence
xAI
Grok 3
“Grok 3 will be made open source in about 6 months.”
24 Aug 2025
Due around February 2026. Musk re-confirmed it that month, answering “Yes” to a direct question on X. As of 8 August 2026 the xai-org organisation still publishes Grok-1 and Grok-2 only — no Grok 3 weights. Broken Third-party evidence for the claim; Official source for the outcome
Apple
Personalized Siri
“It’s going to take us longer than we thought … we anticipate rolling them out in the coming year.”
7 Mar 2025
Still unshipped. Apple’s own Apple Intelligence page presents the features as forthcoming: “Siri AI coming in English later this year.” Other Apple Intelligence features did ship; it is this specific commitment that is outstanding. Partial Company claim via archived report; Official source for the outcome
OpenAI
gpt-oss
“… release a powerful new open-weight language model with reasoning in the coming months.”
31 Mar 2025
Delivered 5–8 August 2025 under Apache 2.0, weights downloadable. Slipped publicly twice first: June, then “later this summer”, then a July postponement for further safety testing. Kept Company claim, archived in the claimant’s own words; Official source for the outcome
xAI
Grok-1
Grok would be open sourced “this week”.
11 Mar 2024
Delivered 17 March 2024 under Apache 2.0. Still available. Kept Third-party evidence for the claim; Official source for the outcome

Verdicts are editorial judgements on delivery against a stated commitment. They are not quality assessments, and they are not weighted by how much each promise mattered.

Two things about the shape of that table are worth saying out loud.

First, after all of that, exactly one outright broken promise survives. We began from the most notorious failures we could recall, applied a test that any of them should have passed, and finished with a single clean miss, one partial, and two kept. That is not a defence of the industry — it is a statement about the difference between what is remembered and what is on the record.

It is worth dwelling on the one that survived, because it is stronger than a missed deadline. Musk named a model and a six-month window in August 2025. When that window expired in February 2026 he was asked directly whether Grok 3 would be open-sourced, and answered “Yes”. Six months after that, the weights still do not exist. The commitment was made, restated at its due date, and remains unmet — which is precisely the shape a promise ledger is built to record, and precisely the shape almost nothing else in this research turned out to have.

Second, the same vendor appears as both the one broken promise and one of the two kept ones. xAI makes dated public commitments more often than most companies, which means it both delivers against them and misses them in public. A ledger like this therefore rewards vagueness: a company that never names a date can never be recorded as late, and the quietest vendor scores best by saying nothing. That is a real limitation of the format, and we would rather state it than let the table imply otherwise.

Why so few

Vendor communication comes in two shapes, and neither can be adjudicated.

The first is announcement as delivery. The thing ships the day it is described. There is no interval during which a promise is outstanding, so there is nothing to check later. Most model launches work this way.

The second is directional sentiment. “We believe in open models.” “We were on the wrong side of history.” These are real positions and they shape expectations, but they name no deliverable and no date, so no outcome can falsify them.

The ledgerable shape — a specific thing, by a stated time — is the exception, not the rule. That is why Musk’s six-month Grok 3 commitment stands out: it named a model and a timeframe, which almost nothing else does.

Where we were wrong

Two corrections, both ours

We have already described the first: this site published the Kimi K3 Modified MIT claim and had to correct it. It is one of the three counter-examples in this piece.

The second happened during this research. We initially rejected OpenAI from the ledger, reasoning that its open-weight announcement and delivery were simultaneous. That was wrong, and it was a research failure rather than a judgement call — we had anchored on a February 2025 remark about being “on the wrong side of history” and never found the actual commitment, posted on 31 March 2025, four months before the model shipped. A second search found it. A negative finding needs the same rigour as a positive one, and ours did not have it.

What this does not show

Four rows is not a dataset. This is a pattern observed while trying to build something else, and it should be read that way.

  • The candidates were not randomly sampled. We started from the failures we could think of, which is exactly the selection most likely to be distorted by memory.
  • All four claims are now archived, but only one is archived in the claimant’s own words. The OpenAI row links a snapshot of Altman’s post that we fetched and confirmed contains the quoted phrase. The other three link archived reports carrying the vendor’s verbatim words.
  • Posts on X archive badly, and that is a structural problem for this kind of work. The Internet Archive holds a capture of Musk’s Grok 3 post that preserved only the page shell and none of the text, and holds no capture at all of his March 2024 post. Where a claim lives on X, an archived news report quoting it is often the more durable artefact than the post itself — an uncomfortable inversion of the usual hierarchy.
  • Apple never published its statement. It was issued to reporters, so there is no Apple-hosted original to archive; only the outcome sits on Apple’s own pages. A great deal of what the industry “said” exists only in someone else’s quotation marks.
  • “No dated promise exists” is not a defence of anyone. Meta’s Behemoth expectation was created by Meta’s own announcement, even though no date was attached. The finding is about what can be adjudicated, not about who behaved well.

How to check us

  1. Every promise above is quoted with its date. Search the quote.
  2. Every outcome was verified on a primary surface — the vendor’s own model repository or product page — not on reporting about it.
  3. Where the claim and the outcome come from different evidence classes, the table says so.
  4. If you find a dated, quotable commitment we missed, we want it. That is the hard part of this work, and the list above is certainly incomplete.

We are continuing to collect rows against the same three-part test. If the ledger grows enough to stand on its own, it will be published as one. For now the honest result is the pattern: the gap between “everyone knows they promised it” and “they said it, on this date” is wide, and it is where nearly all of the story turned out to be.