Methodology v1.0
Methodology
How the Kingy AI Referral Index selects studies, tiers them, classifies their disagreements, and derives the reconciled ranges. Versioned, with every change logged.
What this index is
Published estimates of how much traffic AI assistants send to websites disagree by more than an order of magnitude. Depending on which study you cite, AI referrals are a rounding error or a structural shift. Both claims are made confidently, and both are usually made without reference to the other.
The Kingy AI Referral Index is a standing reconciliation of those studies. It does three things: it catalogues the primary research, it classifies why the studies disagree using a fixed taxonomy, and it publishes the ranges that survive that reconciliation. It is versioned and archived, so a figure cited from it stays checkable after the number changes.
It is not a new measurement. We do not run a panel or a crawler. The one exception is our own first-party data, which is published as a transparency exhibit and is deliberately kept out of the reconciled figures.
What gets included
A record enters the catalogue when it makes a quantified, public claim about AI-driven referral traffic, click impact, or the growth of either. We catalogue the claim as published, including claims we do not find credible — those are tiered, not omitted. Silent exclusion would let us shape the picture by choosing what to leave out.
For each record we store the source organisation, the study name, a one-line method summary, the sample, the measurement window, the metric type, the headline value, a denominator note, a caveat, and a link to the primary source.
Values are recorded exactly as published
The headline value is stored verbatim: 0.37%, 5–15%, 25×, <1%. We do not round it, convert it, or restate it in units the study did not use. This has three consequences that are visible on the hub:
- A published range stays a range. It is drawn as a span across its full width and is never collapsed to a midpoint. A midpoint asserts a precision the study did not claim.
- A bounded claim stays bounded. “
<1%” is charted as the span from zero to one, not as the point one. - An unquantified claim is not quantified. Where a source published prose rather than a figure, the record is catalogued and appears in the downloadable dataset, but it is excluded from the chart and counted in a note beneath it. The count is shown rather than the record being quietly dropped.
- A comparison is not a value. Pew’s headline finding is that users clicked a search result on 8% of visits where an AI summary appeared, against 15% where none did. That is two conditions, not one number. We store it as published and exclude it from the chart, because charting the 8% alone would state the post-condition as the finding and silently discard the comparison that is the finding. The same applies to figures published as ratios in words, such as “1 in 10”.
These refusals are enforced in code rather than left to editorial care, and every one of them was added after a real published figure was misread during construction. We would rather the chart be visibly incomplete than quietly wrong.
The same rule governs units. Each chart panel is drawn on a single scale, and a record published in a different unit from the rest of its panel — percentage points against percentages, say, or a percentage against a multiple — is excluded from that panel and included in the same visible count. Plotting them together would manufacture a comparison the sources never supported.
Tiers
Tier A is a study whose method, sample, denominator and measurement window can each be established from a primary source. Only Tier A records feed the reconciled ranges and the chart’s solid series.
Tier B is quarantined. A record lands here when one of those four cannot be established — most often a vendor-published figure with no disclosed methodology, or a forecast presented in the same register as a measurement. Tier B records are catalogued in full, rendered hatched and grouped separately, and excluded from every reconciled figure.
Quarantine is not an accusation. It is not a claim that a Tier B figure is wrong, and several will likely turn out to be roughly right. It is a statement that we cannot check it, and therefore will not lean on it.
Separately from tier, each record carries a verification status. Verified against primary source means we have read the source document itself. Pending primary verification means the figure currently rests on secondary reporting; those records are marked wherever they appear, including on the chart.
Method buckets
Every record is assigned to one of five buckets describing how the measurement was taken. The bucket, not the headline number, is usually what explains a disagreement.
- 1 · Behavioral — opt-in panels of real users, with behaviour observed directly.
- 2 · Clickstream — passively collected browsing streams, typically via extensions or SDKs.
- 3 · Site-side — a publisher’s own server-side or analytics data for its own properties.
- 4 · Network / CDN — request logs from infrastructure sitting in front of many sites.
- 5 · Self-report / forecast — figures a party published about itself, and projections.
How the reconciled ranges are derived
Reconciliation is done per category — publishers and news, B2B and tech, retail and commerce — because the underlying rates genuinely differ by category, and much of the apparent disagreement between studies is really a disagreement about which sites they looked at.
The procedure is:
- Restrict to Tier A. Tier B and pending-verification records are set aside entirely.
- Group by category and by bucket. Where a category’s Tier A records cluster inside one bucket, we say so — a range derived from a single measurement method is a weaker claim than one that survives across methods.
- Account for the Seven Deltas. Each delta is a recurring, identifiable difference in what was measured. Where a delta can be adjusted for from published detail, we adjust and record the adjustment. Where it cannot, we widen the range rather than pick a point.
- Publish a range, with its n. Every card states how many Tier A studies stand behind it. We do not publish a central estimate, because the honest output of reconciling disagreeing measurements is an interval.
- Say when there is not enough evidence. Where a category has too few Tier A records to reconcile meaningfully, the card says so instead of showing a narrow range that would not survive one more study.
The reconciled ranges are a judgement built on published evidence, not a computation that would produce the same answer automatically. That judgement is what the tiering, the bucketing and the delta taxonomy exist to make inspectable — and the full dataset is downloadable so a reader can disagree with our reconciliation using our own records.
Corrections
Corrections are logged, not applied silently. If a record is wrong, we fix the record and append a dated entry to the changelog below. If the method itself changes, the version number increments and the change is described. Archived waves are never edited after publication — a citation that pointed at a figure keeps pointing at that figure.
If you believe a record misstates your study, tell us and we will check it against the primary source.
Cadence
Waves are quarterly. At each wave the current hub is archived at a permanent URL, the wave stamp increments, and any method changes are appended below. Between waves we publish shorter reaction pieces when a significant new study lands; those link back here rather than restating the index.
Licence
The index data is published under CC BY 4.0. You may republish, chart and build on these figures, including commercially, provided you credit Kingy AI and link to the index. Attribution and the link are both required.
Changelog
This methodology is versioned. Changes are appended here with the date they took effect; earlier waves are never retroactively altered.
No changes yet — v1.0 is the first published version.