Open dataset · refreshed hourly

The AI Citation Index

Everyone argues about what AI search rewards. Almost nobody measures it. This is the source side of the ledger — every domain six answer engines actually cited, taken from 82,567 answers we collected ourselves, published in full and free to reuse.

Citations
356,920
across 37,639 cited answers
Distinct domains
43,954
every host an engine returned
Prompts
7,920
in 578 scans, 74 accounts
Engines
6
April 2026 – September 2026

CC BY 4.0 — republish anything here with a link to rankmonster.ai/sources.

Finding 01

The domains AI answers actually cite

Ranked by breadth — how many separate accounts, in unrelated industries, saw the domain cited in their own scan. Raw citation volume is the wrong sort key: a domain quoted 1,400 times inside three scans of one niche tells you about the niche, not about the engines. By that measure the most universal source in AI answers is reddit.com, which showed up in scans belonging to 71 of 74 accounts.

#DomainSource typeAccounts that saw it citedPromptsCitations
1reddit.comCommunity & forums712,4557,800
2linkedin.comSocial networks692,2258,988
3youtube.comVideo652,0457,412
4en.wikipedia.orgEncyclopedias & reference637871,832
5facebook.comSocial networks566561,630
6medium.comVendor & independent sites486601,645
7forbes.comNews & media46406861
8indeed.comVendor & independent sites46310851
9g2.comReviews & directories455481,756
10gartner.comReviews & directories44363763
11instagram.comSocial networks39256563
12clutch.coReviews & directories384181,096
13gitnux.orgStatistics aggregators35304547
14wifitalents.comStatistics aggregators35247386
15zipdo.coStatistics aggregators34177284
16worldmetrics.orgStatistics aggregators33260451
17finance.yahoo.comNews & media33115225
18trustpilot.comReviews & directories32126463
19sciencedirect.comAcademic & government32162376
20cbinsights.comVendor & independent sites32149346
21prnewswire.comNews & media32137258
22pmc.ncbi.nlm.nih.govAcademic & government31238541
23google.comVendor & independent sites3189300
24f6s.comReviews & directories303281,011
25designrush.comReviews & directories30269587
The prompt sets belong to real customers, so the corpus leans toward the categories they sell into. Breadth ranking is the correction for that, not a cure — read the top of this table as “sources that generalise”, and the citation column as volume inside our sample.
Finding 02

Every engine has a different diet

Share of each engine’s citations by the kind of page it pointed at. The spread is the story: these are not six front-ends over one index, they are six different retrieval philosophies. Perplexity sends 3.5% of its citations to forums and community threads; Claude sends 0.3%.

Source typeChatGPTClaudeGeminiPerplexityAI OverviewsGrokAll engines
Community & forums
Threads real people wrote — Reddit, Quora, Stack Exchange, Hacker News.
3.0%0.3%0.8%3.5%2.8%1.5%2.5%
Video
Video platforms, where the citation points at a watch page or transcript.
0.2%0.1%0.7%2.8%6.7%1.0%2.1%
Social networks
Profile and post URLs on the big social graphs.
0.4%1.1%0.2%4.7%5.0%3.0%3.2%
Reviews & directories
Software review marketplaces, agency directories and ratings sites.
3.1%4.0%1.0%2.8%1.1%1.8%2.6%
Encyclopedias & reference
Wikipedia and the other general-reference encyclopedias.
2.4%0.7%0.3%0.4%0.3%0.7%0.6%
News & media
Publishers and newswires, from national press to trade titles.
0.5%0.8%1.0%0.8%0.6%1.5%0.8%
Academic & government
Journals, preprint servers, .edu, .gov and intergovernmental bodies.
4.8%3.5%2.4%2.8%1.1%3.4%3.0%
Statistics aggregators
Sites whose whole product is scraped statistics round-ups.
0.3%0.3%0.1%0.9%0.1%0.1%0.5%
Vendor & independent sites
Residual — everything the lists above didn’t recognise. Not shaded: it would flatten every other row.
85.2%89.2%93.6%81.3%82.3%86.9%84.7%
0% of the engine’s citations7%+
Statistics farms rival the encyclopedias

A hand-picked list of 1.9k citations to sites whose entire product is scraped statistics round-ups comes to 0.54% of everything — against 0.62% for Wikipedia and every other general reference work combined. And that stats figure is a floor: the bucket only counts the farms we recognised by name. Whatever else it says, it says numbers get cited.

85% is still ordinary websites

The residual bucket — company sites, docs, blogs, personal pages — is the largest single category by a distance. The platforms get the headlines; the long tail of unremarkable pages gets the citations. That is the part you can actually compete for.

Source types are matched against curated domain lists, so every named bucket is a lower bound and “Vendor & independent sites” absorbs everything unrecognised. The lists are in the CSV.
Finding 03

Ranking in one engine barely predicts the next

Give two engines the identical prompt in the identical scan, and compare the sets of domains they cite. The number below is the mean Jaccard overlap — shared domains as a share of all domains either engine used. Gemini and ChatGPT agree on 4.5% of their sources. The best-agreeing pair on the board, Grok and ChatGPT, manages 15.1%. There is no such thing as “ranking in AI”.

OverlapChatGPTClaudeGeminiPerplexityAI OverviewsGrok
ChatGPT6.1%4.5%6.4%9.6%15.1%
Claude6.1%8.7%13.3%10.0%10.5%
Gemini4.5%8.7%9.1%9.6%11.7%
Perplexity6.4%13.3%9.1%13.1%12.9%
AI Overviews9.6%10.0%9.6%13.1%12.4%
Grok15.1%10.5%11.7%12.9%12.4%
no shared sources15% shared

How often each engine cites anything at all

Before overlap can matter, an engine has to show its sources. AI Overviews attaches sources to 85% of its answers; ChatGPT to 30%.

EngineAnswers with sourcesAnswersCitationsSources per cited answer
AI Overviews85%3,42027,5999.5
Perplexity67%17,314176,21215.3
Grok45%17,27336,9764.7
Claude38%9,53046,64912.8
Gemini37%17,52838,0945.8
ChatGPT30%17,50231,3906.1
Pairs are only reported where both engines cited sources on at least 100 of the same prompts. Engines are queried independently and answer non-deterministically, so some of this gap is run-to-run variance rather than a stable difference in index — we have not yet published a repeat-scan variance study to separate the two.
Finding 04

AI does not cite new content

For every page we saw cited and could fetch, we read the publish date out of its own markup and compared it to the first time we recorded an engine citing it. The median page was 196 days old. Only 9% were under a month old; 69% were under a year. Publish-and-wait is not a strategy, but neither is judging a page after three weeks.

25th percentile
88 days
Median
196 days
75th percentile
471 days
Pages measured
2,052
with a readable publish date
Page age at first observed citationShare of cited pagesPages
Under a week old2.8%57
1 week – 1 month5.9%122
1 – 3 months16.9%346
3 – 12 months43.4%890
1 – 3 years19.2%393
Over 3 years11.9%244
This is an upper bound on time-to-first-citation, not a measurement of it. We discover a citation when a scan happens to run, so the true first citation sits at or before the age recorded here. It also only covers pages that publish a machine-readable date and let us fetch them — 12% of cited pages block our crawler outright and are missing from this cut.
Finding 05

What a cited page is made of

We fetch the pages engines cite and pull their structure apart. This is the composite of 4,627 of them — not what SEO advice says a citable page should look like, but what the pages winning citations right now actually contain.

Median length
1,683 words
Median statistics
15
discrete numbers on the page
Quotable paragraphs
10
self-contained, liftable
Question headings
3
headings phrased as a question
Has an FAQ block
62%
Has an author byline
37%
Has structured data
81%
any schema.org type
Leads with a TL;DR
16%
the least common trait

Format of the cited page

FormatSharePages
guide59.6%3,469
listicle22.3%1,298
landing6.7%388
comparison table5.5%318
news2.4%142
review1.8%106
docs0.9%53
video0.8%48
Structure is measured, causation is not. These pages have these traits; we have not run the controlled test that would prove the traits earned the citation. Treat this as the shape of the winning population, not a checklist with a guarantee attached.
Method

How this is measured, and what it can’t tell you

The sample

82,567 answers, from 7,920 distinct prompts run across six engines in 578 scans belonging to 74 accounts, between April 2026 – September 2026. Every source URL an engine returned with its answer is recorded; domains keep their subdomain, so en.wikipedia.org and blog.hubspot.com are their own rows.

What we exclude

Microsoft Copilot was retired from the product mid-sample and is dropped everywhere — leaving it in would drag every cross-engine number toward a window the other five don’t share. Nothing else is filtered: we do not remove a brand’s own domain from its own results.

The bias we can’t remove

Prompts come from real customer prompt sets, so the corpus over-represents the categories our customers sell into. Ranking domains by how many separate accounts saw them cited suppresses most of that, but a universal claim from a non-random sample is still a claim about this sample.

What is missing

Engine answers are non-deterministic and we have not yet published a repeat-scan variance study, so treat single-point differences between engines with more suspicion than the large ones. Crawler-frequency data — how often GPTBot and ClaudeBot actually fetch a page — is being collected but does not yet cover enough sites to publish honestly.

FAQ

Questions about the data

Which sources do AI answer engines cite most?
Across 356,920 citations collected from six engines between April 2026 – September 2026, the domain seen by the most separate accounts is reddit.com — cited in scans belonging to 71 of 74 accounts. Ranking by breadth rather than raw volume matters: a domain cited thousands of times inside one niche is a property of that niche, not of the engines.
Does ranking in ChatGPT mean you rank in Gemini?
No. On the same prompt in the same scan, Gemini and ChatGPT share only 4.5% of their cited domains — the weakest agreement of any pair we measure. The strongest pair, Grok and ChatGPT, still only reaches 15.1%. Answer-engine visibility has to be measured per engine.
How long after publishing does AI start citing a page?
Longer than most content calendars assume. Of 2,052 cited pages with a readable publish date, the median was 196 days old when we first recorded an engine citing it, and only 9% were under a month old. That figure is an upper bound — we learn a page is cited when a scan runs, not the moment it happens.
What does a page that AI cites look like?
The median cited page runs 1,683 words, carries 15 discrete statistics and 10 self-contained quotable paragraphs, and 62% of them include an FAQ block. Guides and listicles dominate the format mix.
Can I use this data?
Yes. Every table on this page has a CSV download and the data is free to republish with attribution and a link to rankmonster.ai/sources. If you want a cut we do not publish — a specific category, a longer window, a different slice — email [email protected] and we will usually run it.
How is this measured?
We run fixed prompt sets through six answer engines and record every source URL each engine returns with its answer. This edition covers 82,567 engine answers across 7,920 distinct prompts and 578 scans, from 74 accounts. Citations are counted per engine response; domains keep their subdomain.

This is the industry view. Get yours.

The same pipeline that produced this page will tell you which of these domains cite you, which cite your competitors, and which prompts you are missing from — across all six engines.

Run a free scan