Open dataset · refreshed hourly

The AI Citation Index

Everyone argues about what AI search rewards. Almost nobody measures it. This is the source side of the ledger — every domain six answer engines actually cited, taken from 74,504 answers we collected ourselves, published in full and free to reuse.

Citations
306,916
across 33,551 cited answers
Distinct domains
40,977
every host an engine returned
Prompts
7,573
in 544 scans, 65 accounts
Engines
6
April 2026 – August 2026

CC BY 4.0 — republish anything here with a link to rankmonster.ai/sources.

Finding 01

The domains AI answers actually cite

Ranked by breadth — how many separate accounts, in unrelated industries, saw the domain cited in their own scan. Raw citation volume is the wrong sort key: a domain quoted 1,400 times inside three scans of one niche tells you about the niche, not about the engines. By that measure the most universal source in AI answers is reddit.com, which showed up in scans belonging to 63 of 65 accounts.

#DomainSource typeAccounts that saw it citedPromptsCitations
1reddit.comCommunity & forums632,2826,709
2linkedin.comSocial networks602,0997,300
3youtube.comVideo591,9887,177
4en.wikipedia.orgEncyclopedias & reference577301,653
5facebook.comSocial networks496281,592
6medium.comVendor & independent sites435971,272
7forbes.comNews & media43387780
8indeed.comVendor & independent sites41282730
9gartner.comReviews & directories40344649
10g2.comReviews & directories395111,559
11instagram.comSocial networks37252557
12gitnux.orgStatistics aggregators33299494
13sciencedirect.comAcademic & government33160354
14worldmetrics.orgStatistics aggregators32255424
15finance.yahoo.comNews & media32109206
16clutch.coReviews & directories303841,003
17pmc.ncbi.nlm.nih.govAcademic & government30239512
18wifitalents.comStatistics aggregators30238351
19zipdo.coStatistics aggregators30171254
20f6s.comReviews & directories29319890
21trustpilot.comReviews & directories29117440
22blog.hubspot.comVendor & independent sites29193366
23salesforce.comVendor & independent sites28239491
24cbinsights.comVendor & independent sites28131299
25scribd.comVendor & independent sites2879128
The prompt sets belong to real customers, so the corpus leans toward the categories they sell into. Breadth ranking is the correction for that, not a cure — read the top of this table as “sources that generalise”, and the citation column as volume inside our sample.
Finding 02

Every engine has a different diet

Share of each engine’s citations by the kind of page it pointed at. The spread is the story: these are not six front-ends over one index, they are six different retrieval philosophies. Perplexity sends 3.6% of its citations to forums and community threads; Claude sends 0.3%.

Source typeChatGPTClaudeGeminiPerplexityAI OverviewsGrokAll engines
Community & forums
Threads real people wrote — Reddit, Quora, Stack Exchange, Hacker News.
2.8%0.3%0.8%3.6%2.7%1.5%2.5%
Video
Video platforms, where the citation points at a watch page or transcript.
0.2%0.1%0.8%3.4%7.0%1.0%2.4%
Social networks
Profile and post URLs on the big social graphs.
0.4%1.1%0.1%4.7%4.9%3.0%3.2%
Reviews & directories
Software review marketplaces, agency directories and ratings sites.
3.3%3.9%1.1%2.9%1.1%1.8%2.6%
Encyclopedias & reference
Wikipedia and the other general-reference encyclopedias.
2.6%0.6%0.3%0.4%0.3%0.7%0.7%
News & media
Publishers and newswires, from national press to trade titles.
0.6%0.9%1.0%0.8%0.6%1.5%0.9%
Academic & government
Journals, preprint servers, .edu, .gov and intergovernmental bodies.
5.1%3.2%2.5%3.1%1.0%3.4%3.1%
Statistics aggregators
Sites whose whole product is scraped statistics round-ups.
0.3%0.3%0.1%1.0%0.1%0.1%0.6%
Vendor & independent sites
Residual — everything the lists above didn’t recognise. Not shaded: it would flatten every other row.
84.7%89.6%93.2%80.0%82.3%86.9%84.1%
0% of the engine’s citations7%+
Statistics farms rival the encyclopedias

A hand-picked list of 1.7k citations to sites whose entire product is scraped statistics round-ups comes to 0.56% of everything — against 0.65% for Wikipedia and every other general reference work combined. And that stats figure is a floor: the bucket only counts the farms we recognised by name. Whatever else it says, it says numbers get cited.

84% is still ordinary websites

The residual bucket — company sites, docs, blogs, personal pages — is the largest single category by a distance. The platforms get the headlines; the long tail of unremarkable pages gets the citations. That is the part you can actually compete for.

Source types are matched against curated domain lists, so every named bucket is a lower bound and “Vendor & independent sites” absorbs everything unrecognised. The lists are in the CSV.
Finding 03

Ranking in one engine barely predicts the next

Give two engines the identical prompt in the identical scan, and compare the sets of domains they cite. The number below is the mean Jaccard overlap — shared domains as a share of all domains either engine used. Gemini and ChatGPT agree on 4.8% of their sources. The best-agreeing pair on the board, Grok and ChatGPT, manages 15.1%. There is no such thing as “ranking in AI”.

OverlapChatGPTClaudeGeminiPerplexityAI OverviewsGrok
ChatGPT6.0%4.8%6.6%9.3%15.1%
Claude6.0%8.8%13.5%10.0%10.5%
Gemini4.8%8.8%9.3%9.3%11.7%
Perplexity6.6%13.5%9.3%13.2%12.9%
AI Overviews9.3%10.0%9.3%13.2%12.4%
Grok15.1%10.5%11.7%12.9%12.4%
no shared sources15% shared

How often each engine cites anything at all

Before overlap can matter, an engine has to show its sources. AI Overviews attaches sources to 85% of its answers; ChatGPT to 29%.

EngineAnswers with sourcesAnswersCitationsSources per cited answer
AI Overviews85%2,76823,81510.1
Perplexity64%15,627145,19314.6
Grok50%15,58436,9764.7
Claude36%8,87340,45612.6
Gemini36%15,83932,9215.8
ChatGPT29%15,81327,5556.0
Pairs are only reported where both engines cited sources on at least 100 of the same prompts. Engines are queried independently and answer non-deterministically, so some of this gap is run-to-run variance rather than a stable difference in index — we have not yet published a repeat-scan variance study to separate the two.
Finding 04

AI does not cite new content

For every page we saw cited and could fetch, we read the publish date out of its own markup and compared it to the first time we recorded an engine citing it. The median page was 203 days old. Only 9% were under a month old; 67% were under a year. Publish-and-wait is not a strategy, but neither is judging a page after three weeks.

25th percentile
90 days
Median
203 days
75th percentile
510 days
Pages measured
1,767
with a readable publish date
Page age at first observed citationShare of cited pagesPages
Under a week old2.8%49
1 week – 1 month6.0%106
1 – 3 months16.4%289
3 – 12 months42.2%746
1 – 3 years19.5%344
Over 3 years13.2%233
This is an upper bound on time-to-first-citation, not a measurement of it. We discover a citation when a scan happens to run, so the true first citation sits at or before the age recorded here. It also only covers pages that publish a machine-readable date and let us fetch them — 12% of cited pages block our crawler outright and are missing from this cut.
Finding 05

What a cited page is made of

We fetch the pages engines cite and pull their structure apart. This is the composite of 3,974 of them — not what SEO advice says a citable page should look like, but what the pages winning citations right now actually contain.

Median length
1,681 words
Median statistics
15
discrete numbers on the page
Quotable paragraphs
10
self-contained, liftable
Question headings
3
headings phrased as a question
Has an FAQ block
61%
Has an author byline
38%
Has structured data
80%
any schema.org type
Leads with a TL;DR
16%
the least common trait

Format of the cited page

FormatSharePages
guide60.3%3,023
listicle21.9%1,098
landing6.9%345
comparison table5.1%258
news2.3%117
review1.6%78
docs0.9%47
video0.9%46
Structure is measured, causation is not. These pages have these traits; we have not run the controlled test that would prove the traits earned the citation. Treat this as the shape of the winning population, not a checklist with a guarantee attached.
Method

How this is measured, and what it can’t tell you

The sample

74,504 answers, from 7,573 distinct prompts run across six engines in 544 scans belonging to 65 accounts, between April 2026 – August 2026. Every source URL an engine returned with its answer is recorded; domains keep their subdomain, so en.wikipedia.org and blog.hubspot.com are their own rows.

What we exclude

Microsoft Copilot was retired from the product mid-sample and is dropped everywhere — leaving it in would drag every cross-engine number toward a window the other five don’t share. Nothing else is filtered: we do not remove a brand’s own domain from its own results.

The bias we can’t remove

Prompts come from real customer prompt sets, so the corpus over-represents the categories our customers sell into. Ranking domains by how many separate accounts saw them cited suppresses most of that, but a universal claim from a non-random sample is still a claim about this sample.

What is missing

Engine answers are non-deterministic and we have not yet published a repeat-scan variance study, so treat single-point differences between engines with more suspicion than the large ones. Crawler-frequency data — how often GPTBot and ClaudeBot actually fetch a page — is being collected but does not yet cover enough sites to publish honestly.

FAQ

Questions about the data

Which sources do AI answer engines cite most?
Across 306,916 citations collected from six engines between April 2026 – August 2026, the domain seen by the most separate accounts is reddit.com — cited in scans belonging to 63 of 65 accounts. Ranking by breadth rather than raw volume matters: a domain cited thousands of times inside one niche is a property of that niche, not of the engines.
Does ranking in ChatGPT mean you rank in Gemini?
No. On the same prompt in the same scan, Gemini and ChatGPT share only 4.8% of their cited domains — the weakest agreement of any pair we measure. The strongest pair, Grok and ChatGPT, still only reaches 15.1%. Answer-engine visibility has to be measured per engine.
How long after publishing does AI start citing a page?
Longer than most content calendars assume. Of 1,767 cited pages with a readable publish date, the median was 203 days old when we first recorded an engine citing it, and only 9% were under a month old. That figure is an upper bound — we learn a page is cited when a scan runs, not the moment it happens.
What does a page that AI cites look like?
The median cited page runs 1,681 words, carries 15 discrete statistics and 10 self-contained quotable paragraphs, and 61% of them include an FAQ block. Guides and listicles dominate the format mix.
Can I use this data?
Yes. Every table on this page has a CSV download and the data is free to republish with attribution and a link to rankmonster.ai/sources. If you want a cut we do not publish — a specific category, a longer window, a different slice — email [email protected] and we will usually run it.
How is this measured?
We run fixed prompt sets through six answer engines and record every source URL each engine returns with its answer. This edition covers 74,504 engine answers across 7,573 distinct prompts and 544 scans, from 65 accounts. Citations are counted per engine response; domains keep their subdomain.

This is the industry view. Get yours.

The same pipeline that produced this page will tell you which of these domains cite you, which cite your competitors, and which prompts you are missing from — across all six engines.

Run a free scan