The AI Citation Index
Everyone argues about what AI search rewards. Almost nobody measures it. This is the source side of the ledger — every domain six answer engines actually cited, taken from 74,504 answers we collected ourselves, published in full and free to reuse.
CC BY 4.0 — republish anything here with a link to rankmonster.ai/sources.
The domains AI answers actually cite
Ranked by breadth — how many separate accounts, in unrelated industries, saw the domain cited in their own scan. Raw citation volume is the wrong sort key: a domain quoted 1,400 times inside three scans of one niche tells you about the niche, not about the engines. By that measure the most universal source in AI answers is reddit.com, which showed up in scans belonging to 63 of 65 accounts.
| # | Domain | Source type | Accounts that saw it cited | Prompts | Citations |
|---|---|---|---|---|---|
| 1 | reddit.com | Community & forums | 63 | 2,282 | 6,709 |
| 2 | linkedin.com | Social networks | 60 | 2,099 | 7,300 |
| 3 | youtube.com | Video | 59 | 1,988 | 7,177 |
| 4 | en.wikipedia.org | Encyclopedias & reference | 57 | 730 | 1,653 |
| 5 | facebook.com | Social networks | 49 | 628 | 1,592 |
| 6 | medium.com | Vendor & independent sites | 43 | 597 | 1,272 |
| 7 | forbes.com | News & media | 43 | 387 | 780 |
| 8 | indeed.com | Vendor & independent sites | 41 | 282 | 730 |
| 9 | gartner.com | Reviews & directories | 40 | 344 | 649 |
| 10 | g2.com | Reviews & directories | 39 | 511 | 1,559 |
| 11 | instagram.com | Social networks | 37 | 252 | 557 |
| 12 | gitnux.org | Statistics aggregators | 33 | 299 | 494 |
| 13 | sciencedirect.com | Academic & government | 33 | 160 | 354 |
| 14 | worldmetrics.org | Statistics aggregators | 32 | 255 | 424 |
| 15 | finance.yahoo.com | News & media | 32 | 109 | 206 |
| 16 | clutch.co | Reviews & directories | 30 | 384 | 1,003 |
| 17 | pmc.ncbi.nlm.nih.gov | Academic & government | 30 | 239 | 512 |
| 18 | wifitalents.com | Statistics aggregators | 30 | 238 | 351 |
| 19 | zipdo.co | Statistics aggregators | 30 | 171 | 254 |
| 20 | f6s.com | Reviews & directories | 29 | 319 | 890 |
| 21 | trustpilot.com | Reviews & directories | 29 | 117 | 440 |
| 22 | blog.hubspot.com | Vendor & independent sites | 29 | 193 | 366 |
| 23 | salesforce.com | Vendor & independent sites | 28 | 239 | 491 |
| 24 | cbinsights.com | Vendor & independent sites | 28 | 131 | 299 |
| 25 | scribd.com | Vendor & independent sites | 28 | 79 | 128 |
Every engine has a different diet
Share of each engine’s citations by the kind of page it pointed at. The spread is the story: these are not six front-ends over one index, they are six different retrieval philosophies. Perplexity sends 3.6% of its citations to forums and community threads; Claude sends 0.3%.
| Source type | Claude | Gemini | AI Overviews | Grok | All engines | ||
|---|---|---|---|---|---|---|---|
Community & forums Threads real people wrote — Reddit, Quora, Stack Exchange, Hacker News. | 2.8% | 0.3% | 0.8% | 3.6% | 2.7% | 1.5% | 2.5% |
Video Video platforms, where the citation points at a watch page or transcript. | 0.2% | 0.1% | 0.8% | 3.4% | 7.0% | 1.0% | 2.4% |
Social networks Profile and post URLs on the big social graphs. | 0.4% | 1.1% | 0.1% | 4.7% | 4.9% | 3.0% | 3.2% |
Reviews & directories Software review marketplaces, agency directories and ratings sites. | 3.3% | 3.9% | 1.1% | 2.9% | 1.1% | 1.8% | 2.6% |
Encyclopedias & reference Wikipedia and the other general-reference encyclopedias. | 2.6% | 0.6% | 0.3% | 0.4% | 0.3% | 0.7% | 0.7% |
News & media Publishers and newswires, from national press to trade titles. | 0.6% | 0.9% | 1.0% | 0.8% | 0.6% | 1.5% | 0.9% |
Academic & government Journals, preprint servers, .edu, .gov and intergovernmental bodies. | 5.1% | 3.2% | 2.5% | 3.1% | 1.0% | 3.4% | 3.1% |
Statistics aggregators Sites whose whole product is scraped statistics round-ups. | 0.3% | 0.3% | 0.1% | 1.0% | 0.1% | 0.1% | 0.6% |
Vendor & independent sites Residual — everything the lists above didn’t recognise. Not shaded: it would flatten every other row. | 84.7% | 89.6% | 93.2% | 80.0% | 82.3% | 86.9% | 84.1% |
A hand-picked list of 1.7k citations to sites whose entire product is scraped statistics round-ups comes to 0.56% of everything — against 0.65% for Wikipedia and every other general reference work combined. And that stats figure is a floor: the bucket only counts the farms we recognised by name. Whatever else it says, it says numbers get cited.
The residual bucket — company sites, docs, blogs, personal pages — is the largest single category by a distance. The platforms get the headlines; the long tail of unremarkable pages gets the citations. That is the part you can actually compete for.
Ranking in one engine barely predicts the next
Give two engines the identical prompt in the identical scan, and compare the sets of domains they cite. The number below is the mean Jaccard overlap — shared domains as a share of all domains either engine used. Gemini and ChatGPT agree on 4.8% of their sources. The best-agreeing pair on the board, Grok and ChatGPT, manages 15.1%. There is no such thing as “ranking in AI”.
| Overlap | Claude | Gemini | AI Overviews | Grok | ||
|---|---|---|---|---|---|---|
| — | 6.0% | 4.8% | 6.6% | 9.3% | 15.1% | |
| Claude | 6.0% | — | 8.8% | 13.5% | 10.0% | 10.5% |
| Gemini | 4.8% | 8.8% | — | 9.3% | 9.3% | 11.7% |
| 6.6% | 13.5% | 9.3% | — | 13.2% | 12.9% | |
| AI Overviews | 9.3% | 10.0% | 9.3% | 13.2% | — | 12.4% |
| Grok | 15.1% | 10.5% | 11.7% | 12.9% | 12.4% | — |
How often each engine cites anything at all
Before overlap can matter, an engine has to show its sources. AI Overviews attaches sources to 85% of its answers; ChatGPT to 29%.
| Engine | Answers with sources | Answers | Citations | Sources per cited answer |
|---|---|---|---|---|
| AI Overviews | 85% | 2,768 | 23,815 | 10.1 |
| 64% | 15,627 | 145,193 | 14.6 | |
| Grok | 50% | 15,584 | 36,976 | 4.7 |
| Claude | 36% | 8,873 | 40,456 | 12.6 |
| Gemini | 36% | 15,839 | 32,921 | 5.8 |
| 29% | 15,813 | 27,555 | 6.0 |
AI does not cite new content
For every page we saw cited and could fetch, we read the publish date out of its own markup and compared it to the first time we recorded an engine citing it. The median page was 203 days old. Only 9% were under a month old; 67% were under a year. Publish-and-wait is not a strategy, but neither is judging a page after three weeks.
| Page age at first observed citation | Share of cited pages | Pages |
|---|---|---|
| Under a week old | 2.8% | 49 |
| 1 week – 1 month | 6.0% | 106 |
| 1 – 3 months | 16.4% | 289 |
| 3 – 12 months | 42.2% | 746 |
| 1 – 3 years | 19.5% | 344 |
| Over 3 years | 13.2% | 233 |
What a cited page is made of
We fetch the pages engines cite and pull their structure apart. This is the composite of 3,974 of them — not what SEO advice says a citable page should look like, but what the pages winning citations right now actually contain.
Format of the cited page
| Format | Share | Pages |
|---|---|---|
| guide | 60.3% | 3,023 |
| listicle | 21.9% | 1,098 |
| landing | 6.9% | 345 |
| comparison table | 5.1% | 258 |
| news | 2.3% | 117 |
| review | 1.6% | 78 |
| docs | 0.9% | 47 |
| video | 0.9% | 46 |
How this is measured, and what it can’t tell you
The sample
74,504 answers, from 7,573 distinct prompts run across six engines in 544 scans belonging to 65 accounts, between April 2026 – August 2026. Every source URL an engine returned with its answer is recorded; domains keep their subdomain, so en.wikipedia.org and blog.hubspot.com are their own rows.
What we exclude
Microsoft Copilot was retired from the product mid-sample and is dropped everywhere — leaving it in would drag every cross-engine number toward a window the other five don’t share. Nothing else is filtered: we do not remove a brand’s own domain from its own results.
The bias we can’t remove
Prompts come from real customer prompt sets, so the corpus over-represents the categories our customers sell into. Ranking domains by how many separate accounts saw them cited suppresses most of that, but a universal claim from a non-random sample is still a claim about this sample.
What is missing
Engine answers are non-deterministic and we have not yet published a repeat-scan variance study, so treat single-point differences between engines with more suspicion than the large ones. Crawler-frequency data — how often GPTBot and ClaudeBot actually fetch a page — is being collected but does not yet cover enough sites to publish honestly.
Questions about the data
- Which sources do AI answer engines cite most?
- Across 306,916 citations collected from six engines between April 2026 – August 2026, the domain seen by the most separate accounts is reddit.com — cited in scans belonging to 63 of 65 accounts. Ranking by breadth rather than raw volume matters: a domain cited thousands of times inside one niche is a property of that niche, not of the engines.
- Does ranking in ChatGPT mean you rank in Gemini?
- No. On the same prompt in the same scan, Gemini and ChatGPT share only 4.8% of their cited domains — the weakest agreement of any pair we measure. The strongest pair, Grok and ChatGPT, still only reaches 15.1%. Answer-engine visibility has to be measured per engine.
- How long after publishing does AI start citing a page?
- Longer than most content calendars assume. Of 1,767 cited pages with a readable publish date, the median was 203 days old when we first recorded an engine citing it, and only 9% were under a month old. That figure is an upper bound — we learn a page is cited when a scan runs, not the moment it happens.
- What does a page that AI cites look like?
- The median cited page runs 1,681 words, carries 15 discrete statistics and 10 self-contained quotable paragraphs, and 61% of them include an FAQ block. Guides and listicles dominate the format mix.
- Can I use this data?
- Yes. Every table on this page has a CSV download and the data is free to republish with attribution and a link to rankmonster.ai/sources. If you want a cut we do not publish — a specific category, a longer window, a different slice — email [email protected] and we will usually run it.
- How is this measured?
- We run fixed prompt sets through six answer engines and record every source URL each engine returns with its answer. This edition covers 74,504 engine answers across 7,573 distinct prompts and 544 scans, from 65 accounts. Citations are counted per engine response; domains keep their subdomain.
This is the industry view. Get yours.
The same pipeline that produced this page will tell you which of these domains cite you, which cite your competitors, and which prompts you are missing from — across all six engines.
Run a free scan