Where AI Citations Come From (Mostly Not Your Site)
Across 25 million links cited by AI engines, 84% pointed at domains the brand doesn't own. Here's what that changes about where you spend your AEO effort.
The standard AEO checklist is almost entirely about your own website: headings, schema, llms.txt, answer blocks, freshness dates. The citation data says that is a minority of the game. Muck Rack's May 2026 edition of What Is AI Reading? analysed more than 25 million links cited by ChatGPT, Claude, and Gemini across 17 industries, and attributed 84% of those citations to sources outside the cited brand's own domain. Paid and advertorial content accounted for 0.3%.
Key takeaways
- 84% of AI citations pointed at domains the brand does not own. Across three editions of the study going back to July 2025 the figure has held between 82% and 89%, which makes it a property of how these systems source information rather than an artefact of one model update.
- Read the taxonomy before the headline. Muck Rack's "earned media" bucket includes journalism, academic work, government and NGO sources, encyclopedic sites, social platforms, and other companies' corporate content. It is closer to "not yours and not paid for" than to "press coverage."
- Journalism alone was 27% of citations, and stayed in a 25–27% band across all three editions.
- The mix is engine-specific. Wikipedia was 7.8% of ChatGPT's citations but 0.6% of Google AI Overviews' and Perplexity's in Profound's 680-million-citation analysis; Reddit ran at 6.6% on Perplexity against 1.8% on ChatGPT.
- Query type shifts it again. Trend questions cited journalism in 46% of responses, more than twice the rate of how-to and comparative questions.
- None of this makes on-site work worthless. It puts a ceiling on it, and the ceiling is lower than most AEO checklists imply.
What the 84% actually counts
The number is easy to misread in a flattering direction, so start with the definition. Muck Rack groups cited sources into earned, owned, and paid, and its earned category covers journalism, academic research, government and NGO material, encyclopedic sites like Wikipedia, social platforms, and third-party corporate content.
That last item is the one to notice. Another company's blog post about your category counts as earned media from your point of view. So does a Wikipedia article, a Reddit thread, a university page, and a government dataset. The honest reading of 84% is not "press coverage drives AI answers." It is 84% of citations pointed somewhere you cannot edit.
Muck Rack is a PR software company, and its framing leans toward the journalism slice of that bucket, which is the slice its product addresses. Worth holding in mind. It is also worth noting that the broader reading survives the discount: whether the off-site share is 84% or somewhat lower under a stricter taxonomy, it is not a rounding error, and it has been stable for a year across three separate measurement runs.
| Category | Share of citations | What it includes |
|---|---|---|
| Earned | 84% | Journalism, academic, government and NGO, encyclopedic, social platforms, third-party corporate |
| Journalism (subset of earned) | 27% | Editorial coverage from news outlets |
| Paid and advertorial | 0.3% | Sponsored placements, advertorials |
The 0.3% deserves a moment on its own. Whatever else is true about AI search, you cannot buy your way into the citation set. That is a meaningful difference from every paid channel marketers are used to reaching for when an organic number goes the wrong way.
Independent studies point the same direction
One vendor study with a self-serving frame would not be enough. Two more, with different methods and different commercial incentives, land in the same place.
Semrush tracked the top 25 most-cited domains weekly across ChatGPT Search, Google AI Mode, and Perplexity for 13 weeks, covering 230,000-plus prompts and over 100 million citations. The names at the top were Reddit, Wikipedia, LinkedIn, YouTube, and Google's own properties. Not a single brand's product site among them.
Profound's analysis of roughly 680 million citations found commercial .com domains taking 80.4% of all citations and .org 11.3%, with the leading individual sources again being platforms and publishers rather than vendors.
Our own 42-day citation study shows the same thing from the other end. Google AI Mode cited about 24 domains a day and ChatGPT about 17, and for a single broad prompt — "alternatives to Nike for high-performance running" — Google drew on 445 distinct domains over six weeks. Any one brand's own site is one entry in a pool that size, competing against hundreds of pages written about the category rather than by a participant in it.
The mix is different on every engine
"Go off-site" is not a strategy until you know which off-site. The engines diverge more than they converge.
| ChatGPT | Google AI Overviews | Perplexity | Claude | |
|---|---|---|---|---|
| Share of responses with citations | 96% | — | — | 55% |
| Average citations per response | ~5 | — | — | ~13 |
| Wikipedia share of citations | 7.8% | 0.6% | 0.6% | — |
| Reddit share of citations | 1.8% | 2.2% | 6.6% | — |
| YouTube in top sources | No | 1.9% | 2.0% | — |
| Character of the source mix | Encyclopedic and factual | Balanced across social and professional sources | Community discussion | Cites least often, most sources when it does |
Citation rates are from Muck Rack (Gemini sat at 82% of responses and about 8 citations each); domain shares are from Profound. Blanks are where the studies do not report a comparable figure, which is itself a reminder that no single dataset covers every engine cleanly.
Two practical consequences. First, a category whose buyers use Perplexity should treat community presence as core infrastructure, while one whose buyers use ChatGPT should care disproportionately about how reference sources describe the category. Second, these shares move: inside a single 13-week window Semrush watched Reddit's presence in ChatGPT responses fall from roughly 60% to about 10%, and Wikipedia's from around 55% to under 20%. Build a strategy on one engine's current source preferences and you have built on sand.
Query type moves it again
Within a single engine, what gets cited depends on what was asked. Muck Rack found industry trend queries cited journalism in 46% of responses, more than double the rate for how-to and comparative queries, which lean on reference material and brand-owned pages instead.
This is the finding that makes the whole thing actionable, because it means your own prompt mix decides where off-site effort pays. A brand whose tracked prompts are mostly "how do I do X" questions has more on-site leverage than the 84% headline suggests. A brand whose prompts are mostly "what is happening in this market" questions has almost none, and needs editorial coverage. Most brands have both and have never separated them.
So separate them. The prompt set you already track can be split by question type in an afternoon, and the split tells you which half of the work to fund.
What this does not prove
Being straight about the limits matters more than the headline here, because "your website doesn't matter" is a conclusion the data does not support and one that would cost you real visibility.
These are citation counts across all queries in an industry, not brand-specific commercial queries. Nobody has published the equivalent breakdown for the narrow, high-intent prompts that decide a shortlist in your category, and there is good reason to expect vendor documentation and pricing pages to do better there than they do across a general corpus. That experiment is the one you can run yourself, on your own prompts, and it is described in the next section.
The categories are also softer than the percentages imply. A bucket that contains both a Reuters investigation and a competitor's blog post is not measuring one thing, and the published summaries do not give a clean own-domain share to set against the 84%.
Most importantly, on-site work is a precondition, not a competitor to off-site work. If AI crawlers cannot reach your pages, or your content is unparseable, or your facts are stale, you forfeit the citations you would otherwise win — including the how-to and comparative queries where brand-owned pages do get cited. Google's own guidance on AI features is explicit that its generative features run on core Search ranking and that there is no separate playbook. Structured data and clear answer structure remain worth doing. They are just not the whole job, and the checklist industry has been selling them as if they were.
Digiday's July 2026 survey of agency practitioners lands on the same balance. Nate King of RPA put the overlap with conventional SEO at "80 to 90% of the same tactics" while stressing that the remaining "10 to 20% is super critical" — and VML's Heather Physioc was blunter about the vendors promising more: "Anyone promising anything… is a liar."
How to audit your own citation source mix
The general studies tell you the shape of the problem. Only your own data tells you where to act, and the audit is a couple of hours of work.
- Export every cited domain from your tracked prompts, on every engine, including the runs where your brand was never mentioned. Those runs are the informative ones: they show which sources are carrying an answer you are absent from.
- Group the domains by type. Your own site, editorial outlets, forums and communities, review platforms, reference sites, video, competitors' own domains. A ranked list of 300 domains is data; seven buckets with percentages is a decision.
- Compute your own domain's share of all citations. That number is the ceiling on what on-site optimisation can reach for these prompts. If it is 6%, you now know what the other 94% of your effort should be aimed at.
- Rank the third-party domains by citation count. The top ten or twenty are your real target list. It will be category-specific and it will not match Reddit-Wikipedia-LinkedIn, because your buyers ask narrower questions than a general corpus samples.
- Open the cited pages and read them. Are you listed? Described accurately? Priced correctly? Present at all? A review profile three versions out of date and a comparison thread that omits you are concrete, fixable gaps, and fixing them is faster than earning a new citation from scratch.
- Re-run it quarterly. Source preferences shift fast enough that a target list built a year ago is pointing at the wrong surfaces.
Then work the list. The channels that move off-site citations are the unglamorous ones: accurate review-platform profiles, genuine participation in the communities where your category is discussed, relationships with the journalists who actually cover your market, and documentation good enough that other people cite it. None of it is a hack, which is precisely why it holds. As Digiday's practitioners noted, the manipulation-shaped tactics — keyword stuffing, link schemes, llms.txt as a magic file — are the ones showing the least effect.
Where this leaves the checklist
On-site AEO became the default advice because it is the part of the problem a brand fully controls, and controllable work is easier to sell, scope, and report on. The citation data is a reminder that the controllable part is not the large part.
The reframe worth carrying: your website is how you become citable, and the rest of the web is what decides whether you get cited. Both halves need budget, and for most brands the second half currently has none.
Elmo is an open-source, self-hosted AI visibility platform that runs your prompt sets across ChatGPT, Claude, Gemini, Perplexity, and Google's AI surfaces, recording every citation, mention, and competitor named alongside you. Because it stores every cited domain and the data is yours to query directly, the source-mix audit above is a grouped query over your own results rather than a research project.
For the fundamentals, start with AI citations and answer engine optimization, then how to track your brand in AI search. For how the surfaces themselves are growing, see AI Overviews' share of US searches. For the vocabulary, see the AI search glossary.
Frequently asked questions
What share of AI citations come from a brand's own website?
A minority. Muck Rack's May 2026 edition of "What Is AI Reading?" analysed more than 25 million links cited by ChatGPT, Claude, and Gemini across 17 industries and attributed 84% of citations to earned media, meaning sources outside the brand's own domain. The figure has stayed between 82% and 89% across three editions since July 2025.
Does that mean on-site AEO is pointless?
No. It means on-site work has a ceiling, not that it has no value. Your pages still have to be crawlable, parseable, and accurate or you cannot be cited even on the queries where you would win, and Google's own guidance is that its AI features run on core Search ranking. The finding argues for rebalancing effort toward off-site surfaces, not abandoning your own site.
Which websites do AI engines cite most?
Across published studies the recurring names are Reddit, Wikipedia, LinkedIn, YouTube, and large editorial outlets. Semrush tracked 100 million-plus citations over 13 weeks in 2025 and found those domains consistently at the top. But the mix differs sharply by engine, and your own category's list will differ from any published one.
Do different AI engines cite different kinds of sources?
Substantially. In Profound's analysis of roughly 680 million citations, Wikipedia accounted for 7.8% of ChatGPT's citations but only 0.6% of Google AI Overviews' and Perplexity's, while Reddit ran at 6.6% on Perplexity against 1.8% on ChatGPT. Engines also cite at different rates: Muck Rack recorded citations in 96% of ChatGPT responses, 82% of Gemini's, and 55% of Claude's.
Does the type of question change which sources get cited?
Yes. Muck Rack found industry trend queries cited journalism in 46% of responses, more than double the rate for how-to and comparative queries, which lean instead on reference material and brand-owned pages. Your prompt mix therefore determines which off-site surfaces are worth the investment.
How do I find which third-party sites shape my AI visibility?
Export the cited domains from your tracked prompts on every engine, including the runs where your brand was not mentioned, then group them by type and rank them by citation count. The top off-site domains for your prompts are your target list. Open the cited pages and check whether they represent you accurately.