AI Source Analysis
When an assistant describes your brand, it is repeating something it read. Source analysis tells you what it read: which pages the engines lean on, how much each one carries, and which sources are doing that job for your competitors instead.
Mention tracking tells you whether an engine named you. Source analysis tells you what made it name you, by capturing the citations under each answer, grouping them by source type, and showing which sources are carrying your competitors and not you.
How source attribution works in AI answers
An engine builds an answer from two layers. One is what it learned during training: a settled impression of your brand formed before anyone asked the question. The other is retrieval, the live fetch it performs while writing, pulling in pages that exist on the web right now.
Only the second layer leaves a trace. When an engine names or links a page under its answer, that citation can be captured and counted. The training layer shapes how you are framed without linking to anything, which is why an honest attribution report describes influence rather than proof, and why a brand can be described confidently from sources that appear nowhere in the citation list.
That split is also why the two layers respond at different speeds. A new page can change a retrieval-based answer within days. Nothing you publish this week changes what the model already learned.
Words you'll see, in plain English
A source the engine links or names under its answer. The visible half of attribution, and the only half most tools can see.
The wider question of which sources shaped an answer, including ones the engine drew on without linking.
The live fetch an engine performs mid-answer. Pages reached here can change an answer within days.
What the model learned before the question was asked. It shapes framing and changes only when the model is refreshed.
How much of your visibility traces to one domain. A high number on a single source is a concentration risk.
A source your competitors are cited from and you are not. Usually the most actionable row in the report.
How the analysis runs
Four steps, and the first one is why this cannot be done retroactively.
Capture the answers, with their citations
A fixed prompt set runs across the engines on a schedule and every answer is stored with whatever sources it named or linked. Because an answer is generated per request and never published, this has to be captured as it happens.
Resolve and categorise each source
Cited URLs are resolved to domains and grouped by type: your own site, encyclopaedic references, news and press, review and comparison sites, industry publications, community and social.
Weigh how much each one carries
Frequency across the prompt set, which engines lean on it, and whether the citation accompanies a favourable or unfavourable description. A source cited once in passing is not a source the model is built on.
Surface the gaps
Sources cited for competitors and not for you, and sources citing outdated facts about you. Those two lists are what turn a report into a work queue.
Source types tracked
Grouping matters because different question types pull on different categories.
Your own site
Product, pricing and documentation pages. The one category you control outright, and usually not the one carrying the most weight.
Encyclopaedic references
Wikipedia and Wikidata. Heavily relied on for who a company is, and often the oldest description of you still in circulation.
News and press
Coverage, announcements and analyst write-ups. Recency matters here more than anywhere else.
Reviews and comparisons
Category roundups and review platforms. The sources engines reach for most on shortlist questions.
Industry publications
Trade press and specialist outlets. Narrower reach, disproportionate trust inside a category.
Community and social
Forums, Q&A threads and developer communities. Where an old complaint can outlive the thing it described.
What a source report shows
An illustration of the shape these reports take, not a customer story.
- • Your product page
- • Your documentation
- • One encyclopaedic entry, three years old
- • Two independent category roundups
- • A review platform listing
- • A trade publication feature
The finding is the asymmetry, not the count. Both brands are cited. Yours is cited almost entirely from pages you wrote yourself, which engines weigh less on a recommendation question precisely because you wrote them. The work that follows is not more of your own content; it is presence on the third-party sources already carrying the answer.
See what the engines are citing about you
A free scan returns the answers verbatim, with the sources named under them.
What to do with a source report
The report is only worth reading if the next move changes because of it.
Start with the gap, not the volume
The useful finding is rarely that you have few citations. It is that a specific source cites your competitor on the questions that matter and does not cite you.
Match the asset to the source type
If a category roundup carries the weight, another blog post does not answer it. If encyclopaedic references do, the work is factual accuracy and notability, not content volume.
Correct at the source, not on your own site
When an engine cites an outdated third-party page about you, updating your own site does not reach it. The correction has to land where the engine is actually reading.
Watch concentration
Visibility resting on one domain is fragile. If that page changes or drops, so does your presence in the answers built on it.
Source tracking checklist
- A fixed prompt set is running and its answers are stored with their citations
- Cited domains are grouped by type, not read as one flat list
- Sources are separated into ones you own and ones you do not
- Competitor citations are captured from the same runs as your own
- Outdated third-party pages describing your brand are logged as their own work item
- Concentration is checked: no single domain carries most of your citations unnoticed
- The report is reviewed monthly, not only after a campaign
Frequently asked questions
How does source attribution work in AI-generated answers?
An engine builds an answer from two layers: what it learned in training, and what it retrieves live while writing. Attribution is visible only for the second layer, where the engine names or links the pages it used. Those citations are what can be captured and counted; the training layer shapes framing without leaving a link, which is why attribution reports describe influence rather than proof.
How can I track brand citations and source attribution inside AI answers?
By running a fixed set of buyer questions across the engines on a schedule and storing every answer together with the sources it cited. An AI answer is generated per request and never published, so nothing can be crawled after the fact. The citations have to be captured as the answers are produced, and the prompt wording kept fixed for week-over-week movement to mean anything.
What is AI source analysis?
It is the practice of reading which websites and pages AI engines lean on when they describe or recommend your brand, then grouping those sources by type and weight. It answers a different question from mention tracking: not whether you were named, but what made the engine name you.
Which sources do AI models cite most?
It varies by question type rather than by brand. Shortlist and comparison questions pull heavily on review platforms and category roundups; identity questions lean on encyclopaedic references and your own site; anything time-sensitive favours recent news. That is why sources are categorised rather than counted as one list.
Can I see sources for engines that do not show citations?
Partially, and it is worth being precise about the limit. Where an engine exposes citations they are captured directly. Where it does not, the same prompts are run and the answers compared against known sources, which identifies likely influences rather than proving them. A report should distinguish the two, not blend them.
An AI keeps citing an outdated page about us. What do we do?
Correct it where the engine is reading, not only on your own site. If the citation points at a third-party review, comparison or encyclopaedic entry, updating your product page will not reach it. Get the source itself corrected, then watch whether the citation list changes over the following weeks.
How is source analysis different from citation tracking?
Citation tracking counts how often engines link you. Source analysis asks the wider question of which sources are shaping the answer at all, including the ones carrying your competitors, and groups them so the pattern is legible. The two are usually read together.
Related features
Competitor Tracking
Which rivals get named, and which sources put them there
AI Visibility Score
Mention, position and citation as one trend line
AI Model Comparison
How differently each engine sources its answers
For the definition and the metric behind it, see AI citation tracking and source attribution.
Understand Your AI Source Coverage
See which sources engines cite about your brand, and which ones are carrying your competitors instead.