Citations & Sources

AI Source Analysis

When an assistant describes your brand, it is repeating something it read. Source analysis tells you what it read: which pages the engines lean on, how much each one carries, and which sources are doing that job for your competitors instead.

TL;DR

Mention tracking tells you whether an engine named you. Source analysis tells you what made it name you, by capturing the citations under each answer, grouping them by source type, and showing which sources are carrying your competitors and not you.

6
source types tracked
2
layers, only one cites
0
crawlable answer pages

How source attribution works in AI answers

An engine builds an answer from two layers. One is what it learned during training: a settled impression of your brand formed before anyone asked the question. The other is retrieval, the live fetch it performs while writing, pulling in pages that exist on the web right now.

Only the second layer leaves a trace. When an engine names or links a page under its answer, that citation can be captured and counted. The training layer shapes how you are framed without linking to anything, which is why an honest attribution report describes influence rather than proof, and why a brand can be described confidently from sources that appear nowhere in the citation list.

That split is also why the two layers respond at different speeds. A new page can change a retrieval-based answer within days. Nothing you publish this week changes what the model already learned.

Words you'll see, in plain English

Citation

A source the engine links or names under its answer. The visible half of attribution, and the only half most tools can see.

Attribution

The wider question of which sources shaped an answer, including ones the engine drew on without linking.

Retrieval

The live fetch an engine performs mid-answer. Pages reached here can change an answer within days.

Training data

What the model learned before the question was asked. It shapes framing and changes only when the model is refreshed.

Source share

How much of your visibility traces to one domain. A high number on a single source is a concentration risk.

Citation gap

A source your competitors are cited from and you are not. Usually the most actionable row in the report.

How the analysis runs

Four steps, and the first one is why this cannot be done retroactively.

01

Capture the answers, with their citations

A fixed prompt set runs across the engines on a schedule and every answer is stored with whatever sources it named or linked. Because an answer is generated per request and never published, this has to be captured as it happens.

02

Resolve and categorise each source

Cited URLs are resolved to domains and grouped by type: your own site, encyclopaedic references, news and press, review and comparison sites, industry publications, community and social.

03

Weigh how much each one carries

Frequency across the prompt set, which engines lean on it, and whether the citation accompanies a favourable or unfavourable description. A source cited once in passing is not a source the model is built on.

04

Surface the gaps

Sources cited for competitors and not for you, and sources citing outdated facts about you. Those two lists are what turn a report into a work queue.

Source types tracked

Grouping matters because different question types pull on different categories.

Your own site

Product, pricing and documentation pages. The one category you control outright, and usually not the one carrying the most weight.

Encyclopaedic references

Wikipedia and Wikidata. Heavily relied on for who a company is, and often the oldest description of you still in circulation.

News and press

Coverage, announcements and analyst write-ups. Recency matters here more than anywhere else.

Reviews and comparisons

Category roundups and review platforms. The sources engines reach for most on shortlist questions.

Industry publications

Trade press and specialist outlets. Narrower reach, disproportionate trust inside a category.

Community and social

Forums, Q&A threads and developer communities. Where an old complaint can outlive the thing it described.

What a source report shows

An illustration of the shape these reports take, not a customer story.

Tracked question
"which tools should I be considering in this category?"
Cited for you
  • • Your product page
  • • Your documentation
  • • One encyclopaedic entry, three years old
Cited for a competitor
  • • Two independent category roundups
  • • A review platform listing
  • • A trade publication feature

The finding is the asymmetry, not the count. Both brands are cited. Yours is cited almost entirely from pages you wrote yourself, which engines weigh less on a recommendation question precisely because you wrote them. The work that follows is not more of your own content; it is presence on the third-party sources already carrying the answer.

See what the engines are citing about you

A free scan returns the answers verbatim, with the sources named under them.

Run a free scan

What to do with a source report

The report is only worth reading if the next move changes because of it.

01

Start with the gap, not the volume

The useful finding is rarely that you have few citations. It is that a specific source cites your competitor on the questions that matter and does not cite you.

02

Match the asset to the source type

If a category roundup carries the weight, another blog post does not answer it. If encyclopaedic references do, the work is factual accuracy and notability, not content volume.

03

Correct at the source, not on your own site

When an engine cites an outdated third-party page about you, updating your own site does not reach it. The correction has to land where the engine is actually reading.

04

Watch concentration

Visibility resting on one domain is fragile. If that page changes or drops, so does your presence in the answers built on it.

Source tracking checklist

  • A fixed prompt set is running and its answers are stored with their citations
  • Cited domains are grouped by type, not read as one flat list
  • Sources are separated into ones you own and ones you do not
  • Competitor citations are captured from the same runs as your own
  • Outdated third-party pages describing your brand are logged as their own work item
  • Concentration is checked: no single domain carries most of your citations unnoticed
  • The report is reviewed monthly, not only after a campaign

Frequently asked questions

How does source attribution work in AI-generated answers?

An engine builds an answer from two layers: what it learned in training, and what it retrieves live while writing. Attribution is visible only for the second layer, where the engine names or links the pages it used. Those citations are what can be captured and counted; the training layer shapes framing without leaving a link, which is why attribution reports describe influence rather than proof.

How can I track brand citations and source attribution inside AI answers?

By running a fixed set of buyer questions across the engines on a schedule and storing every answer together with the sources it cited. An AI answer is generated per request and never published, so nothing can be crawled after the fact. The citations have to be captured as the answers are produced, and the prompt wording kept fixed for week-over-week movement to mean anything.

What is AI source analysis?

It is the practice of reading which websites and pages AI engines lean on when they describe or recommend your brand, then grouping those sources by type and weight. It answers a different question from mention tracking: not whether you were named, but what made the engine name you.

Which sources do AI models cite most?

It varies by question type rather than by brand. Shortlist and comparison questions pull heavily on review platforms and category roundups; identity questions lean on encyclopaedic references and your own site; anything time-sensitive favours recent news. That is why sources are categorised rather than counted as one list.

Can I see sources for engines that do not show citations?

Partially, and it is worth being precise about the limit. Where an engine exposes citations they are captured directly. Where it does not, the same prompts are run and the answers compared against known sources, which identifies likely influences rather than proving them. A report should distinguish the two, not blend them.

An AI keeps citing an outdated page about us. What do we do?

Correct it where the engine is reading, not only on your own site. If the citation points at a third-party review, comparison or encyclopaedic entry, updating your product page will not reach it. Get the source itself corrected, then watch whether the citation list changes over the following weeks.

How is source analysis different from citation tracking?

Citation tracking counts how often engines link you. Source analysis asks the wider question of which sources are shaping the answer at all, including the ones carrying your competitors, and groups them so the pattern is legible. The two are usually read together.

Related features

For the definition and the metric behind it, see AI citation tracking and source attribution.

Limited: 7-Day Free Trial

Understand Your AI Source Coverage

See which sources engines cite about your brand, and which ones are carrying your competitors instead.

No credit card requiredSetup in under 5 minCancel anytime