Reference
How to track which sources an LLM cites
Which is often. One 2026 study found ChatGPT returned zero citations on 57.8% of runs, and Semrush’s clickstream analysis put ChatGPT’s web search rate at 34.5% of queries in February 2026. Any citation report you read covers the minority of answers where the engine looked something up. That is a useful minority. It is not the whole picture.
What is an LLM citation, and is it the same thing on every engine?
No. Three engines, three quite different contracts.
Perplexity gives you citations as a first-class product. Its Search API returns ranked results with title, URL, snippet and date. Its Agent API returns web-grounded answers with built-in citations and a structured search results object.
Anthropic makes citations non-optional. Its web search tool documentation states citations are always enabled, and each result carries URL, title and up to 150 characters of the cited text. Do not confuse this with Anthropic’s separate Citations API, which grounds answers in documents you supply.
OpenAI distinguishes what it showed from what it read. Its web search guide returns inline citations plus a sources field, and states explicitly that unlike inline citations, which show only the most relevant references, sources returns the complete list of URLs the model consulted. If you are only counting visible citations on OpenAI, you are undercounting influence.
Google is the outlier. There is no citation API. What exists is Search Console’s generative AI performance report, and its documentation makes the limits plain: impressions only, no clicks, no CTR, no queries, no position, with AI Overviews and AI Mode combined rather than separated.
Inline citations are not the full source list
Worth restating because it changes how you read every vendor number. On OpenAI, inline citations are a filtered subset and sources is the full set. A tool scraping the visible answer sees the first. A tool using the Responses API can see both. When two tools disagree about your citation count, this is often why, and neither is lying.
Why do two tools report completely different citation counts?
Capture method. UI capture, API capture and clickstream panels are three different populations. SE Ranking states its AI visibility tracker uses direct UI-based monitoring so results reflect live behaviour. API-based tools get structure and permission instead. Both are defensible. They will not match.
Why the Bing API retirement still shows up in your data
Microsoft retired the Bing Search APIs on 11 August 2025, decommissioning existing instances and pointing customers at Grounding with Bing in Azure AI Agents. That is not a drop-in swap: the old API returned structured results to your application, the replacement feeds context to an agent that writes an answer. Tooling built before that date was rebuilt, not migrated, and comparisons across that boundary are not clean.
Variance. The April 2026 study found day-to-day Jaccard similarity for cited sources of 0.34 to 0.42, lower than the 0.45 to 0.59 it measured for brand mentions. Sources are the least stable thing on the page. It also reported a mean Gini coefficient of 0.715, meaning citations concentrate heavily on a few domains.
Sampling depth. Of eleven vendors whose public pages we checked on 2026-08-20, exactly one publishes how many times it samples a prompt: Evertune, at 100 samples per prompt per model. None publishes a confidence interval. That includes us: Rovoki samples each question once per engine today and publishes no band. Given the 0.34 to 0.42 figure above, single-run citation counts should be read as directional.
Which LLM citation tracking approach should you choose?
Ranked by how directly the method reaches the citation data, which is the axis that matters for this specific job. It is not a ranking of overall product quality.
| # | Approach | Cost | Citation fidelity | Best for |
|---|---|---|---|---|
| 1 | Perplexity Search / Agent API | Usage-based | Highest. Structured source URLs as a product surface | Teams with an engineer and a specific question |
| 2 | Anthropic web search tool | $10 per 1,000 searches | Very high. Citations always on, with cited text spans | Auditing Claude specifically |
| 3 | OpenAI Responses API | Usage-based | High, if you read sources and not just inline citations | Seeing everything ChatGPT consulted |
| 4 | Profound | $99/mo Starter | Strong. The 35,000-URL G2 citation study was captured in it | Citation data without building |
| 5 | SE Ranking | $129 + $89/mo monthly, or $103.20 + $71.20 annual | Strong. Stores cached copies of AI answers | Reading the actual answers cheaply |
| 6 | Ahrefs Brand Radar | Inconsistent across their own pages | Strong on Google surfaces | Existing Ahrefs customers |
| 7 | Semrush AI Toolkit | $117.33/mo SEO plan (bundled), or $165.17/mo AI Toolkit | Good, broad location coverage | Existing Semrush customers |
| 8 | Rovoki | See pricing | Publishes the full raw answer behind every counted question | Checking a parse rather than trusting it. n=1, no bands yet |
| 9 | Otterly.ai | $29/mo | Adequate at the price | Smallest ongoing budget |
| 10 | Evertune | $800/mo | Deepest sampling of any vendor | Anyone who needs the count to survive scrutiny |
| 11 | Google Search Console | Free | Impressions only, AI Overviews and AI Mode combined | The only first-party Google signal that exists |
Read the top three rows honestly: if you have engineering time and one clear question, the APIs are better than buying a tool. They are cheaper, they return more, and nobody is between you and the data. Tools earn their price on continuity, alerting and multi-engine coverage, not on access.
What can citation tracking not tell you?
No citation does not mean no influence. Most ChatGPT answers come from weights rather than retrieval. Your brand can be named confidently in an answer that cites nothing at all. A citation report scores that as a zero.
Citations do not equal recommendation. Being the source an engine read while recommending a competitor is a real and common outcome. Count named-and-recommended separately from cited, or you will celebrate the wrong number.
Correlation is not a lever. Ahrefs measured 75,000 brands and found branded web mentions correlated with AI Overview visibility at 0.664 against 0.218 for backlinks, with YouTube mentions at 0.737 for ChatGPT. Ahrefs says plainly this is correlation, not causation, and the sample is limited to domains above DR 40.
Two popular fixes have been tested and failed. JSON-LD moved AI Overview citations by minus 4.6% against matched controls. And of the valid llms.txt files Ahrefs found, 97% were fetched by nothing at all, with Google on record that it does not use the file.
What do you do with the source list once you have it?
Read it as a target list. The domains that appear repeatedly in your category are the pages the engine will keep reading next month, and they are usually a small set: the Gini figure above says so.
In practice that tends to mean ranked listicles, which Evertune found accounted for 63% of citations across roughly 400 million citations, noting Search Engine Land labels that piece sponsored vendor content. It also tends to mean community pages, at 17.1% of cited domains in the G2 study. Getting onto the pages already in your citation report is a more direct route than trying to displace them.
Sources, all checked 2026-08-20
- Schulte, Bleeker and Kaufmann, Don't Measure Once, April 2026
- Semrush, ChatGPT clickstream search insights
- Perplexity, Search API quickstart
- Perplexity, Agent API quickstart
- Anthropic, web search tool documentation
- Anthropic, Citations API
- OpenAI, web search tool guide
- Google Search Central, generative AI performance reports
- Google Search Console help, generative AI report
- Microsoft, Bing Search API retirement
- SE Ranking AI visibility tracker
- Evertune pricing
- Profound pricing
- Ahrefs Brand Radar
- Semrush AI visibility
- Otterly.ai pricing
- Indig and Johnson via Search Engine Land, community signals
- Ahrefs, AI Overview brand correlations
- Ahrefs, AI brand visibility correlations
- Ahrefs, schema and AI citations
- Ahrefs, llms.txt study
- Search Engine Land, Google on llms.txt
- Thinking Machines Lab, defeating non-determinism in LLM inference
- Evertune via Search Engine Land, listicles study (sponsored)
- G2 via PR Newswire, AEO category growth