Reference
AI visibility tracking: what a real measurement looks like
That is a competitor’s product and we are naming it first, because the number that decides whether a visibility dashboard is trustworthy is the one almost no vendor prints.
What makes a visibility number trustworthy?
Not the engine count. Not the size of the prompt library. The number of times each question was actually asked. AI answers are not stable. Ask the same question twice on the same afternoon and the set of brands named partly changes. This is not a bug in any particular tool, it is a property of the systems being measured, and it means a percentage from a single run is closer to an anecdote than a metric.
The number nobody prints
Every product in this category sells you a percentage. Almost none of them tells you the sample size behind it. That single missing figure is the difference between a measurement and a decoration, and it is the first thing to ask for on a sales call. Evertune states theirs: up to 100 samples per prompt per model, across 11 models, on a Pro plan at $800 a month. Whatever else is true of them, they answered the question.
What n does to a percentage
The best public evidence comes from SparkToro. Six hundred volunteers ran 12 prompts across three engines, 2,961 runs in total. There is under a 1 in 100 chance that any two responses to the same prompt return the same list of brands, and roughly 1 in 1,000 that they return that list in the same order.
Rand Fishkin’s conclusion is the important half, and it is more optimistic than the headline: visibility measured as a percentage across dozens to hundreds of prompts, run multiple times, is a reasonable metric. Repetition is what converts instability into signal. The academic framing agrees: Schulte, Bleeker and Kaufmann argue visibility must be characterised as a distribution rather than a single-point outcome. “Don’t measure once” is the title, and it is also the whole finding.
Which AI visibility tracking tools are worth a buyer’s shortlist?
Ranked on the strength of the measurement, not on price or feature count. Prices read from vendor pages on 2026-08-20.
| # | Tool | Price | The measurement claim | Honest caveat |
|---|---|---|---|---|
| 1 | Evertune | $800/mo Pro | Up to 100 samples per prompt per model, 11 models | Entry price rules out most teams |
| 2 | Profound | $99/mo annual, $399 Growth | Consumer panel plus Conversation Explorer on observed volume | Starter tracks ChatGPT only |
| 3 | Ahrefs Brand Radar | $398/mo select, $699 all | Real monthly prompts, plus YouTube, TikTok and Reddit | Observed demand can be thin in narrow niches |
| 4 | Scrunch | $250/mo annual Starter | Personas, geo and industry segments, live bot crawl feed | Sampling frequency is not stated publicly |
| 5 | Semrush | $117.33/mo SEO plan (bundled), or $165.17/mo AI Toolkit | Bundled tracking backed by a 126M-prompt index | Prompt allowances are per-day caps, not depth |
| 6 | Peec AI | Not readable on their page | Daily tracking, prompt tagging, competitor benchmarking | We could not verify a price, so we publish none |
| 7 | SE Ranking | $103.20 + $71.20/mo annual, or $129 + $89 monthly | 200 prompts, daily, on top of a Core plan | An add-on, not a dedicated instrument |
| 8 | Otterly.ai | $29/mo Lite | Daily tracking, 4 engines, 50+ countries. Claude and Gemini are paid add-ons | Daily n=1. Excellent value, thin per-day evidence |
| 9 | Rovoki | See pricing | 4 engines reported separately, 18 countries, 31 languages, every raw answer published | n=1 per run today, no confidence band shipped |
| 10 | AthenaHQ | Free tier, $295/mo Starter | 9 models on Starter, credit-metered | Credits make depth a budgeting decision |
Why does every tool give you a different score?
Because they are not measuring the same thing, and mostly not the same universe. Semrush’s 2026 index analysed 126 million US prompts across four surfaces, covering more than 1,200 brands. Only 36 of those brands appeared in the top 100 most-mentioned list on every platform in every month of the window. The same study found ChatGPT cites an average of 15 sources per response while Gemini cites about 3.
If a brand’s presence differs that much between platforms, then two tools covering different engine mixes will disagree by construction, before you get to prompt sets, country settings, language, run counts, or how each vendor defines a mention. Dan Taylor’s argument is that most of these tools present a closed sandbox as though it were the open web.
The practical consequence: never compare a score from one tool to a score from another. Compare a tool to itself over time, on a fixed prompt set, and insist the vendor tells you when that prompt set changed.
What does Evertune do better, and where does it not fit?
Better: they publish their method and their research, and the research is checkable. Their citation study covered roughly 25,000 of the most-cited URLs across six engines, finding listicles at 63% of all citations, with the caveat that Search Engine Land labels the piece sponsored vendor content. A separate exercise produced the cited-page anatomy half this industry now quotes: about 941 words, 4 H2s, 2 H3s, 15 external links, 10 images.
Where it does not fit: $800 a month is not a starting point for a company with one product and twenty questions to track. At that scale, Otterly at $29 or AthenaHQ’s free tier will tell you the same broad thing for a hundredth of the money, and the honest recommendation is to start there and graduate.
It is also worth remembering what the number is for. AI referrals are still around 1% of total website visits, though they convert at roughly 7%. And Pew found users click a link inside a Google AI summary about 1% of the time. Visibility tracking is a leading indicator of being named, not a traffic report. Anyone selling it as the second thing is selling you the wrong chart.
What are we not able to do yet?
Rovoki measures ChatGPT, Perplexity, Gemini and Claude, reported separately, across 18 countries and 31 languages or worldwide, returns a Visibility Standing from 0 to 100, and publishes the raw answer behind every question it counted.
It samples n=1 per run and does not publish confidence bands. By the standard set out at the top of this page, that makes a single Rovoki run directional rather than conclusive, and we would rather say so here than have you discover it in month two. Bands are what we are building toward. They are not shipped. Until they are, Evertune’s sampling depth is a real advantage over us and we are not going to pretend otherwise.
What we do have is auditability. Every counted answer is published in full with its citations, so a number you disagree with can be opened rather than argued about. If your blocker is depth, buy depth. If your blocker is that nobody can check the dashboard, that is the problem we built for.
One last piece of context on where to spend effort: across 75,000 brands, unlinked branded web mentions correlated with AI visibility at 0.664 on ChatGPT against 0.266 for domain rating, with YouTube mentions the strongest single signal at about 0.737. The authors are explicit that correlation is not causation. Getting written about is doing more work than getting linked to.
Sources, all checked 2026-08-20
- Evertune pricing
- Evertune AI search statistics
- Profound pricing
- Ahrefs Brand Radar
- Scrunch pricing
- Semrush pricing
- Peec AI
- SE Ranking AI Visibility Tracker
- Otterly.ai pricing
- AthenaHQ pricing
- SparkToro, AI recommendation inconsistency
- Schulte, Bleeker and Kaufmann, Don't Measure Once
- Semrush 2026 AI Visibility Index
- Search Engine Land, the problem with AI share of voice
- Search Engine Land on Evertune's citation study (sponsored)
- Similarweb, AI search statistics
- Pew Research Center, clicks and AI summaries
- Ahrefs, AI brand visibility correlations