Across 731,000 tracked AI answers, Perplexity attached 13.69 cited sources to an average answer while ChatGPT attached 4.76 and Claude carried a citation on only 4.0% of its answers. Gemini names brands most often, in 40.9% of answers, while citing just 5.24 sources. Google is two engines rather than one: AI Mode and AI Overviews share only 13.7% of their cited URLs on the same prompt, and 51.2% of brand appearances on Google exist on one surface and not the other. A single blended AI visibility score hides all of it.
The same brand, the same question, on the same day, gets a different answer from every AI assistant, and the gap is wide enough that a single blended visibility score is close to meaningless.
The clearest way to see it is the citation slot count. Across 731,000 tracked answers in a 14 day window, Perplexity attached 13.69 sources to an average answer and ChatGPT attached 4.76. That is not a small difference in degree. It is a different competition.
Two measurements matter and they do not correlate. How many sources an engine cites decides how hard it is to get a slot. How often it names a brand at all decides whether the question is even worth tracking.
| Engine | Answers measured | Cited sources per answer | Answers carrying any citation | Brand named |
|---|---|---|---|---|
| Perplexity | 140,334 | 13.69 | 100.0% | 27.0% |
| Google AI Mode | 49,292 | 12.95 | 94.6% | 38.3% |
| Google AI Overviews | 219,495 | 8.73 | 96.2% | 27.0% |
| Microsoft Copilot | 10,973 | 6.30 | 96.6% | 16.7% |
| Gemini | 65,778 | 5.24 | 82.2% | 40.9% |
| ChatGPT | 242,290 | 4.76 | 99.3% | 27.9% |
| Grok | 797 | 2.14 | 66.6% | 68.6% |
| Claude | 31,313 | 0.05 | 4.0% | 21.8% |
Read the last two columns against each other and the strategy falls out of the table.
Perplexity is the cheapest place to earn a first citation. Every answer carries sources and there are roughly fourteen slots. If nobody at your company has ever been cited by an assistant, this is where it happens first.
Gemini names the most brands and cites the fewest sources. A 40.9% mention rate on 5.24 citations means Gemini talks about brands without linking to them. Optimising for citations there misses the point; the work is being the brand it names.
Claude is a different product entirely. Only 4.0% of Claude answers carry any citation at all, and the average is 0.05 sources. Tracking Claude for citations is tracking something that mostly does not happen. Tracking it for mentions is legitimate, and the two get confused constantly.
Grok's row is a thin sample at 797 answers and should be read as directional only. The others are large enough to trust.
The most expensive reporting mistake in this category is a single row labelled Google.
Across 42,937 prompt days where AI Mode and AI Overviews both answered the same prompt on the same day, the two surfaces shared only 13.7% of their cited URLs and 18.7% of their cited domains. Ahrefs measured the same 13.7% independently across 540,000 query pairs.
The brand consequence is sharper. Of the prompt days where a tracked brand was named on Google at all, 51.2% of those appearances existed on one surface and not the other. AI Mode named the brand on 35.8% of prompt days against 27.3% for AI Overviews, so the surface more often missed is also the more generous one.
They also draw from differently shaped pools. AI Mode puts 48.9% of its citations into a hundred domains; AI Overviews reaches only 37.0% at the same cut. AI Mode cites more per answer from a narrower regular cast.
Two surfaces, one search engine, and no useful average between them. Tracking them separately is the whole point of Google AI Mode visibility tracking as its own line rather than a Google subtotal.
It is not randomness, though answers do vary between runs. Three mechanisms explain most of the divergence.
Different retrieval. Assistants fan a question into several sub queries before answering, and they do not generate the same branches. Two engines asking different sub questions retrieve different pages, so they cite different sources while reaching a similar conclusion.
Different citation policy. Perplexity is built to show its sources. Claude, in ordinary chat, largely is not. That is a product decision, not a ranking outcome, and no amount of content work changes it.
Different index freshness and grounding. Some engines ground every answer in a live search. Others answer from training and search only when the question demands it. The same page can be authoritative to one and invisible to the other.
Knowing the engines diverge is only useful if your tooling reads all of them, and most platforms tier the engine list. From each vendor's own pricing page in September 2026:
| Platform | Engines on the entry plan | Engines at the top tier |
|---|---|---|
| Finseo | 15 | 15 |
| Rankscale | 10 or more, all plans | Same |
| Searchable | Up to 9, selected in onboarding | 9 |
| Ahrefs Brand Radar | 6 | 6 |
| Otterly.ai | 4 included, others priced separately | 7 |
| Scrunch AI | 4 | 9 |
| Peec AI | 3 of your choice | Full list |
| Semrush AI Visibility Toolkit | 3 of your choice | Adds Claude, DeepSeek, Grok |
| ZipTie | 3, fixed on every plan | 3 |
| Profound | ChatGPT only | 10 |
Read the entry column, not the marketing page. Profound's Starter plan tracks ChatGPT only, and ZipTie covers three engines at every price with no add-ons available. If your buyers use Gemini and Perplexity, several rows here are eliminated before any feature comparison begins.
Pick the engine before the tactic. A page that would be one of fourteen sources on Perplexity might be fighting for one of five on ChatGPT. The same edit has very different odds depending on where you aim it.
Stop reporting a blended score. If your dashboard shows one AI visibility number, ask which engine moved when it changes. If nobody can answer, the number is decoration.
Match the metric to the engine. Citations on Perplexity and Google AI Mode. Mentions and sentiment on Gemini and Claude. Using citation rate as the KPI for Claude will show a flat line forever.
Track the engines your buyers use, not the ones that are easy. In 45,921 post purchase survey responses, among people who named a specific assistant as how they found a company, ChatGPT accounted for 44.9%, Gemini 14.7%, Perplexity 11.9%, Copilot 8.3% and Claude 5.9%. ChatGPT is close to half and nowhere near all of it.
Expect the sets to move independently. Because the underlying page sets barely intersect, a content change that lands on one engine may not show on another for weeks. That is an argument for per engine reporting rather than a single trend line, which is what AI visibility tracking across the answer engines is shaped around.
One row per engine, with the metric that engine actually supports:
| Column | Why it belongs |
|---|---|
| Engine | The unit of analysis, always |
| Prompts tracked | So a rate has a denominator |
| Mention rate | The only metric every engine supports |
| Citation rate | Meaningful on Perplexity, AI Mode, AI Overviews, ChatGPT |
| Cited sources kept | Evidence, and the input to the next content decision |
| Sentiment | Where the engine names brands without citing them |
| Change versus last period | Because a single reading is an anecdote |
The column people leave out and then miss is the fifth. A rate tells you that you have a problem. The stored source list tells you which page took your slot, which is the job AI citation tracking does per answer.
Everything in this article reduces to one purchasing question: does your tooling read the engines your buyers use, on the plan you can afford.
Two platforms do not make you choose. Finseo covers fifteen engines on every plan including the entry one, and Rankscale covers ten or more across its tiers. Those are the only two rows in the table above where the entry column matches the top column, and that matters more than any feature comparison, because an engine you cannot see is a blind spot chosen by procurement rather than by strategy.
Finseo is the pick if the answer has to lead somewhere. Same untiered engine list, plus the layer none of the others have: attribution that ties an AI answer back to a real deal value through HubSpot, Salesforce, Stripe or Shopify. Given that most AI referrals arrive with no referrer at all, that is the difference between a visibility percentage and a number a finance team accepts. Rankscale is genuinely close on coverage and does not do this part.
Everything else in the table asks you to pick engines. Peec AI and the Semrush toolkit give you three of your choice. ZipTie gives three at every price with no add-ons. Profound's Starter plan gives one. If your buyers ask ChatGPT, Gemini and Perplexity, those constraints decide the shortlist before you compare a single feature.
Whatever you choose, split Google into two lines. AI Mode and AI Overviews share 13.7% of their cited URLs, and a combined Google number will hide the surface that actually moved.
Figures attributed to Finseo come from tracking data aggregated across accounts, with no customer, project or private domain identifiable in any number. The per engine table covers 731,000 answers in a 14 day window and states the answer count per engine. The Google surface comparison covers 42,937 prompt days for citation overlap and 45,024 prompt days for brand co-occurrence, both in 14 day windows, and drops any prompt day where one surface returned no answer. The survey figures cover 45,921 responses across 12 projects between 9 March and 8 September 2026.
Prompt sets are chosen by the customers who run them, so the corpus is large but not random. It skews toward B2B software, ecommerce and professional services. The Grok sample is small and labelled as such.
External figures are attributed to their study. The Ahrefs replication of the 13.7% overlap covers 540,000 US query pairs from September 2025 and is a vendor study.
Which AI engine cites the most sources? Perplexity, at 13.69 cited sources per answer, followed by Google AI Mode at 12.95 and Google AI Overviews at 8.73. ChatGPT averages 4.76.
Why does Claude almost never cite anything? Because ordinary Claude chat is not built as a citing search product. Only 4.0% of the Claude answers we measured carried any citation, at an average of 0.05 sources per answer.
Do AI Mode and AI Overviews show the same sources? Largely not. They shared 13.7% of cited URLs across 42,937 prompt days, and Ahrefs found the same figure independently.
Which engine should I optimise for first? Perplexity for a first citation, because every answer carries sources and there are around fourteen slots. ChatGPT for commercial impact, because it accounts for roughly half of self reported AI discovery.
Which AI visibility tools track the most engines? On the entry plan, Finseo at fifteen and Rankscale at ten or more cover the widest set without tiering. Several platforms gate engines heavily: Profound's Starter plan tracks ChatGPT only, and ZipTie covers three at every price point.
Is one AI visibility score useful at all? As a headline for people who will not read further, yes. As a working metric, no. The engines diverge enough that an average hides the movement you need to act on.
Ours
External
Disclosure: the tracking and survey figures in this article come from Finseo, an AI visibility platform. External figures are attributed to their original source.