A crowded market selling the same promise

Search for AI visibility tools and you'll find dozens of similarly pitched products, Profound, Peec AI, Otterly, Scrunch, Ahrefs Brand Radar, Semrush's AI toolkit, and many more, most promising to show you where your brand ranks inside ChatGPT, Perplexity, and Google's AI. The pitch is consistent: a dashboard, a score, a position. The evidence for how stable that position actually is gets mentioned far less often.

What a 2,961-prompt study actually found

SparkToro, working with Gumshoe.ai, ran the most rigorous public test of this to date: 12 brand recommendation prompts across categories like chef's knives, headphones, cancer care hospitals, and digital marketing consultants, tested 60 to 100 times per platform across ChatGPT, Claude, and Google's AI, 2,961 total runs over November and December 2025 (via Search Engine Journal's coverage of the SparkToro research).

The result: ChatGPT and Google returned the exact same brand list less than 1% of the time across repeated identical prompts. The same list in the same order appeared less than 0.1% of the time. Run the same query twice and you are very unlikely to see the same "ranking" both times.

Worth sitting with

If a tool tells you that you're "#3 in AI visibility for your category," that specific position is very likely to be different the next time anyone, including that same tool, asks the question again. Treat any single-run ranking number as noise, not a fact about your standing.

The one thing that is stable

The same research found something genuinely useful underneath the noise. Even though exact list order was wildly inconsistent, the same handful of brands kept showing up across runs. For headphone recommendations specifically, across 994 varied prompts, the same four brands appeared in 55 to 77% of all responses. That's a real, measurable pattern, just not the one most tools are selling.

Rate of appearance across many repeated prompts is a stable, meaningful metric. Rank position within a single answer is not. A tool that reports "you appeared in 42% of relevant prompts this month, up from 31%" is telling you something real. A tool that reports "you rank #4" is reporting a coin flip.

Why tools disagree with each other

This also explains a pattern anyone who has trialed more than one of these tools has noticed: they rarely agree. Each tool samples a different set of prompts, runs a different query volume, and checks at a different moment against systems that are, per the research above, inherently non-deterministic by design. Two honest tools measuring the same brand can produce genuinely different numbers without either one being wrong. The instability is in the thing being measured, not necessarily in the measurement.

What these tools are genuinely useful for

  • Directional trend over time, appearance rate rising or falling across a consistent prompt set, month over month.
  • Consideration set membership, whether you show up at all across a category, versus a competitor who never does.
  • Which pages get cited, when a tool shows source links, that's observable fact, not a probabilistic ranking.
  • Prompt discovery, seeing the actual phrasing buyers use is valuable regardless of how stable the resulting answer is.

What to stop expecting from them

  • A precise, stable "rank" comparable to a Google position.
  • Agreement between two different tools measuring the same brand.
  • A single snapshot being representative of your real standing.

I ran my own version of this test at a smaller scale, 30 buyer prompts across five engines, and saw the same underlying pattern the SparkToro data explains: consensus on category leaders, near-total disagreement everywhere else. The full breakdown is in the main AI visibility guide.

A proper audit measures your rate of appearance across a real, repeated prompt set, not a single misleading snapshot. That is what I do.

FAQ
Are AI visibility tools accurate?
They are accurate at measuring rate of appearance across many repeated queries, but not at giving you a stable ranking position. A large scale study running 2,961 prompts found AI systems return the exact same brand list less than 1% of the time, and the same list in the same order less than 0.1% of the time.
Why do different AI visibility tools give different numbers for the same brand?
Because the underlying AI responses themselves are highly non-deterministic. Different tools sample different prompts, run different volumes of queries, and query at different times, all of which produce different snapshots of an inherently inconsistent target.
What should I actually track instead of an AI ranking?
Rate of appearance across a consistent, repeated prompt set over time, which the research shows is a meaningfully more stable metric than any single ranking position or list order.

Primary sources: SparkToro and Gumshoe.ai, 2,961-prompt AI recommendation consistency study, via Search Engine Journal. Original 30 prompt, 5 engine experiment (July 2026). Accessed July 25, 2026.