In this exclusive opinion piece, Laura Hamod Barnes, founder and CEO of Connected Media, looks at the flaws in simplistic visibility scores and free AI audits and why she believes it often creates more fear than insight.
Picture this scenario, and for some reading it’s likely already a lived experience.
An email lands in your inbox, subject line ‘Your brand is invisible on ChatGPT’. Ok, that’s definitely not good. Inside is a score out of 100, a few red flags next to your competitors’ names, and an invitation to book a call. It gets forwarded internally with the inevitable question: should we be worried about this?
Earlier this month, the IAB released the first standardised framework for AI visibility measurement, noting there are now more than 20 companies hawking AI visibility tools, each using a different methodology which in turn produce different answers for the same brand, the framework draws a formal line between what it calls ‘decision-grade’ data and merely ‘directional data’. It’s a distinction marketers and brand managers have to pay close attention to.
There’s a general understanding that AI search is becoming a meaningful part of how people discover, compare and evaluate brands and if a business or brand is consistently absent from those environments, there may be a genuine commercial risk.
But before a low visibility score becomes a new budget line, a more important question needs to be asked.
Who created the prompts behind the score, and were they designed around actual customers, category and buying journey? Because, the number on the dashboard means very little without knowing that.
Measurement is still maturing
There’s a good reason to pay attention to AI visibility. What ranks prominently in traditional search doesn’t automatically appear in AI-generated answers, and citation behaviour can vary significantly between platforms.
Researchers at the University of Toronto compared the same search queries across Google, GPT-4o, Gemini, Claude and Perplexity and found overlap with Google’s own top 10 results ranging from just 4 per cent on GPT-4o up to 15.2 per cent on Perplexity, a near four-fold gap between platforms answering identical questions.
That matters because AI visibility is often presented as a single metric when, operationally, it’s anything but.
ChatGPT, Gemini, Perplexity and Copilot aren’t different flavours of the same channel.
They’re different products, built by different companies, pulling from different sources, used by different people for different reasons. Averaging them into one score doesn’t make reporting simpler, it just hides where a brand or business is winning and where it doesn’t exist at all.
A separate academic study analysing 167,551 AI citations across 128 brands, 12 markets and 13 languages found that just 14.3 per cent pointed to the brand’s own website. The other 85.7 per cent were spread across third-party sources: reviews, forums, comparison platforms, publishers and other outside sites.
If an AI visibility audit mostly checks a business or brands’ own site, it’s measuring just a fraction of what’s actually influencing whether your brand gets surfaced, cited or recommended.
In May Google published its first official guidance on optimising for generative AI search. Its position was fairly straightforward. SEO remains relevant, because Google’s AI search experiences still lean on many of the same ranking and quality systems that underpin traditional search.
This is a useful distinction. AI visibility is not a completely separate marketing discipline requiring an entirely new playbook. Much of the underlying work remains the same. Strong technical foundations, useful content, authority, structured information and a credible presence across the wider digital ecosystem. But what’s changing is the measurement layer.
Brands must understand where they appear, how they’re described, which sources are shaping those answers, how results vary by platform, and how visibility shifts over time. While the discipline isn’t entirely new, the monitoring is.
Where AI visibility audits are misleading
A familiar tactic is emerging when it comes to AI search. As AI makes audits and diagnostics faster and cheaper to produce at scale, free reports have become an increasingly common way for some agencies to open a new business conversation. The problem isn’t the audit itself. It’s when the findings are designed to create urgency rather than genuinely help a marketer understand what matters.
A well-designed diagnostic can be genuinely useful, surfacing missed demand, weak category association, thin third-party authority, or places where competitors are pulling ahead, the pinchpoint comes around methodology. But, prompt sets are often thrown together quickly. However, meaningful prompt architecture, built around a brand’s actual buyers, takes considerably more thought, which is exactly why the quick version tends to be what a prospect sees, and the properly built one only shows up once they’re signed as a client.
For example, a generic audit might test a premium skincare brand against searches built around ‘cheap’, ‘budget’ or ‘affordable’. An ecommerce-only retailer could get measured against location-based queries. A B2B company might be assessed using consumer language its buyers would never actually use. The model can answer every one of those prompts correctly, but the resulting score can still be commercially meaningless.
Put the wrong questions in, and it’s easy to manufacture a visibility problem that is not there. Which is exactly why marketers should be looking at how the audit was built, not just how bad the result looks.
While anyone can build a cheap AI visibility tool, the real value is in understanding the customer journey well enough to write the right prompts; separating meaningful patterns from noise; connecting AI visibility to search behaviour, content strategy, brand authority, media investment, and ultimately, commercial outcomes. That takes careful judgement and collaboration.
The difference is rarely the dashboard, it’s the thinking behind it.
AI visibility monitoring can only become a genuinely valuable addition when the methodology is robust enough to provide truly meaningful results.
Otherwise, marketers are just looking at another score built to create urgency without enough, or even any, context behind it to act on.

