To measure your AI search visibility, you need to track three primary metrics: citation frequency (how often AI systems mention your brand), citation accuracy (whether AI says correct things about you), and citation sentiment (whether AI recommends you positively, neutrally, or negatively). Unlike traditional SEO where Google Search Console provides clear data, AI search measurement requires a combination of manual testing, specialist tracking tools, and systematic prompt monitoring across ChatGPT, Perplexity, Google AI Overviews, and Claude. This guide covers the metrics, tools, and routines you need.
Primary Metrics for AI Search Visibility
| Metric | What It Measures | How to Track | Target |
|---|---|---|---|
| Citation frequency | How often your brand appears in AI answers | Automated tools + manual sampling | Increasing month-on-month |
| Citation accuracy | Whether AI states correct facts about you | Manual review of AI outputs | 95%+ accuracy |
| Citation sentiment | How positively AI frames your brand | Manual review + sentiment analysis | Positive or neutral |
| Query coverage | % of target queries where you appear | Track a defined set of prompts | 30%+ for initial goals |
| Competitor share | Your citations vs competitor citations | Comparative tracking across prompts | Greater than or equal to competitors |
| Platform coverage | Which AI platforms cite you | Cross-platform monitoring | Present on 3+ platforms |
The Challenge: AI Search Measurement Is Not Like SEO Measurement
Traditional SEO measurement is straightforward. Google Search Console tells you your rankings, impressions, and clicks. Google Analytics shows you traffic and conversions. The data is reliable and consistent.
AI search measurement is fundamentally different:
No central analytics platform. There is no “AI Search Console” that shows you where you appear across ChatGPT, Perplexity, and Google AI Overviews.
Non-deterministic results. The same prompt can produce different results each time. AI outputs vary based on conversation history, model version, and even time of day.
Multiple platforms. You need to monitor ChatGPT, Perplexity, Google AI Overviews, and Claude separately. Each uses different data sources and evaluation criteria.
No click-through data. When an AI system cites your business, there is often no referral link that shows up in your analytics. The citation itself is the value, because it shapes the user’s perception before they visit any website.
Tools for Measuring AI Search Visibility
Specialist GEO Tracking Tools
| Tool | Platforms Tracked | Key Features | Approximate Price |
|---|---|---|---|
| Peec AI | ChatGPT, Perplexity, Google AI Overviews | Brand mention tracking, competitor comparison, sentiment analysis | From £200/month |
| Otterly | ChatGPT, Perplexity, Google AI Overviews | Prompt monitoring, citation tracking, trend analysis | From £150/month |
| Profound | ChatGPT, Perplexity, Google AI Overviews, Claude | Deep AI citation analysis, entity mapping | Custom pricing |
SEO Platforms With AI Features
| Tool | AI Feature | Limitation |
|---|---|---|
| Semrush | AI Overview tracking | Google AI Overviews only, so it does not track ChatGPT or Perplexity |
| Ahrefs Brand Radar | Brand mention monitoring | Monitors traditional web mentions, not AI-specific citations |
| Moz | AI Overview visibility | Limited to Google AI Overviews |
Manual monitoring, which no tool replaces
No tool captures everything, so manual monitoring still matters:
- ChatGPT: Test 10-20 target prompts weekly. Use new conversation sessions each time.
- Perplexity: Test the same prompts and note which sources are cited.
- Google: Search target queries and check for AI Overview presence.
- Claude: Test key prompts to check for brand mentions and accuracy.
Setting Up Your Monitoring Routine
Weekly Monitoring (30-45 minutes)
Prompt testing (20 minutes). Test your top 10 target prompts across ChatGPT, Perplexity, and Google. Record whether you appear, what is said, and whether it is accurate.
Competitor check (10 minutes). Note which competitors appear for the same prompts. Track any changes from previous weeks.
Accuracy review (5 minutes). Flag any inaccurate information for entity correction.
Log results. Record everything in a tracking spreadsheet or your GEO tool’s dashboard.
Monthly Analysis (1-2 hours)
Citation trend analysis. Are your citations increasing, decreasing, or stable month-on-month?
Query gap analysis. Which target prompts still do not return your business? What entity signals might be missing?
Competitor movement. Have any competitors gained or lost AI visibility?
Accuracy audit. Review all AI outputs about your business for accuracy and consistency.
Content impact assessment. Which content changes correlated with increased citations?
Quarterly Strategic Review (Half day)
- Full entity audit. Re-audit all entity signals across the web.
- Strategy adjustment. Based on 3 months of data, adjust your GEO priorities.
- Target prompt expansion. Add new prompts to your monitoring set based on emerging client queries.
- ROI assessment. Correlate AI visibility changes with business outcomes (enquiries, leads, revenue).
What movement actually looks like
Most benchmark tables for this are invented. Here are the figures from our own published case studies instead, with the query set size, because a citation share means nothing without knowing what it is a share of.
| Engagement | Query set | Citation share | Elapsed |
|---|---|---|---|
| Midlands IFA practice | 45 | 0% to 31% | 90 days |
| Commercial law firm | 75 | 4% to 41% | 12 months |
| B2B SaaS | 180 | 9% to 38% | 14 months |
Two things worth drawing out. The law firm moved 4% to 12% to 26% to 41% across four quarters, and the first three months produced almost no citation movement at all because they went on entity foundations. That flat opening is normal and it is where most programmes get abandoned.
And citation share is not linear with effort. A 180-query set moving 29 points is a bigger piece of work than a 45-query set moving 31.
The four checks almost nobody runs
Everything above is the AI side. This is the analytics side, and it is where we find most of the measurement errors, including our own. Every figure here comes from a twelve month pass across 56 of our own analytics properties.
1. Split search traffic by engine, not by channel
GA4’s Traffic acquisition report defaults to channel, and the Organic Search channel bundles Google, Bing, Yahoo, DuckDuckGo and Ecosia into one row. So the default view of the only tool that can show you the split actively hides it. Change the dimension to session source.
What it showed across our estate over 90 days, with datacentre sessions removed:
| Engine | Sessions | Conversions | Rate |
|---|---|---|---|
| 2,184 | 168 | 7.7% | |
| bing | 1,025 | 64 | 6.2% |
| yahoo | 295 | 13 | 4.4% |
| ai assistants | 223 | 10 | 4.5% |
| duckduckgo | 184 | 30 | 16.3% |
| ecosia | 68 | 5 | 7.4% |
DuckDuckGo converted at more than twice Google’s rate, and Bing delivered 47% of Google’s session volume. On one individual property the split was far more extreme: 830 search sessions, 91 enquiries, and Google accounted for 35 sessions and 2 of the enquiries.
Three of our properties are majority non-Google and the rest are Google-dominated. You cannot predict which from the outside, which is the argument for measuring rather than assuming.
2. Track distinct pages earning impressions, not total impressions
This is the most useful metric on this page and the one least likely to be on your dashboard.
Total impressions can hold steady while the number of pages producing them collapses, because the surviving pages absorb the queries. On one of our properties:
| Window | Pages earning any impression | Impressions |
|---|---|---|
| 15 to 28 June | 2,648 | 19,051 |
| 1 to 14 July | 360 | 1,753 |
| 1 to 14 August | 101 | 538 |
| 15 to 28 August | 7 | 179 |
Weekly impressions went 8,912 to 812 in a single week. Put a 28-day rolling average over that and the cliff becomes a slope, which is precisely what a rolling average is built to do. On a large template-driven site, site-wide impression totals are close to useless as a health metric.
Page count moves straight away. Put it next to clicks.
3. Check whether your events are actually counting
An event can appear in the GA4 events list, fire correctly every time, and never reach a conversion report, because reaching one requires it to be marked as a key event separately.
On one of our directories that gap was worth 180 phone clicks over twelve months. On a directory, somebody tapping a practice’s phone number is the enquiry. Every one fired, sat in the events list, and reached no report anybody reads. Its sister site, same stack and same build, had it marked, so the two were never comparable.
Across our estate, 8 of 33 live properties have any key event configured at all. One has 17,122 sessions a quarter and none.
Open GA4 admin, list what you count as a conversion, then check that list against everything the site can actually do.
4. Take the datacentre traffic out before you calculate anything
Across 56 properties, 4,127 of 15,796 sessions came from cloud regions: Council Bluffs, Boardman, Ashburn, Glenview, The Dalles and Singapore. That is 26% of everything, and closer to half on two of our directories. Four and five second average engagement, effectively no conversions.
GA4 excludes known bots by default, but that list is the IAB one and it does not catch these.
Look at Direct traffic broken down by city, with average engagement time beside it. If somewhere you have never sold to is sending four second sessions, that is your answer. Strip them out and one of our directories stopped looking mediocre and started converting organic traffic at 14.2%.
A fifth, if you deploy often: IndexNow does not reach Google
IndexNow covers Bing, Yandex, Seznam and Naver. Google has never adopted it. A lot of deploy pipelines now ping IndexNow and treat indexing as handled.
On one of our sites that meant every deploy notified Bing and nothing notified Google. A round of title and description rewrites shipped in July, Bing had them within days, and Google has still not seen them on a single page. If IndexNow is the whole of your indexing strategy, you have a Bing strategy.
Why average position will mislead you
Worth its own warning, because it is on the front page of every Search Console report.
On one of our directories, Google clicks went 148 to 453 across the 90 days either side of one day’s structural work. Average position over the same window went 25.7 to 28.7, in the wrong direction. Both numbers are right.
The site had gone from 1,121 pages earning impressions to 1,666. Every new page entered the index low, because new pages always do, and dragged the mean down while the established pages improved and took the clicks.
Average position is an unweighted mean across every query you appear for. A site that gets broader will almost always see it fall. A site that prunes its weakest pages will see it rise while earning less. Use clicks, click-through rate, and the count of pages earning any impression. The full numbers are in the case study.
Connecting AI Visibility to Business Outcomes
The ultimate measure of GEO success is business impact. Track these downstream metrics:
Branded search volume. As AI systems mention your brand more frequently, branded searches in Google should increase.
Direct website traffic. Users who encounter your brand in AI answers may navigate directly to your website.
Enquiry source. Ask new prospects how they found you. “I asked ChatGPT” or “I saw you in an AI search” are increasingly common responses.
Lead quality. AI-referred leads often have higher intent because the AI has already named your business as the answer to their specific problem.
How we run this
Measurement sits inside the Synaptic Authority Engine as one of the 13 pillars, and the reporting covers citation frequency, accuracy, sentiment and competitor share across ChatGPT, Perplexity, Google AI Overviews and Claude, alongside the four analytics checks above.
The reason those four checks are on this page at all is that we found every one of them on our own properties first. We publish the numbers from our own estate rather than anonymised client data, because our own sites are the ones we can show you the exports for.
For the tooling side of this, we compared the options in the best GEO tracking tools for regulated sectors.
A tool measures. It does not move the number. Peec AI versus a managed agency covers where that line sits.
Frequently Asked Questions
Is there a free way to check my AI search visibility? Yes, manual testing. Search for your business and service terms in ChatGPT (free tier), Perplexity (free tier), and Google (check for AI Overviews). This takes time but costs nothing.
How often should I check my AI visibility? Weekly prompt testing and monthly detailed analysis is the recommended minimum. Quarterly strategic reviews ensure your approach stays aligned with evolving AI search behaviour.
Which matters more, appearing in ChatGPT or Google AI Overviews? Both matter. Google AI Overviews currently reach more users because they appear in regular Google searches. ChatGPT and Perplexity reach users who have specifically chosen to use AI for research. The ideal position is visibility across all platforms.
Can AI visibility be gamed? Not sustainably. AI systems are designed to detect and discount manipulative signals. Genuine entity authority, meaning consistent information, verifiable credentials and authentic third-party mentions, is the only reliable path to sustainable AI visibility.
How do I know if an enquiry came from an AI citation? Ask. Add “How did you hear about us?” to your enquiry forms with “AI search (ChatGPT, Perplexity, etc.)” as an option. Many businesses are surprised by how many prospects select this option.
One symptom worth knowing how to read before you start: if your unassigned or direct traffic in GA4 has been climbing month after month with no campaign change behind it, a good part of that is frequently AI referrals the platform cannot attribute, because many assistants strip the referrer. Why is my unassigned traffic growing in GA4? covers how to separate that from the mundane causes, which matters because the mundane causes are more common and cheaper to fix.
Want a professional assessment of your AI search visibility? Book a free AI visibility audit and we will show you exactly where you stand across every AI platform, or see how our AI SEO agency closes the gaps for you.
Related: for why the buying path itself became harder to observe, and what still works to measure it, see AI-powered customer journeys. For the sequencing decisions these metrics are meant to inform, see AI visibility strategy. For two worked examples where visibility fell and enquiries rose at the same time, with the month-by-month figures, see traffic with intent.