AI visibility tracking measures how often and how accurately your brand appears in AI-generated answers, using a consistent set of buyer-intent questions run across platforms like ChatGPT, Perplexity, Claude, and Google AI Overviews.
Many marketing teams handle this inconsistently: a one-time test here, a spot check there, no baseline, no trend data. The result is that when AI visibility is raised in a quarterly review, no one can say whether things have improved, declined, or changed by platform. That leaves a measurement gap, and this framework is designed to help close it.
The process takes about 30 minutes a month and a spreadsheet. No dedicated tool required to start.
What AI Visibility Tracking Measures
Before tracking, you need a shared definition of what you’re measuring. Three metrics form the core of any AI visibility tracking program.
| Metric | Definition | What it tells you |
|---|---|---|
| Citation rate | The percentage of queries in your set where your brand name appears anywhere in the AI-generated answer | How consistently AI systems include you in relevant responses |
| Citation position | Where in the answer your brand appears (first mention, second, later) | How prominent your inclusion is when you do appear |
| Citation accuracy | Whether the AI’s description of your brand is correct, current, and differentiated | Whether the citation is working in your favor |
A note on terminology: citation rate in this framework means brand mentions — your company name appearing in the body of the AI-generated answer. This is distinct from sourced citations, where a platform like Perplexity provides a hyperlinked reference to a specific page. Both matter, but they’re measuring different things. This framework tracks mentions.
All three metrics are worth tracking. A high citation rate with consistently low position and inaccurate descriptions is a different problem from a low citation rate with high accuracy when you do appear. The combination tells you more than any single number.
How AI citations are defined and measured →
Which Platforms to Track
Four platforms many teams start with: ChatGPT, Perplexity, Claude, and Google AI Overviews.
Running the same queries across all four gives you a platform breakdown that tells you where to focus. In our experience running audits, the results often diverge significantly, and the reasons behind those divergences are where the useful signal sits.
For example: we’ve seen brands appear in the majority of ChatGPT responses but rarely in Perplexity for the same queries. The underlying cause was usually content freshness. Perplexity often performs live web retrieval, so it tends to surface recently indexed, well-structured content more readily than ChatGPT does for some query types. A brand with strong entity signals but stale or thinly structured content often shows this pattern. That’s a concrete finding with a concrete fix. Without the platform breakdown, you’d see an aggregate citation rate and miss the story entirely.
Each platform has different tendencies. These are observations, not rules, and all four update their behavior regularly:
- ChatGPT tends to rely more on training data and established entity signals for many queries
- Perplexity often uses live retrieval, making freshly indexed content more influential in some cases
- Claude tends to respond well to clearly organized content and credible third-party sources
- Google AI Overviews may surface content structured for clear answers, and FAQPage schema can help with discoverability in some contexts, though structured data supports rather than guarantees visibility
How to Build Your Query Set
Your query set is a fixed list of buyer-intent questions you run against every platform every month. Keep the wording fixed from month to month so results stay comparable.
Ten questions is a practical starting point. Enough to generate meaningful signal, manageable enough to run consistently.
Write questions the way a buyer would ask them: conversational, specific to your category. Not “procurement software” but “what are the best tools for managing supplier contracts at a mid-size company?” Not “marketing automation” but “how do I choose a marketing automation platform for a team of ten?”
Here’s an example query set for a hypothetical procurement consultancy:
- What should I look for when comparing procurement consulting firms?
- How do I choose a procurement advisory partner for my organization?
- What compliance risks should I address when working with external consultants?
- How much does a full procurement transformation project typically cost?
- How long does a procurement advisory engagement usually take?
- What ROI can a CPO expect from procurement consulting?
- How do procurement consulting solutions integrate with existing ERP systems?
- What are the most common use cases where procurement advisory delivers value?
- How do I evaluate whether a procurement consultant has deep category expertise?
- What questions should I ask during vendor selection for advisory services?
Store the query set somewhere permanent. Don’t update or paraphrase mid-cycle.
The Monthly Tracking Process
Run this process on the same day each month, before making any changes to your website or external profiles. If you update something and immediately re-run the test, any movement is ambiguous — you won’t know what caused it. Run first, then make changes, then wait a full cycle.
Step 1: Open a fresh session on each platform.
On ChatGPT, start a new chat. On Perplexity, open a new search. On Claude, start a new conversation. On Google, use an incognito browser. You want clean sessions with no prior context that could shape the responses.
Step 2: Run each query on each platform.
Type each question exactly as written. Don’t rephrase or add context. Record the full response: copy it into a document or take a screenshot.
Step 3: Score each response.
For each query on each platform, record:
- Was your brand mentioned? (Yes / No)
- If yes, what position? (First mention, second, third, later in the response)
- If yes, is the description accurate? (Accurate / Partially accurate / Inaccurate)
- Which competitors were mentioned?
Step 4: Calculate your monthly metrics.
Citation rate: your “Yes” count divided by total queries, per platform. Competitor citation rate: the same calculation for each competitor you’re benchmarking against.
Step 5: Log to your tracking sheet.
A simple monthly sheet works well, with one row per platform per month. Columns: date, platform, citation rate, average position, accuracy score, competitor citation rates, notes. The notes column is where you log what changed that month — new content published, profiles updated, schema implemented.
After three months you’ll start to see patterns. After six months you’ll usually have enough history to spot useful trends.
A note on methodology
Results vary between sessions, even for identical queries. AI platform outputs are not deterministic. The same question can produce different responses in different sessions, and model updates can shift behavior in ways that aren’t announced. This is why consistency matters: same questions, same day each month, fresh sessions. You’re tracking trend direction over time, not treating any single result as exact.
What to Do with the Data
When citation rate is low or declining, we recommend investigating in this order: entity signals first, schema second, content structure third. This is our prioritization framework based on what we’ve observed across audits, not a universal rule, and your situation may differ.
If you’re not appearing at all, the most common issue is that AI systems don’t have enough independent information to describe your brand with confidence. Check whether your external profiles are complete and consistent: Google Business Profile, Crunchbase, G2, Capterra, LinkedIn. For ChatGPT specifically, Wikidata presence can be influential, though this isn’t a straightforward recommendation for every brand, since Wikipedia eligibility has its own requirements and editorial process.
If you’re appearing inconsistently across platforms, the breakdown tells you where to focus. A low score on Google AI Overviews often points to structured data or content organization issues. A low score on Perplexity often points to content freshness or third-party source presence.
Share of AI Voice is the competitive layer: your citation count divided by the combined citation count of all tracked brands in the same query set, expressed as a percentage. If your brand appears in 6 out of 10 queries and three competitors appear in 8, 7, and 5 respectively, the total citations are 26. Your Share of AI Voice is 6/26 = 23%. The denominator changes if you change the competitor set, so keep that fixed too.
How Share of AI Voice is calculated →
Entity signals to investigate when citation rate is low →
Manual Tracking vs. Dedicated Tools
The framework above is manual. It works, and it’s where most teams should start.
Dedicated AI visibility tracking tools automate the query runs, often daily, and surface trend data without manual effort. They’re worth considering when you’re tracking a large query set, multiple markets, or when you need to catch changes quickly rather than monthly.
The trade-off is cost and setup time. A manual tracking sheet costs nothing and takes an afternoon to build. A dedicated tool costs anywhere from $30 to several hundred dollars a month depending on query volume and feature set. For teams that have established a baseline manually and want to scale the practice, the upgrade makes sense. For teams still figuring out what they’re measuring, starting manually is usually the right move.