blog

AI prompt tracking: monitoring brand mentions in AI search

Knowing where you ranked in Google used to answer most of the visibility question. Rank third for a keyword and you could track the position, see who sat above you, measure the traffic it produced, and work on moving up.

AI search is less tidy. Ask ChatGPT for the best analytics platforms for a European SaaS company and you get five brands. Ask Gemini the same thing and the list shifts. Perplexity might name three of them and cite entirely different sources. Claude might describe your company accurately and never recommend it. Ask again a week later and the answer has moved.

That is why prompt tracking is becoming part of AI search monitoring. Instead of watching where a page ranks for a keyword, you watch what AI systems actually say when people ask the questions that matter to your business.

What AI prompt tracking is

Prompt tracking means running a defined set of questions across the AI platforms on a schedule and recording what comes back. For a project management company, the set might include:

What are the best project management tools for remote teams?

Which project management software is best for a 50-person SaaS company?

What are good alternatives to Asana?

Compare Asana, Monday.com and ClickUp for an engineering team.

The answer itself is only half of it. You want to know whether your brand appears, which competitors sit beside it, how prominently you are named, what the model says about you, and which sources it leans on. Do that on repeat and a picture of your AI visibility builds up. Opening ChatGPT once, typing your company name, and screenshotting the reply tells you almost none of that.

Knowing you is not recommending you

Ask an AI a direct question:

What is Acme Analytics?

If Acme has any real web presence, the model can probably explain what it does. That tells you the AI knows the brand. Now ask the open question:

What are the best privacy-first analytics platforms for European companies?

If Acme drops out of that answer, you have learned something more useful. The first prompt measures recognition. The second measures discovery.

Two-panel diagram: the prompt "What is Acme Analytics?" tests recognition (the AI can describe you), while "Best privacy-first analytics for European teams?" tests discovery (whether the AI names you unprompted). Discovery is the one that wins new customers.

Recognition asks whether the AI knows you. Discovery asks whether it brings you up on its own. New customers arrive through the second one.

Discovery is where brand mentions in AI search get interesting for a marketing team. You want to know whether you appear when someone researches a problem, weighs options, hunts for alternatives, or builds a shortlist without already knowing your name. Those are the prompts where a stranger can find you.

It looks like rank tracking, but it is not

Calling prompt tracking "rank tracking for AI" is a fine starting point and not quite right. Google returns a ranked list. An AI returns an answer. That answer might recommend three companies evenly, or push one hard and mention four others in passing. You might appear first and be described badly. A competitor might appear lower and get the stronger endorsement. Assigning positions 1, 2, and 3 throws most of that away.

The questions worth asking are closer to these:

  • Were we mentioned?
  • Were we recommended, or just listed?
  • Which competitors appeared?
  • How were we described?
  • Which sources were cited, and was one of them ours?
  • Did the answer change since last time?
  • Which engines include us consistently?

That reads a lot richer than a single ranking number.

The same prompt, five different answers

ChatGPT, Gemini, Perplexity, Claude, and Grok are not one search engine wearing five faces. They run different models, different retrieval systems, different indexes, and different ways of choosing sources and building a reply. You can be strong in one and nearly invisible in another.

Diagram: one prompt fans out to ChatGPT, Gemini, Perplexity, Claude and Grok, and each returns a different result, some naming the brand, some citing a third-party review, some omitting it. Visibility exists separately in each engine.

One prompt, five engines, five answers. Your visibility lives separately in each, so a single check tells you about one of them.

Picture tracking this one:

What are the best consent management platforms for European businesses?

ChatGPT names five companies. Gemini names a different five. Perplexity includes you but cites an independent review instead of your site. Claude only mentions you once the user adds a requirement. Grok lands somewhere else again. There is no single ranking underneath all of that. Voris tracks prompts across all five engines for exactly this reason: one model shows you one version of the market, and several together show where you actually stand.

Which prompts to track

This matters more than how many. A list of 500 weak prompts rarely beats 30 chosen well. Start from the questions a real buyer asks while researching your category, and let them narrow the way a real conversation does.

Some stay broad:

What are the best email marketing platforms?

Some get specific:

What are the best email marketing platforms for B2B SaaS?

Some move toward a decision:

Which email marketing platform is best for a 20-person SaaS company?

Competitor and comparison prompts belong in the set too:

What are the best alternatives to Mailchimp?

Mailchimp vs Brevo vs ActiveCampaign for a European SaaS company.

Wording matters, because AI search is conversational. Buyers add the context that keyword research strips out: company size, geography, industry, budget, technical and privacy requirements, integrations, the use case. Any one of those can change the answer.

Track problems, not just categories

One group of prompts is easy to forget. People do not always know what to search for. Someone might type:

Why is traffic from ChatGPT showing up as direct traffic in my analytics?

They are not looking for an "AI visibility analytics platform." They are describing a problem, and the category comes later in the conversation, if at all. For B2B especially, these problem-led prompts are worth watching, because they show whether you enter the conversation before the buyer has even decided what kind of product they need. That is a different, earlier kind of visibility than making a "best tools" list.

Mentions and citations are not the same

An AI can name your brand without citing your site, and cite your content without recommending your product. Say Perplexity answers a question about AI referral traffic using research from your blog: you earned a citation. Say ChatGPT recommends you among the best AI analytics tools but links to nothing: you earned a mention.

Both count, for different reasons. Citation tracking tells you whether your content is becoming part of the material AI systems build answers from. Mention tracking tells you whether your company is becoming part of the set they choose from. You want to read both.

Competitor visibility beats your own score

Say your brand shows up in 37% of tracked prompts. Is that good? On its own, there is no way to tell. If your biggest competitor sits at 12%, you are doing well. If they sit at 81%, you have a different problem entirely.

Monitoring gets far more useful once competitors are in the set. Then you can ask the questions that lead somewhere: which competitors appear when you do not, which prompts they consistently own, whether they win on their own content or on third-party mentions, whether some engines favour them, and whether you show up for informational questions but vanish from the commercial ones. Each gap is something concrete to go investigate.

Visibility moves over time

A single check tells you what an AI said in that moment. It says nothing about whether things are improving. Answers shift as models update, retrieval changes, new content lands, old content ages out, and competitors build authority. The trend is worth more than the screenshot.

Say you publish a benchmark report in September. Over the next few weeks your brand starts turning up for six prompts tied to that research. Perplexity begins citing the report. ChatGPT starts folding you into related recommendations. Gemini still does not. Maybe the report did it, maybe the coverage of the report did, maybe something else moved at the same time. You cannot claim cause from this alone, but you have a signal worth chasing, and without a history of tracking you would never have seen the pattern at all.

What the data is actually for

Tracking for its own sake does little. The value is in the gaps you can act on. When competitors appear for a prompt and you do not, read the answer: what is the AI drawing on, which pages does it cite, what claims are attached to your competitors, which topic do you barely cover. Sometimes the fix is content. Sometimes it is documentation, or a clearer comparison page, or stronger third-party coverage, or original research, or more independent reviews.

Sometimes there is nothing obvious to fix. AI search is unpredictable enough that not every change has a clean explanation, and good monitoring shows what happened without pretending to know why every time.

Connect it to what happens on your site

Getting mentioned is useful. Getting cited is useful. What you eventually want to know is whether any of it turns into business. Someone might meet your company in ChatGPT, click through from Perplexity two days later, read three pages, come back direct the next week, and sign up. The AI mention and the signup are the same customer journey, and most analytics tools only see the last steps.

Four-stage pipeline: an AI citation (coral) leads to a visit (violet), then on-site behaviour (sky), then a conversion (teal). Joining AI visibility, citations, and referral traffic to web analytics turns a mention count into a journey.

A mention only earns its keep when you can follow it to a visit and a conversion. Voris joins the citation, the referral, and the on-site behaviour into one journey.

Joining AI visibility, citation tracking, AI referral traffic, and web analytics is what moves you from "AI mentioned us 184 times" to "these are the topics where AI recommends us, these are the assistants sending people over, and this is what those people do once they arrive." The second one reads like a business metric.

How often to monitor

There is no universal cadence. A fast-moving market rewards daily tracking; a stable B2B category is usually fine on weekly. Consistency matters more than frequency. Running 50 prompts today and a different 50 next month makes comparison hard. A stable core set gives you a baseline, and you add prompts as products, competitors, and language change.

Grouping prompts by intent helps too. Informational prompts show whether your expertise is visible. Category prompts show whether you are discovered. Comparison prompts show how you sit against competitors. High-intent prompts show whether you reach the answers closest to a decision. A single visibility percentage spread across all of those hides more than it tells you.

Treat it as trend data, not truth

New analytics categories tempt everyone to make the numbers look more exact than they are. AI search does not cooperate. Answers vary, models change, retrieval changes, location and context nudge the reply, and the same prompt can read differently on the next run. That does not make prompt tracking useless. It makes it trend data. One appearance today and gone tomorrow is noise. Steady presence across several relevant prompts, several engines, and several weeks starts to look like visibility.

From rankings to conversations

Search visibility used to be easy to picture: a results page, your site somewhere on it. AI search replaces the page with a conversation, and in that conversation your company can be recommended, compared, criticised, cited, left out, or introduced to someone who had never heard of you, all without a single blue link on screen.

So the question stops being "where do we rank?" and becomes "what does AI say about us while customers are deciding, and what does it say when we are not there to watch?" Prompt tracking is how you get to read the answer.

Anna van Bergeijk, Head of Brand. Writes the blog and reads the replies.

read next

ChatGPT vs Perplexity for B2B discovery

More B2B buyers now start in a chatbot than on Google. ChatGPT and Perplexity surface your brand in very different ways, and only one shows its sources.