okaeo

Guide

How to track your brand in AI answers

A repeatable method for measuring whether ChatGPT, Gemini, Perplexity and AI Overviews name you, where you sit inside the answer, and which pages they cite

About 9 minutes · Updated 2026-09-18

Why asking ChatGPT yourself does not work

Almost every team starts the same way: someone opens ChatGPT, types the category question, sees a competitor named first, and forwards the screenshot. It is a useful shock, but it is not a measurement, and building a plan on it goes wrong in four specific ways.

  • Answers are not deterministic. Ask the same question twice and you can get two different brand lists. A single sample cannot tell a real change from normal variance.
  • Your account is personalised. Memory, past chats and location shape what you see, so what you observe is not what a new buyer observes.
  • One prompt is not a category. Buyers ask dozens of differently-phrased questions, and your visibility usually varies enormously between them.
  • One engine is not the market. Retrieval stacks differ, so being strong in Perplexity says little about Gemini.

The fix is not more careful manual checking. It is sampling: a fixed set of prompts, run across multiple engines, on a schedule, with the raw answers stored so you can go back and read what actually changed.

Step 1, build a prompt set that mirrors the buying journey

Your prompt set is the instrument. Everything you report later inherits its biases, so it is worth an afternoon of real thought rather than a brainstorm of branded keywords.

Cover four stages, and keep the wording conversational. People type keywords into a search box and full sentences into an AI, so a prompt set that reads like a keyword list is already measuring the wrong thing.

  • Awareness, no brand named: "what tools can show whether AI mentions my company"
  • Consideration, category comparison: "best AEO platforms for a B2B SaaS marketing team"
  • Evaluation, specific constraints: "which AI visibility tools track DeepSeek and Qwen", "typical pricing for AI brand monitoring"
  • Decision, branded: "is Okaeo any good", "Okaeo vs doing it manually"

Forty to sixty prompts per brand is a workable starting point. Below thirty, weekly noise swamps the signal. And once the set is set, freeze it: changing prompts mid-quarter invalidates every trend line you were building.

Step 2, choose engines by where your buyers actually are

Engine choice is a market decision, not a technical one. A European B2B brand and a consumer brand selling into China do not need the same coverage, and tracking engines your buyers never open just adds cost and noise.

  • Almost always: ChatGPT and Google AI Overviews, which together account for the bulk of AI answer exposure in most Western markets.
  • Usually: Gemini, Perplexity and Claude, where research-heavy and technical buyers concentrate.
  • If you sell into enterprises: Copilot, because it sits inside the Microsoft tools those buyers already have open.
  • If you sell into Chinese-speaking markets: DeepSeek, Qwen and Kimi, which behave differently enough that Western results do not transfer.

Step 3, measure five things, not one

"Are we in AI" is not a metric. These five are, and each one leads to a different action.

  • Mention rate: the share of answers that name you at all. This is your floor.
  • Rank inside the answer: named first reads as a recommendation, named fifth reads as a footnote.
  • Citations: which of your URLs the engine actually linked to. This tells you which pages are doing the work.
  • Share of voice: your mention rate against each competitor's, on the same prompts. This is the number leadership will ask for.
  • Sentiment and accuracy: how you are described. A mention that misstates your pricing is a negative, not a win.

Store the full answer text, not just the metrics. Six weeks later, when share of voice drops four points, the only way to explain it is to read the answers from both periods side by side.

Step 4, sample on a schedule and resist daily readings

Daily sampling is right for collection and wrong for interpretation. Answer engines change output day to day for reasons that have nothing to do with you: index refreshes, model updates, retrieval changes. Collect daily, but read weekly averages, and only act on a change that holds for two consecutive weeks.

The exception is a genuine event: a launch, a funding announcement, a competitor's campaign, a bad review going viral. Around those, daily readings are worth watching closely, because a wrong claim can propagate across engines within days.

Step 5, turn the data into three decisions

Most AI visibility dashboards die because nobody knows what to do on Monday morning. Three questions turn the numbers back into work.

  • Where are we absent but our category is being answered? Those prompts are content briefs, and they are usually the highest-value ones you own nothing on.
  • Who is cited in our place, and on what page? Read that page. Nine times out of ten it answers the question more directly than yours does, not better.
  • Where are we mentioned but described wrongly? That is a fact-supply problem on your own site, and it is usually the fastest thing on this list to fix.

Common questions

Collect data daily so you have a dense record, but interpret weekly averages. Answer engines vary day to day for reasons unrelated to your brand, so only act on changes that hold for two consecutive weeks.

Forty to sixty prompts per brand is a workable starting point, spread across awareness, consideration, evaluation and decision stage questions. Below about thirty prompts, week-to-week noise overwhelms the signal.

A manual check is a useful shock but not a measurement. Answers are non-deterministic, your account is personalised by memory and location, and one prompt in one engine cannot represent a category. Reliable tracking needs a fixed prompt set sampled repeatedly across several engines.

Sometimes, but traffic understates the effect. Many AI answers end without a click, so a brand can be described favourably in thousands of answers and see almost none of it in analytics. Measure mentions, rank inside the answer, citations and share of voice directly rather than inferring them from sessions.

Now see it on your own brand