July 28, 2026 · 9 min read

How to Track Whether AI Models Mention Your Brand

ChatGPT and Perplexity recommend products every day. Here is the actual method for measuring whether yours comes up, including what we built and what broke.

Quick answer: You track AI brand mentions by running a fixed set of buyer-intent prompts against each model on a schedule, recording whether your brand appears, in what position, and which sources the model cited. There is no equivalent of Search Console for this, so you either build the pipeline or buy a tool. Below is the method we run, including the parts that took several attempts to get right.


Why this became worth measuring

A meaningful share of product research now happens inside a chat window. Someone evaluating tools asks ChatGPT what the best options are, gets five names, and investigates two of them. If you are not in the five, you were never in the consideration set, and no analytics tool will tell you it happened.

This is the part that makes it uncomfortable. When you lose a Google ranking you can see it. When you lose an AI recommendation you see nothing at all. There is no impression count, no position, no referrer. The traffic simply does not arrive and you have no way to know it existed.

Which is the whole reason to measure it deliberately.


Why this is harder than SEO tracking

Four properties make AI mentions genuinely difficult to measure, and any tool or method claiming precision here is overselling.

Answers are non-deterministic. Ask the same model the same question twice and you can get different brands. A single check tells you almost nothing. You need repetition to see a distribution rather than a sample.

There is no official surface. Google gives you Search Console. OpenAI, Anthropic, and Perplexity give you nothing comparable. Everything is inferred from the outside.

Personalization varies. Memory, custom instructions, and account history all shape answers. Your own account is the worst possible place to test, because the models have been reading your site all year.

Models update without notice. A result from March may not reproduce in July, and nothing announces the change. Trend data matters more than any single measurement.


The method

Step 1: Build a prompt set that reflects real buying behavior

This is the step people rush and it determines everything downstream.

Write 20 to 40 prompts in four groups:

Category prompts. What someone asks before they know any brands.

  • "best AI marketing tools for a solo founder"
  • "how do I do marketing without hiring anyone" Comparison prompts. Mid-evaluation.
  • "Okara vs NoimosAI"
  • "alternatives to hiring a marketing agency" Problem prompts. Where the buyer describes a situation instead of a category.
  • "I have a SaaS product and no time for marketing, what should I use"
  • "how do I get my startup mentioned on Reddit without getting banned" Brand prompts. What the models say when asked about you directly.
  • "what is Okara"
  • "is Okara any good" Problem prompts are the ones most people skip and they matter most. Buyers describe symptoms far more often than they name categories, and the brand that gets recommended in response to a symptom is the one that wins the consideration set.

Brand prompts serve a different purpose. They tell you whether the models understand what you sell at all, which for us turned out to be the more urgent finding.

Step 2: Query the models cleanly

Run every prompt against each model you care about. Currently that means ChatGPT, Claude, Perplexity, Google AI Mode, and Gemini, though the list keeps changing.

Three rules that matter more than they sound:

Use the API, not the app. App sessions carry memory and personalization. API calls with no system prompt are the closest you get to a neutral observer.

Run each prompt several times. Five is a workable floor. You are measuring frequency of appearance, not presence or absence. "We appear in 3 of 5 runs" is real information. "We appeared" is not.

Log the full response, not just a yes or no. The competitors named alongside you and the sources cited are more actionable than your own mention.

Step 3: Record four things per run

FieldWhy it matters
Mentioned (yes/no)The headline number, tracked as a rate across runs
Position in the listBeing fifth of five is close to not being mentioned
Sentiment and framing"Okara is a budget option" and "Okara is the best fit for founders" are different outcomes
Cited sourcesThe most actionable field by a distance

The citation field is where the leverage sits. If a model recommends three competitors and cites the same two listicles each time, you now know exactly which pages to get into. That is a concrete task, unlike "improve our GEO."

Step 4: Automate it

Manual checks work for a first pass and then stop scaling. We built ours on the DataForSEO LLM mentions API running on a schedule, writing to a database, with a dashboard on top showing mention rate over time per prompt per model.

If you would rather not build, several tools now cover this. Profound, Athena HQ, and Scrunch all track LLM visibility, and Semrush includes AI visibility tracking on its higher tiers. Any of them beats not measuring.


What we found, including the part we did not expect

Three things came out of running this on ourselves.

Brand prompts exposed a problem the category prompts hid. We were tracking whether we appeared for "best AI marketing tools." We were not tracking what models said when asked "what is Okara." The answer, it turned out, was that several described us as a private AI chat product with encrypted conversations and open-source model access.

That is not what we sell. It was what we used to sell, and the models had learned it from third-party directory listings that were never updated, plus our own site still carrying legacy pages.

No amount of GEO work on category prompts would have fixed that, because the problem was entity definition rather than visibility. If the models do not know what you are, ranking for your category is not the constraint.

Check your brand prompts before anything else. If the answer to "what is [your product]" is wrong, fix that first. Directory listings, your own legacy content, and schema markup are the levers, and none of them are content marketing.

Cited sources clustered far more than expected. Across dozens of runs, the same handful of roundup articles kept appearing as citations. Vendor pages, including our own, appeared far less often than third-party comparisons. The practical implication is that getting included in other people's listicles moves the needle more than publishing your own.

Consistency across models is low. A brand strong in Perplexity is often invisible in ChatGPT. Perplexity leans on live retrieval, so recent content and clean citation structure help. ChatGPT leans more on training data, so historical presence and volume of third-party mention matter more. These need different work.


What to do with the data

If your mention rate is near zero on category prompts: the constraint is third-party presence, not your own content. Get into the roundups your competitors are cited in. Find them in your own citation logs.

If you are mentioned but positioned poorly: the models have a view of you and it is unflattering or too narrow. Usually this traces to how you are described on comparison sites and directories rather than on your own site.

If your brand prompts return the wrong description: stop all other GEO work and fix the entity. Update directory listings, add unambiguous schema, publish a clean definitional page, and remove or de-index legacy content describing a product you no longer sell. This is the highest-leverage GEO work available and almost nobody does it, because it does not look like marketing.

If you are strong in one model and absent in another: treat them as separate channels. Perplexity rewards fresh, well-cited content. ChatGPT rewards accumulated third-party presence. The same work does not serve both.


Frequently asked questions

How often should I check AI brand mentions? Weekly is enough for most teams. Daily produces noise, since non-determinism means small movements are usually nothing. Monthly is too slow to catch a model update.

Can I use Google Search Console for this? No. Search Console covers Google Search including AI Overviews to a limited extent, but it gives you nothing for ChatGPT, Claude, or Perplexity. Those surfaces have no publisher-facing analytics at all.

Does traffic from AI models show up in analytics? Some does. Perplexity and ChatGPT pass referrer data when a user clicks through, so you can see it in your analytics. What you cannot see is the far larger set of conversations where you were recommended and the user did not click, or where you were not recommended at all.

What is a good AI mention rate? There is no benchmark, and anyone quoting one is guessing. Measure your own trend and your rate relative to named competitors on the same prompts. Relative position is the only meaningful number here.

Do I need a tool or can I do this manually? Manual works for a first assessment. Take 20 prompts, run each 5 times across 3 models, log the results in a spreadsheet. That is 300 queries and an afternoon, and it will tell you more than most GEO audits. Automate once you want a trend rather than a snapshot.

Which models should I track? ChatGPT and Perplexity at minimum, since they carry the most product-research volume. Add Google AI Mode if your category has strong search intent, and Claude if you sell to developers or technical buyers.

Does llms.txt affect whether models mention me? Evidence is thin so far. It is cheap to implement and unlikely to hurt, but treat claims about its impact skeptically until there is better data. Entity clarity and third-party citations demonstrably matter more.