Okara
Back to Blog
By Fatima Rizwan · Published July 4, 2026 · 12 min read

What Is Share of Voice in AI Search (and How to Improve Yours)

Learn how to calculate response, mention, and citation share across AI-generated answers and build a repeatable measurement process you can trust

Share of voice in AI search describes how often a brand appears in generated answers for a defined set of prompts. It can help you see where a brand is present, absent, mentioned, or cited. But the phrase does not have one standard formula.

That caveat matters. A tool might call brand mentions divided by competitor mentions "share of voice." Another might report the percentage of responses that mention the brand. A third might blend mentions, citations, position, and sentiment into a proprietary score. Those numbers are not interchangeable.

Useful measurement starts by naming the metric, fixing the denominator, and recording how the answers were collected.

Three measures that should not be conflated

Before calculating anything, define the scope: platforms, models, prompts, competitors, countries, languages, dates, and number of repeat runs. Then choose a measure that answers the question you care about.

Response share

Response share measures presence across answer runs:

response share = valid response runs that mention the brand / all valid response runs in scope

Suppose you test 40 prompts three times on one platform. That creates 120 response runs. If your brand appears in 30, its response share is 25%.

The denominator is response runs, not prompts and not total brand mentions. A brand mentioned five times in one answer still counts once for that run. This measure answers, "In what portion of the sampled answers did the brand appear?"

State how you handle refusals, errors, blank answers, and answers that do not address the prompt. Excluding them without a written rule can move the result.

Mention share

Mention share compares a brand with a specified competitor set:

mention share = the brand's mention units / mention units for all tracked brands

The counting unit needs its own definition. A sensible default is one unit per brand per response. If an answer mentions Acme four times and Bravo once, each gets one mention unit. This avoids rewarding repetitive wording. If you count every textual occurrence instead, say so.

Imagine the sampled answers contain 100 brand-response mention units across the tracked set, and 30 belong to your brand. Your mention share is 30%. Change the competitor list and the denominator changes, so comparisons are only valid when that list stays fixed.

This measure answers, "Of the tracked brands that appeared, what portion of the appearances belonged to us?" It does not tell you how often no tracked brand appeared.

Citation share

Citation share measures linked sourcing:

citation share = citations to the brand's owned domains / all external citations in scope

If the responses contain 80 external citations and 10 point to your domains, your citation share is 12.5%. Set rules for duplicate links, repeated footnotes, subdomains, syndicated pages, and non-clickable source names before collecting data.

A related metric, citation response rate, uses response runs as the denominator:

citation response rate = valid response runs citing the brand / all valid response runs

Citation share and citation response rate answer different questions. Keep both if links matter to your reporting.

None of these formulas is the universal definition of AI share of voice. They are transparent options. Pick the one that fits the decision, publish the denominator beside the result, and resist combining unlike measures into a single unexplained percentage.

Build a prompt sample you can defend

A random collection of prompts is easy to assemble and hard to interpret. Start with the questions customers ask during research, comparison, and purchase. Support tickets, sales calls, internal site search, and Search Console query data can all help you build the list.

Group prompts by intent so one large cluster does not dominate the result. A software company might separate problem-solving questions, category research, comparisons, and branded questions. Record the number of prompts in each group and decide whether each group receives equal weight or is weighted by a documented business reason.

Keep the wording stable. "Best accounting software for a freelancer in Canada" is a different prompt from "What bookkeeping app should I use?" Both may belong in the sample, but substituting one for the other breaks the comparison.

Branded prompts need special care. A model mentioning your company after being asked about it says little about competitive visibility. Report branded and unbranded results separately.

Control the conditions of every run

Generated answers can change even when the prompt does not. Providers update models, retrieval systems choose different sources, and account or location signals can alter the output. One answer is an observation, not a stable ranking.

For each run, record:

  • The exact prompt, including punctuation and any system instructions
  • The provider, product, model or mode, and whether web search was enabled
  • The date and time
  • Country, language, and location settings
  • Whether the session was logged in and whether memory or personalization was active
  • Whether the prompt ran in a fresh conversation
  • The full answer, named brands, and cited URLs

Use the same setup each time you compare periods. If a provider retires a model or changes a mode, note the break in the series rather than treating the next result as directly comparable.

Location and account controls are not administrative trivia. A recommendation for a logged-in user in London may differ from one for a signed-out user in Toronto. If those audiences matter, treat them as separate samples.

Repeat runs and report uncertainty

Run each prompt more than once under the same recorded conditions. Repeats reveal how often a brand moves in and out of the answer. They do not remove variance, but they make it visible.

The number of repeats should reflect the decision at stake and the resources available. Publish the count instead of claiming that a particular number is always sufficient. A small sample can be useful for finding obvious gaps, but it cannot support a precise claim about the whole market.

When reporting a proportion, include its numerator and denominator: "30 of 120 response runs, or 25%." For larger studies, add an appropriate confidence interval and describe the sampling design. Closely related prompts and repeated runs may not be independent, so a basic binomial interval can understate uncertainty. If the analysis will guide substantial spending, involve someone who can account for clustering and weighting.

Avoid treating a small month-to-month movement as proof that an edit worked. The difference may come from ordinary sampling variation, a model update, a source refresh, or changed test conditions. Look for sustained movement across repeated samples, then check the underlying answers and citations.

How to read monitoring tool scores

AI visibility tools save time by running prompts, extracting brand names, and storing citations. Their dashboards can be useful, especially when the alternative is an inconsistent spreadsheet.

The headline scores are usually proprietary metrics, though. Vendors may use different prompt sets, models, run frequencies, competitor lists, weighting, and definitions of a mention. Some combine citation frequency, answer position, sentiment, or other signals. Two products can assign different scores to the same brand without either calculation being wrong.

Before relying on a score, ask:

  • What is the numerator and denominator?
  • Which prompts, models, locations, and dates are included?
  • How many repeat runs are made?
  • Are mentions counted once per answer or every time the name appears?
  • How are citations, errors, and missing answers handled?
  • Can the raw answers and cited URLs be exported?
  • Did the methodology change during the reporting period?

Use a proprietary score to track a consistent panel within that tool. Do not present it as an industry-wide measure or compare it directly with a differently defined score.

What can improve visibility

Share-of-voice monitoring is most useful when it leads back to the source material. Review the answers where competitors appear and inspect the pages that were cited. You may find that your site does not address the question, gives an outdated answer, or makes claims without evidence.

Work on the gaps you can verify:

  1. Publish accurate pages for questions your customers actually ask.
  2. Put the direct answer near the relevant heading instead of burying it in an introduction.
  3. Support factual claims with first-party documents, original research, or clearly attributed evidence.
  4. Keep product details, pricing, and limitations current.
  5. Make important content available as crawlable text.
  6. Earn legitimate coverage by giving publishers, customers, and researchers information worth discussing.

Third-party pages can be useful evidence about a company, but no particular directory, review site, or forum is a guaranteed visibility lever. Create or maintain a profile when it helps customers and accurately represents the product. Do not buy fabricated reviews or seed disguised endorsements.

Structured data has a narrower role than some AI search advice suggests. Use supported schema when it accurately describes content visible on the page and qualifies the page for a relevant search feature. Google says AI Overviews and AI Mode have no additional technical requirements and need no special schema. Adding FAQPage markup does not by itself make an answer system mention or cite a page.

Crawler access: search is not training

Crawler controls differ by provider and by purpose. Search crawlers discover material for search experiences. Training bots collect or control content for model development. User-triggered agents may retrieve a page only after someone asks for it. Grouping all of them under "AI crawler" leads to bad robots.txt advice.

For Google Search, Googlebot governs crawling for Search, including pages that may appear in AI Overviews and AI Mode. A page must be indexed and eligible to appear in Search with a snippet. Google-Extended is a separate control for certain Gemini training and grounding uses described by Google. It does not control inclusion or ranking in Google Search, AI Overviews, or AI Mode.

OpenAI documents OAI-SearchBot for search, GPTBot for potential model training, and ChatGPT-User for user-initiated actions. Perplexity likewise documents PerplexityBot for search and Perplexity-User for user requests. Anthropic publishes separate controls for Claude-SearchBot, ClaudeBot, and Claude-User.

Allowing one bot does not require allowing another. Decide separately whether you want search discovery, user-requested access, and model training. Then check the provider's current documentation, robots.txt, and any CDN or firewall rules. Bot names and policies can change.

Crawler access only makes retrieval possible. It does not guarantee indexing, a mention, a citation, or a favorable description.

Use the data as a diagnostic

The useful part of this exercise is rarely the headline percentage. The answer archive can show which customer questions you have not covered, which claims need better support, and where outdated information is being repeated.

Compare results within stable prompt groups. Review the cited sources. Check whether the answer is accurate, not merely whether your name appears. A visible but incorrect recommendation can be worse than no mention at all.

Connect this monitoring with ordinary business measures such as qualified visits, assisted conversions, sales conversations, and customer research. AI share of voice measures sampled answer visibility. It does not prove reach, persuasion, or revenue.

Where Okara fits

Disclosure: Okara is our product.

Okara can run a defined prompt set across supported AI answer products and track mentions and citations over time. Its scores, like those from other monitoring vendors, are proprietary measures rather than universal standards. Review the underlying responses and methodology before using any score in a decision.

Frequently asked questions

It is a family of measures for brand visibility in sampled AI-generated answers. Depending on the method, it may refer to the share of response runs containing a brand, the brand's portion of tracked mentions, or its portion of citations. The definition and denominator should appear with the number.

Is share of voice the same as a GEO score?

Not necessarily. A share-of-voice metric can use a simple, disclosed ratio. A GEO score is often a vendor's composite of several signals. Check the vendor's methodology rather than assuming the labels mean the same thing.

How many prompts should I track?

There is no universal minimum. Use enough prompts to cover the customer questions and intent groups that matter, then disclose the sample. More prompts do not fix a biased list. A smaller, well-documented panel is often more useful than a large collection of vague or repetitive prompts.

Why did my result change when I used the same prompt?

Generated answers are variable. Retrieval results, model updates, location, account state, personalization, and ordinary sampling variation can all affect the answer. Use repeat runs and hold the other conditions steady before interpreting a change.

Should I track every AI platform in one number?

Usually not. Report each platform and model separately first because their products and citation behavior differ. If you publish a combined figure, explain the weighting. A simple average gives a small platform the same influence as a large one, which may not match your audience.

Does schema improve AI search share of voice?

There is no general guarantee that adding schema increases AI mentions or citations. Structured data can help search systems understand eligible page features when it matches visible content and follows provider guidance. Google does not require special schema for AI Overviews or AI Mode.

Does blocking Google-Extended remove a site from AI Overviews?

No. Google says Google-Extended does not affect inclusion or ranking in Google Search. Googlebot and Search eligibility govern Google Search features, including AI Overviews and AI Mode.

How quickly should share of voice improve?

There is no dependable fixed timeline. Crawling, indexing, source selection, model changes, the query set, and the work performed all affect the result. Record the change date and watch repeated samples rather than promising movement within a set number of weeks or months.

Sources