lobodrocks
AUDIT METHODOLOGY

How we measure share of voice in AI answers

The AI Visibility Audit runs 30 buyer-journey questions through ChatGPT, Gemini, Perplexity and Google AI Overviews/AI Mode, five times each in fresh Spain-located sessions, and scores mention rate, citation rate and position-weighted share of voice against five named competitors, each with a 95% confidence interval. It is a measured baseline, delivered in five days as a report, a dataset and a ten-step plan. Replace the closing clause with: "so we commit to the part that is ours: the measurement, the plan and the re-measurement."

By Dmytro LobodUpdated

A measurement you cannot check is an opinion with a chart. This page is the method behind the AI Visibility Audit: what we ask, where, how often, what we record, how the numbers are calculated and how wide the error bars are.

Replace with: "Every third-party figure on this page is named and dated. The arithmetic is shown in full further down, in a worked example." (Keeps the confident frame and the single neutral "worked example" label, and is true.)

What the audit measures

One question: when a buyer in Spain asks an AI assistant something in your category, how often is your brand in the answer, how high up, and who else is there. Four numbers answer it, each per engine and per question group, with a 95% confidence interval:

  • Mention rate: the share of runs in which your brand is named.
  • Citation rate: the share of runs in which a URL on your domain is cited.
  • Position-weighted share of voice: your slice of the answers against five named competitors, with earlier positions weighted more.
  • Source concentration: which third-party pages the engines lean on in your category, and how concentrated that list is.

The ghost-citation rate, the technical check and the ten-step plan exist to explain those numbers and to move them.

The question set: 30 questions, four groups

The unit of measurement is a question a buyer would actually type or say. We build 30 with you in a 45-minute session, then freeze the set. They follow the buyer journey:

Group Count What it covers Example shape
Category 10 "Who does this?" with no brand named best [service] in Barcelona; [service] agency Spain
Problem 10 The problem before the buyer knows the category how do I [outcome]; what does [service] cost
Comparison 6 Alternatives and "versus" questions [brand A] vs [brand B]; agency or freelancer for [service]
Brand 4 Questions that name you what is [brand]; [brand] reviews; [brand] prices

The split flexes (a consumer brand needs more comparison questions, a B2B service more problem questions), but all four groups are always present, because they behave differently. Semrush and Growth Memo's June 2026 ghost-citations study found comparative questions produce brand mentions in 43.3% of appearances against 18% for informational ones. Measure only category questions and you underestimate yourself; measure only brand questions and you flatter yourself.

How the 30 are built:

  1. We start from the five things buyers ask about any category, per Semrush's July 2026 topic-authority study: definition, comparison, alternatives, use case and purchase consideration.
  2. You bring what buyers already say: sales-call notes, support questions, Search Console queries, objections.
  3. Questions are short and conversational, the way people talk to an assistant. The June 2026 Semrush study found short conversational prompts produce 30 to 50 times more brand mentions than long structured ones.
  4. Questions are in the language your buyers use: Spanish for most clients in Spain, English where you sell abroad.
  5. Only the brand group names you; competitors appear only in comparison questions.
  6. You sign off the set and it does not change inside an engagement. In the AI Visibility Program any change goes into a change log, so month three is comparable with month one.

Four engines, five runs, Spain

The engines. ChatGPT with search on, Gemini, Perplexity, and Google AI Overviews together with AI Mode. The choice follows where Spanish users are: SE Ranking's June 2026 study of 101,574 Spanish sites puts ChatGPT at 70% of AI referral traffic, Gemini at 12%, Perplexity at 12%, Copilot at 3% and Claude at 2%. Google's AI features sit on top of the search engine most of Spain still uses, and AI Mode has been live in Spain since October 2025. Claude can be added on request.

Five runs per question. The same question produces different answers on different runs. Kevin Indig's June 2026 piece on prompt tracking in Search Engine Land reports an AirOps study of 815,000 prompt-page pairs in which only 2.2% of ChatGPT citations survived three runs; he calls single-run reporting astrology. Five runs per question per engine is the floor that turns a sighting into a rate. That is 150 runs per engine and 600 per audit.

Fresh sessions. Every run starts clean: no login, no memory, no earlier conversation, no custom instructions. An assistant that remembers yesterday's question is not measuring the market.

Spain locale. Runs are made from a Spanish location with Spanish language settings, and the Google surfaces on google.es. The engines localise their sources: in Ahrefs' July 2026 analysis of 2.26 million Spanish prompts, El País was the fifth most-cited domain on ChatGPT and led AI Mode citations in Spain. Model version and run date are recorded for every run, and the report says whether an engine was run through its API or its consumer app, since the two can differ.

Five competitors, chosen by rule

Share of voice needs a denominator, and a denominator chosen to flatter you is worthless. The rule:

  1. You name the brands you lose deals to. Usually more than five.
  2. We run a pilot pass of ten category questions and list the brands the engines themselves name.
  3. The final five come from the overlap. At least two must have appeared in the pilot; otherwise the benchmark measures your wishlist, not the market.
  4. The set is fixed for the engagement and written into the report. Any later change is logged with the reason.

What we record per run

Each of the 600 runs is one row in the dataset you receive: the full answer text, so anything can be re-scored later; engine, model version, date, question and group; every brand mentioned, in order; every brand cited, with URL; your own mention and position; your own citation URL; sentiment towards your brand by a fixed rule (recommended, listed neutrally, advised against); a ghost-citation flag; and the list of cited domains.

The metrics, with formulas

Mention rate = runs where your brand is named ÷ all runs, per engine.

Citation rate = runs where a URL on your domain is cited ÷ all runs, per engine.

Position-weighted share of voice. In each run, every brand among the six (you plus five competitors) gets points by position in the answer:

Position in the answer Weight
First brand named 1.0
Second 0.6
Third 0.4
Fourth or later 0.2
Not named 0

Share of voice = your points across all runs ÷ the points of all six brands across all runs. Position matters: Semrush's July 2026 study found 74% of users choose the top-mentioned brand.

Source concentration = citations to the ten most-cited domains ÷ all citations, per engine. Next to the percentage we list those domains and the pages that drive competitor mentions. That list is the outreach plan, and the part of the report a PR agency can act on directly.

A worked example

One engine, 30 questions, 5 runs: 150 runs.

  • Your brand is named in 48 runs: mention rate 48 ÷ 150 = 32%.
  • A URL on your domain is cited in 21 runs: citation rate 21 ÷ 150 = 14%. In 9 of those 21 the brand was not named: ghost-citation rate 9 ÷ 21 = 43%.
  • Positions of your 48 mentions: 10 first, 14 second, 12 third, 12 later. Points: 10 × 1.0 + 14 × 0.6 + 12 × 0.4 + 12 × 0.2 = 25.6.
  • Competitor A is named in only 40 runs, but 30 of them first: 30 × 1.0 + 6 × 0.6 + 4 × 0.4 = 35.2 points.
  • All six brands total 118.0 points. Share of voice: you 25.6 ÷ 118.0 = 21.7%; Competitor A 35.2 ÷ 118.0 = 29.8%.
  • The engine cited 1,950 URLs across the 150 runs; the ten most-cited domains account for 897 of them: source concentration 46%.

You are named more often than Competitor A and still hold less share of voice, because A leads the answers. A mention count hides that.

The 95% confidence interval, and why single runs are astrology

Every rate is a sample. For mention and citation rates we use the standard interval for a proportion: rate ± 1.96 × √(rate × (1 − rate) ÷ runs).

In the example, the 32% mention rate over 150 runs carries ±7.5 points, so the honest statement is "between roughly 25% and 40%". The 14% citation rate carries ±5.6 points: "between 8% and 20%". For share of voice, a ratio of weighted sums, we resample the runs (a bootstrap) and report the 2.5th and 97.5th percentiles.

Three rules follow:

  1. Per engine, never blended. Indig's June 2026 analysis shows that averaging engines and reasoning levels creates artefacts of up to 18 points. A pooled figure appears only below the per-engine ones, with its own caveat.
  2. A three-point move is not news. With an interval of ±7, a change from 32% to 35% is noise, and the monthly report says so.
  3. The interval is a floor. The formula treats runs as independent, which flatters precision slightly, since five runs of one question are not five questions.

A single run has no interval at all. Reported as a score, it is astrology with a logo.

The ghost-citation rate

Being cited and being named are different events, and the engines split them differently. Semrush and Growth Memo's June 2026 study of 3,981 brand appearances across four engines and 14 countries found 61.7% of AI citations were ghost citations: URL listed, brand not named. ChatGPT cited brands' URLs in 87% of appearances but named them in 20.7%; Gemini did the reverse, naming brands 83.7% of the time and citing them 21.4%.

So we report both rates and the ghost-citation rate between them. A high ghost rate usually means informational pages that engines quote but that never say who you are; the fix is comparison and priced pages, which earn mentions.

The technical check

It is short, because most of what moves the numbers is content and third-party coverage, not markup: Muck Rack's May 2026 analysis of 25 million links found 84% of AI citations point to earned media, and paid content got 0.3%. We check what can block you and what helps the engines work out who you are:

  • Crawlability. robots.txt allows OAI-SearchBot, PerplexityBot, Google-Extended and Claude's crawlers; content is server-rendered, not JavaScript-only; Bing Webmaster Tools and IndexNow are set up, since the Bing index feeds Copilot and part of ChatGPT retrieval.
  • Schema. Organization or ProfessionalService, Person for the founder, Service with offers, FAQPage where there are questions, and sameAs pointing to the brand's own profiles. Google's guidance, updated July 2026, says no markup is required for AI features; schema is entity hygiene, not a ranking lever.
  • Entity. One name, one description, one entity home, corroborated by LinkedIn, Google Business Profile, Wikidata and the directories the engines cite.
  • llms.txt. We check that it exists and is accurate, and we say plainly that engines ignore it: Google's July 2026 guide states Search ignores llms.txt, and SE Ranking's November 2025 study of about 300,000 domains found no relationship between having one and being cited. Keep it for agents and people; expect zero citation effect. (We keep one ourselves, for the same reasons.)
  • One URL per buyer question. Surfer's December 2025 study found pages ranking for the fan-out sub-queries an engine issues are 161% more likely to be cited. One page cannot rank for thirty questions.

Each item is a pass or a written-out fix. There is no "AI readiness score out of 100"; one score hides which item is blocking you.

The ten-step plan and the deliverable

Ten steps because that is what a team can own in a quarter. The order is always the same; the specifics come from your data: entity fixes; crawler access and Bing/IndexNow; schema corrections; one page per buyer question, starting where nobody owns the answer; a comparison page that names competitors fairly; an original-data page the trade press can quote; FAQ blocks and visible dates; the outreach list from source concentration; reviews and directory profiles on the sites the engines cite; and a measurement cadence that says what a real change looks like given the interval.

The deliverable: the report per engine and question group with intervals, the 600-row dataset as CSV, the question set and competitor rule, the technical checklist, the cited-sources list, the ten steps with owner and effort, and a 45-minute readout.

What we commit to

  • Numbers you can audit yourself. Every rate comes per engine with its 95% confidence interval, and the 600-run dataset ships with the report, so any figure in it can be recalculated from the raw answers.
  • Movement, not a guaranteed spot. Google AI Mode replaces 56% of its cited sources every week and only 2.2% of ChatGPT citations survive three runs (Indig, June 2026). Anyone promising a fixed position in ChatGPT is selling something nobody controls. We commit to the baseline, the plan and the re-measurement that shows whether it worked.
  • A baseline built to become a trend. One audit is one point in time. Monthly waves with the same question set turn it into a trend, which is what the AI Visibility Program is for.
  • Shortlists and lead quality. AI referrals were 0.30% of traffic to Spanish sites between January and April 2026 (SE Ranking). The value is being on the shortlist when a buyer asks, not the sessions.
  • Both halves of the picture. With 61.7% of citations being ghost citations, being cited and being recommended are different events, so the report gives mention rate, citation rate and the ghost rate between them.

Timeline and price

Five working days from a signed-off question set. Day 0: the 45-minute brief that produces the questions and competitors. Days 1 and 2: the 600 runs. Day 3: scoring, intervals and the technical check. Day 4: the plan. Day 5: the readout.

Price: fixed fee, quoted within 24 hours, EU invoice, prices ex-VAT. You own the dataset and the report. Agencies can run the audit referral, co-branded or white-label: see how we work with partners, or the packages for the rest of the studio's work. Terms used here are defined in the AI visibility glossary.

Questions

Why 30 questions and not 300?

Thirty questions, five runs and four engines already produce 600 answers, enough to put a usable confidence interval on each rate per engine. More questions widen coverage but not precision; more runs per question tighten the interval. If your category is broad, we split it into two sets of 30 rather than diluting one. The AI Visibility Program can add questions later, with the change logged.

Why five runs per question in fresh sessions?

Because the same question gets different answers on different runs. Kevin Indig's June 2026 tracking method found that after three ChatGPT runs only 2.2% of citations persisted, and he calls single-run reports astrology. Five runs per question per engine is the floor at which a rate becomes a rate rather than an anecdote. Fresh sessions mean no login, no memory and no earlier chat that could steer the answer.

Which engines do you use, and why not Claude or Copilot?

ChatGPT with search, Gemini, Perplexity and Google AI Overviews/AI Mode, all set to Spain. SE Ranking's June 2026 Spain data puts ChatGPT at 70% of AI referral traffic to Spanish sites, Gemini and Perplexity at 12% each, Copilot at 3% and Claude at 2%, and Google's AI features sit on top of the search engine most of Spain still uses. Claude can be added on request; we leave it out by default because of that 2%.

Can you guarantee that my brand appears in ChatGPT?

No, and be wary of anyone who does. Google AI Mode replaces 56% of its cited sources every week and only 2.2% of ChatGPT citations survive three runs, so a promise of fixed placement is a promise about a moving target. What we sell is a measured baseline, a plan built on what the engines actually cite in your category, and monthly re-measurement with error bars if you continue into the AI Visibility Program.

What exactly do I receive after five days?

A report with the four metrics per engine and per question group, each with its confidence interval; the full dataset of 600 runs as a CSV, with answer text, brands, URLs and flags per run; the signed-off question set and the competitor rule; the technical checklist with pass/fix items; the list of cited sources that drive competitor mentions; the ten-step plan; and a 45-minute readout call. Model versions and run dates are printed on the cover.

How much does the audit cost, and can an agency resell it?

It is a fixed fee, quoted within 24 hours once we know your category and markets; no hourly rates. Agencies can run it referral, co-branded or white-label, and the first meeting with a PR, branding, event, media or SEO agency includes a free audit for one of its clients. Details are on the partners page.

Sources

  1. Kevin Indig, Search Engine Land, How to make prompt tracking much more accurate 2026-065 runs per prompt; 2.2% citation persistence after three ChatGPT runs; AI Mode replaces 56% of sources weekly; single runs as astrology; 18-point blending artefacts
  2. Semrush and Growth Memo, ghost citations study 2026-063,981 appearances, 115 prompts, 14 countries, 4 engines; 61.7% ghost citations
  3. Semrush and Growth Memo, ChatGPT topic authority study 2026-071,094 US categories, 50,000 brands; the five buyer questions; 74% choose the top-mentioned brand
  4. Muck Rack, What Is AI Reading? (third edition) 2026-0525M+ links across ChatGPT, Claude and Gemini
  5. SE Ranking, Tráfico de IA en España 2026-06101,574 sites with SE Ranking analytics; AI referral shares by engine; AI referrals 0.30% of traffic
  6. Ahrefs ES, Cuáles son los medios más citados por la IA en España 2026-072.26M Spanish prompts, six engines
  7. Google Search Central, Optimizing your website for generative AI features on Google Search 2026-07Published May 2026, updated July 2026; states that Google Search ignores llms.txt and that no markup is required
  8. SE Ranking llms.txt study, via Search Engine Journal 2025-11About 300,000 domains; 10.13% adoption; no relationship with citation frequency
  9. Surfer fan-out study, via Search Engine Land 2025-1210,000 keywords; pages ranking for fan-out queries 161% more likely to be cited
  10. Human Level, AI Mode arrives in Spain 2025-10

Checked at the date shown; figures move.

Talk to Lobod

Start with the numbers. Then we build.

Tell me what you're launching, or what isn't working. Within 24 hours you get either the package that fits, with a price, or a five-day Research Sprint proposal.

Dima@lobods.comLinkedIn

Barcelona · Spain & EU · EN / ES · replies same day
Letters

Twice a month, something worth reading about AI visibility, launches and what actually moves a brand.

Double opt-in: one letter asks you to confirm, and nothing arrives until you do. One click at the bottom of any letter leaves for good.