A measurement you cannot check is an opinion with a chart. This page is the method behind the AI Visibility Audit: what we ask, where, how often, what we record, how the numbers are calculated and how wide the error bars are.
Replace with: "Every third-party figure on this page is named and dated. The arithmetic is shown in full further down, in a worked example." (Keeps the confident frame and the single neutral "worked example" label, and is true.)
What the audit measures
One question: when a buyer in Spain asks an AI assistant something in your category, how often is your brand in the answer, how high up, and who else is there. Four numbers answer it, each per engine and per question group, with a 95% confidence interval:
- Mention rate: the share of runs in which your brand is named.
- Citation rate: the share of runs in which a URL on your domain is cited.
- Position-weighted share of voice: your slice of the answers against five named competitors, with earlier positions weighted more.
- Source concentration: which third-party pages the engines lean on in your category, and how concentrated that list is.
The ghost-citation rate, the technical check and the ten-step plan exist to explain those numbers and to move them.
The question set: 30 questions, four groups
The unit of measurement is a question a buyer would actually type or say. We build 30 with you in a 45-minute session, then freeze the set. They follow the buyer journey:
| Group | Count | What it covers | Example shape |
|---|---|---|---|
| Category | 10 | "Who does this?" with no brand named | best [service] in Barcelona; [service] agency Spain |
| Problem | 10 | The problem before the buyer knows the category | how do I [outcome]; what does [service] cost |
| Comparison | 6 | Alternatives and "versus" questions | [brand A] vs [brand B]; agency or freelancer for [service] |
| Brand | 4 | Questions that name you | what is [brand]; [brand] reviews; [brand] prices |
The split flexes (a consumer brand needs more comparison questions, a B2B service more problem questions), but all four groups are always present, because they behave differently. Semrush and Growth Memo's June 2026 ghost-citations study found comparative questions produce brand mentions in 43.3% of appearances against 18% for informational ones. Measure only category questions and you underestimate yourself; measure only brand questions and you flatter yourself.
How the 30 are built:
- We start from the five things buyers ask about any category, per Semrush's July 2026 topic-authority study: definition, comparison, alternatives, use case and purchase consideration.
- You bring what buyers already say: sales-call notes, support questions, Search Console queries, objections.
- Questions are short and conversational, the way people talk to an assistant. The June 2026 Semrush study found short conversational prompts produce 30 to 50 times more brand mentions than long structured ones.
- Questions are in the language your buyers use: Spanish for most clients in Spain, English where you sell abroad.
- Only the brand group names you; competitors appear only in comparison questions.
- You sign off the set and it does not change inside an engagement. In the AI Visibility Program any change goes into a change log, so month three is comparable with month one.
Four engines, five runs, Spain
The engines. ChatGPT with search on, Gemini, Perplexity, and Google AI Overviews together with AI Mode. The choice follows where Spanish users are: SE Ranking's June 2026 study of 101,574 Spanish sites puts ChatGPT at 70% of AI referral traffic, Gemini at 12%, Perplexity at 12%, Copilot at 3% and Claude at 2%. Google's AI features sit on top of the search engine most of Spain still uses, and AI Mode has been live in Spain since October 2025. Claude can be added on request.
Five runs per question. The same question produces different answers on different runs. Kevin Indig's June 2026 piece on prompt tracking in Search Engine Land reports an AirOps study of 815,000 prompt-page pairs in which only 2.2% of ChatGPT citations survived three runs; he calls single-run reporting astrology. Five runs per question per engine is the floor that turns a sighting into a rate. That is 150 runs per engine and 600 per audit.
Fresh sessions. Every run starts clean: no login, no memory, no earlier conversation, no custom instructions. An assistant that remembers yesterday's question is not measuring the market.
Spain locale. Runs are made from a Spanish location with Spanish language settings, and the Google surfaces on google.es. The engines localise their sources: in Ahrefs' July 2026 analysis of 2.26 million Spanish prompts, El País was the fifth most-cited domain on ChatGPT and led AI Mode citations in Spain. Model version and run date are recorded for every run, and the report says whether an engine was run through its API or its consumer app, since the two can differ.
Five competitors, chosen by rule
Share of voice needs a denominator, and a denominator chosen to flatter you is worthless. The rule:
- You name the brands you lose deals to. Usually more than five.
- We run a pilot pass of ten category questions and list the brands the engines themselves name.
- The final five come from the overlap. At least two must have appeared in the pilot; otherwise the benchmark measures your wishlist, not the market.
- The set is fixed for the engagement and written into the report. Any later change is logged with the reason.
What we record per run
Each of the 600 runs is one row in the dataset you receive: the full answer text, so anything can be re-scored later; engine, model version, date, question and group; every brand mentioned, in order; every brand cited, with URL; your own mention and position; your own citation URL; sentiment towards your brand by a fixed rule (recommended, listed neutrally, advised against); a ghost-citation flag; and the list of cited domains.
The metrics, with formulas
Mention rate = runs where your brand is named ÷ all runs, per engine.
Citation rate = runs where a URL on your domain is cited ÷ all runs, per engine.
Position-weighted share of voice. In each run, every brand among the six (you plus five competitors) gets points by position in the answer:
| Position in the answer | Weight |
|---|---|
| First brand named | 1.0 |
| Second | 0.6 |
| Third | 0.4 |
| Fourth or later | 0.2 |
| Not named | 0 |
Share of voice = your points across all runs ÷ the points of all six brands across all runs. Position matters: Semrush's July 2026 study found 74% of users choose the top-mentioned brand.
Source concentration = citations to the ten most-cited domains ÷ all citations, per engine. Next to the percentage we list those domains and the pages that drive competitor mentions. That list is the outreach plan, and the part of the report a PR agency can act on directly.
A worked example
One engine, 30 questions, 5 runs: 150 runs.
- Your brand is named in 48 runs: mention rate 48 ÷ 150 = 32%.
- A URL on your domain is cited in 21 runs: citation rate 21 ÷ 150 = 14%. In 9 of those 21 the brand was not named: ghost-citation rate 9 ÷ 21 = 43%.
- Positions of your 48 mentions: 10 first, 14 second, 12 third, 12 later. Points: 10 × 1.0 + 14 × 0.6 + 12 × 0.4 + 12 × 0.2 = 25.6.
- Competitor A is named in only 40 runs, but 30 of them first: 30 × 1.0 + 6 × 0.6 + 4 × 0.4 = 35.2 points.
- All six brands total 118.0 points. Share of voice: you 25.6 ÷ 118.0 = 21.7%; Competitor A 35.2 ÷ 118.0 = 29.8%.
- The engine cited 1,950 URLs across the 150 runs; the ten most-cited domains account for 897 of them: source concentration 46%.
You are named more often than Competitor A and still hold less share of voice, because A leads the answers. A mention count hides that.
The 95% confidence interval, and why single runs are astrology
Every rate is a sample. For mention and citation rates we use the standard interval for a proportion: rate ± 1.96 × √(rate × (1 − rate) ÷ runs).
In the example, the 32% mention rate over 150 runs carries ±7.5 points, so the honest statement is "between roughly 25% and 40%". The 14% citation rate carries ±5.6 points: "between 8% and 20%". For share of voice, a ratio of weighted sums, we resample the runs (a bootstrap) and report the 2.5th and 97.5th percentiles.
Three rules follow:
- Per engine, never blended. Indig's June 2026 analysis shows that averaging engines and reasoning levels creates artefacts of up to 18 points. A pooled figure appears only below the per-engine ones, with its own caveat.
- A three-point move is not news. With an interval of ±7, a change from 32% to 35% is noise, and the monthly report says so.
- The interval is a floor. The formula treats runs as independent, which flatters precision slightly, since five runs of one question are not five questions.
A single run has no interval at all. Reported as a score, it is astrology with a logo.
The ghost-citation rate
Being cited and being named are different events, and the engines split them differently. Semrush and Growth Memo's June 2026 study of 3,981 brand appearances across four engines and 14 countries found 61.7% of AI citations were ghost citations: URL listed, brand not named. ChatGPT cited brands' URLs in 87% of appearances but named them in 20.7%; Gemini did the reverse, naming brands 83.7% of the time and citing them 21.4%.
So we report both rates and the ghost-citation rate between them. A high ghost rate usually means informational pages that engines quote but that never say who you are; the fix is comparison and priced pages, which earn mentions.
The technical check
It is short, because most of what moves the numbers is content and third-party coverage, not markup: Muck Rack's May 2026 analysis of 25 million links found 84% of AI citations point to earned media, and paid content got 0.3%. We check what can block you and what helps the engines work out who you are:
- Crawlability. robots.txt allows OAI-SearchBot, PerplexityBot, Google-Extended and Claude's crawlers; content is server-rendered, not JavaScript-only; Bing Webmaster Tools and IndexNow are set up, since the Bing index feeds Copilot and part of ChatGPT retrieval.
- Schema. Organization or ProfessionalService, Person for the founder, Service with offers, FAQPage where there are questions, and
sameAspointing to the brand's own profiles. Google's guidance, updated July 2026, says no markup is required for AI features; schema is entity hygiene, not a ranking lever. - Entity. One name, one description, one entity home, corroborated by LinkedIn, Google Business Profile, Wikidata and the directories the engines cite.
- llms.txt. We check that it exists and is accurate, and we say plainly that engines ignore it: Google's July 2026 guide states Search ignores llms.txt, and SE Ranking's November 2025 study of about 300,000 domains found no relationship between having one and being cited. Keep it for agents and people; expect zero citation effect. (We keep one ourselves, for the same reasons.)
- One URL per buyer question. Surfer's December 2025 study found pages ranking for the fan-out sub-queries an engine issues are 161% more likely to be cited. One page cannot rank for thirty questions.
Each item is a pass or a written-out fix. There is no "AI readiness score out of 100"; one score hides which item is blocking you.
The ten-step plan and the deliverable
Ten steps because that is what a team can own in a quarter. The order is always the same; the specifics come from your data: entity fixes; crawler access and Bing/IndexNow; schema corrections; one page per buyer question, starting where nobody owns the answer; a comparison page that names competitors fairly; an original-data page the trade press can quote; FAQ blocks and visible dates; the outreach list from source concentration; reviews and directory profiles on the sites the engines cite; and a measurement cadence that says what a real change looks like given the interval.
The deliverable: the report per engine and question group with intervals, the 600-row dataset as CSV, the question set and competitor rule, the technical checklist, the cited-sources list, the ten steps with owner and effort, and a 45-minute readout.
What we commit to
- Numbers you can audit yourself. Every rate comes per engine with its 95% confidence interval, and the 600-run dataset ships with the report, so any figure in it can be recalculated from the raw answers.
- Movement, not a guaranteed spot. Google AI Mode replaces 56% of its cited sources every week and only 2.2% of ChatGPT citations survive three runs (Indig, June 2026). Anyone promising a fixed position in ChatGPT is selling something nobody controls. We commit to the baseline, the plan and the re-measurement that shows whether it worked.
- A baseline built to become a trend. One audit is one point in time. Monthly waves with the same question set turn it into a trend, which is what the AI Visibility Program is for.
- Shortlists and lead quality. AI referrals were 0.30% of traffic to Spanish sites between January and April 2026 (SE Ranking). The value is being on the shortlist when a buyer asks, not the sessions.
- Both halves of the picture. With 61.7% of citations being ghost citations, being cited and being recommended are different events, so the report gives mention rate, citation rate and the ghost rate between them.
Timeline and price
Five working days from a signed-off question set. Day 0: the 45-minute brief that produces the questions and competitors. Days 1 and 2: the 600 runs. Day 3: scoring, intervals and the technical check. Day 4: the plan. Day 5: the readout.
Price: fixed fee, quoted within 24 hours, EU invoice, prices ex-VAT. You own the dataset and the report. Agencies can run the audit referral, co-branded or white-label: see how we work with partners, or the packages for the rest of the studio's work. Terms used here are defined in the AI visibility glossary.