AI visibility has a vocabulary problem: three acronyms for one job, "citation" and "mention" used as if they were the same thing, and file formats sold as strategy. This glossary is the set of terms we use in the AI Visibility Audit, in its methodology and in client reports, defined once so that a report can be read without a call.
It is maintained by Dmytro Lobod at lobod.rocks and was last updated on 9 September 2026. Every number carries its source and date; where a figure is a vendor's own analysis rather than a documented study, it says so. Each entry ends with the Spanish equivalent and common synonyms. If you find an error, write to us; we would rather correct a definition than defend it.
Terms
AI visibility
How often, how prominently and how favourably a brand appears in the answers of AI assistants (ChatGPT, Gemini, Perplexity, Google AI Overviews and AI Mode) when people ask about its category. Measured as mention rate, citation rate and share of voice over a fixed prompt set, per engine, with a confidence interval. It is not traffic: AI referrals were 0.30% of visits to Spanish sites in early 2026 (SE Ranking, June 2026), while 28% of people in Spain used ChatGPT regularly in 2025 (Funcas, January 2025). The value is being on the shortlist, not the click. Also: visibilidad en IA; AI search visibility; brand visibility in LLMs.
GEO (generative engine optimization)
The term comes from a November 2023 paper by Aggarwal and colleagues at Princeton and other universities, "GEO: Generative Engine Optimization", which tested content changes (adding statistics, quotations, citations) and measured their effect on visibility in generated answers. In practice today it means the work of making a brand and its pages more likely to be retrieved, cited and named by generative engines. It overlaps heavily with SEO, since Google states its AI features are rooted in core ranking (July 2026), and with PR, since most citations go to earned media. Also: optimización para motores generativos; generative engine optimisation.
AEO (answer engine optimization)
Optimising content to be the direct answer rather than a link: question-shaped headings, a two-sentence direct answer at the top, FAQ blocks, structured data. The label predates generative AI (it was used for featured snippets and voice assistants) and was adopted for AI answers. The difference from GEO is emphasis: AEO on the answer format, GEO on retrieval and citation. On-site AEO is necessary, not sufficient: Muck Rack found 84% of AI citations point to earned media (May 2026). Also: optimización para motores de respuesta; answer engine optimisation.
LLMO / LLM SEO
Large language model optimization. Same goal as GEO, with more weight on how a brand is represented in a model's own knowledge (its training data and entity descriptions) rather than only in the pages retrieved at answer time. Used mainly by B2B agencies; some position themselves as "LLMO agencies". We treat it as a synonym for GEO and try not to invent a fourth acronym. Also: LLM SEO; optimización para LLM; large language model optimisation.
AI search optimization
The umbrella term Google and the larger SEO platforms use for all of the above. Google's guide "Optimizing your website for generative AI features on Google Search" (published May 2026, updated July 2026) sets the official position: no new machine-readable files, no AI markup, no chunking or AI-specific rewriting; core ranking and quality systems apply; what helps is non-commodity content with a unique point of view, images and video, Business Profile data and structured data as part of overall SEO. Also: optimización para búsqueda con IA; optimising for AI features.
Share of voice in AI answers
A brand's slice of the attention across a fixed set of questions on one engine, against named competitors. lobod.rocks weights brands by position in each answer (first 1.0, second 0.6, third 0.4, later 0.2), sums the points across all runs and divides a brand's points by the points of all six brands measured. It is reported per engine with a 95% confidence interval, never as one blended "AI visibility score": Kevin Indig showed in June 2026 that blending engines and reasoning levels creates artefacts of up to 18 points. Formulas and a worked example are on the methodology page. Also: share of voice en respuestas de IA; AI share of voice; cuota de voz en IA.
Mention vs citation
A mention is the brand named in the answer text. A citation is a URL on the brand's domain listed as a source. The two diverge sharply by engine: in Semrush and Growth Memo's June 2026 study, ChatGPT cited brands' URLs in 87% of appearances but named them in 20.7%, while Gemini named brands 83.7% of the time and cited them 21.4%. At topic level, mentions and citations correlate slightly negatively (−0.229; Semrush, July 2026). Informational pages earn citations; comparison and priced pages earn mentions. Track both. Also: mención frente a cita; brand mention; source citation.
Ghost citation
A citation without a mention: the URL is listed as a source, the brand is not named in the answer. The term comes from Semrush and Growth Memo's June 2026 study of 3,981 brand appearances across 115 prompts, 14 countries and four engines, in which 61.7% of AI citations were ghost citations. Typical of informational pages that engines like to quote but that never say who wrote them. The same study found comparative prompts produce brand mentions 43.3% of the time against 18% for informational ones, which is the fix. Also: cita fantasma; citation without mention.
Query fan-out
An engine's technique of splitting one question into several sub-queries across subtopics and data sources, retrieving for each, then writing one answer. Google names it in its AI features guide (2026). Surfer's December 2025 study of 10,000 keywords found pages ranking for the fan-out sub-queries are 161% more likely to be cited, and 68% of citations came from pages that ranked in the top 10 for neither the main query nor a fan-out sub-query. Ahrefs (March 2026) found only 38% of AI Overview citations come from top-10 pages, down from 76% in July 2025. Practical consequence: one URL per buyer sub-question. Also: fan-out de consultas; expansión de consultas; sub-queries.
Grounding / RAG
Retrieval-augmented generation: the model retrieves documents from a search index at answer time and writes with them, citing what it used. Gemini's "Grounding with Google Search" generates one or more queries, retrieves from the live Google index and returns citations (Google AI for Developers). ChatGPT search and Perplexity do the same with their own indexes. Grounding is why a page published this week can be cited this week, and also why citations churn: a different retrieval gives a different answer. Also: grounding; generación aumentada por recuperación; retrieval-augmented generation.
AI Overviews vs AI Mode
AI Overviews is the AI summary shown above Google's results, rolled out in the US in May 2024, with about eight citations per answer (Muck Rack, May 2026). AI Mode is a full conversational search tab; it launched in Spain in October 2025 (Human Level). Both are rooted in core ranking and both may use query fan-out (Google, 2026). Both are volatile: AI Mode replaces 56% of its cited sources every week (Indig, June 2026), and when Gemini 3 became the AI Overviews default in January 2026 it replaced about 42% of previously cited domains (SE Ranking). Google's products favour newspapers more than chat assistants do; El País leads AI Mode citations in Spain (Ahrefs ES, July 2026). Also: AI Overviews (resúmenes de IA de Google); AI Mode (Modo IA).
Entity and entity home
An entity is a thing an engine can identify unambiguously: an organisation, a person, a product, a place. The entity home, a term Jason Barnard set out in Search Engine Land in March 2026, is the one canonical URL that describes the entity and that every other profile corroborates: LinkedIn, Google Business Profile, Wikidata, Crunchbase, directories. Assistants reconcile entities by consistent name, URL and description across those sources. Two spellings of a founder's name, two founding years or a sameAs pointing at a parent company split the entity and keep an assistant describing you as something you are not.
Also: entidad y entity home; página de referencia de la entidad.
sameAs
A schema.org property stating that the entity on this page is the same as the one at another URL: a LinkedIn page, a Wikidata item, a Crunchbase profile, a Google Maps listing. It asserts identity, not relationship. A parent company belongs in parentOrganization or memberOf, not in sameAs; a sister brand's Instagram is not you. Cheap, and one of the few markup items that genuinely helps entity resolution. Not a ranking lever.
Also: sameAs (propiedad de schema.org); perfiles equivalentes.
Knowledge graph
A database of entities and the relationships between them, which engines consult to answer "who or what is X". Google's Knowledge Graph draws on Wikidata, Wikipedia, Google Business Profile and the web at large. For a young brand the realistic path is a Wikidata item (its bar is verifiability, not Wikipedia's notability), a Business Profile, and the same description on every profile. A Wikipedia article for a small company will be deleted and can poison the entity; Wikidata yes, Wikipedia no. Also: grafo de conocimiento; Knowledge Graph de Google.
OAI-SearchBot
OpenAI's crawler that, in OpenAI's words, "surfaces websites in ChatGPT's search results". Sites that block it are not shown in ChatGPT search answers; robots.txt changes take about 24 hours to apply (OpenAI crawler documentation). It is distinct from GPTBot, which collects training data and can be blocked without affecting search. ChatGPT's retrieval is no longer Bing-only: Peec's September 2026 analysis of ChatGPT server events found a proprietary index alongside at least eight external providers, so being crawlable by OAI-SearchBot matters in its own right. Also: OAI-SearchBot (rastreador de búsqueda de OpenAI).
ChatGPT-User
OpenAI's on-demand fetcher. It retrieves a page when a user's action requires it, for example opening a link inside a conversation, and OpenAI states it is "not used to determine whether content may appear in Search". In your server logs, ChatGPT-User means a person's session fetched a page of yours; OAI-SearchBot means you are being indexed for search. Counting OAI-SearchBot hits monthly is the earliest sign that citations may follow; ChatGPT-User hits are a lagging signal, since they mean someone already reached your page from a ChatGPT answer. Also: ChatGPT-User (agente de usuario de ChatGPT).
PerplexityBot
Perplexity's crawler, which "surfaces and links websites in search results". A second agent, Perplexity-User, fetches on demand and "generally ignores robots.txt" (Perplexity documentation). To appear in Perplexity, allow PerplexityBot and its published IP ranges. Perplexity is the most community-skewed engine, with Reddit at 6.6% of all its citations (Profound, 2025, vendor analysis); in Spain its most-cited domains are YouTube, Spanish Wikipedia, Reddit, TikTok and Instagram (Ahrefs ES, October 2025). Also: PerplexityBot; Perplexity-User.
ClaudeBot
Anthropic's crawler for training data. Two other Anthropic agents matter more for visibility: Claude-SearchBot, which "indexes content to improve search result quality" (blocking it "may reduce your site's visibility"), and Claude-User, which fetches on demand (Anthropic support documentation). Anthropic's Trust Center lists Brave Search as a web-search subprocessor (March 2025). Claude is the most selective citer, with 55% of responses carrying citations (Muck Rack, May 2026), and accounts for 2% of AI referral traffic in Spain (SE Ranking, June 2026). Also: ClaudeBot; Claude-SearchBot; Claude-User.
Google-Extended
A robots.txt token, introduced in September 2023, that controls whether your content may be used for Gemini training and for grounding Gemini's answers. Google states it "does not impact a site's inclusion in Google Search nor is it used as a ranking signal" (Google crawlers documentation). Blocking it removes you from Gemini grounding while leaving Search untouched; allow it if you want to be visible in Gemini. AI Overviews and AI Mode follow normal Google indexing and are not controlled by this token. Also: Google-Extended (token de robots.txt para Gemini).
llms.txt
A proposal by Jeremy Howard in September 2024 for a Markdown file at /llms.txt that summarises a site for language models, with an optional llms-full.txt carrying the full text. The evidence on effect is consistent: Google's guide (July 2026) states that Search ignores llms.txt; SE Ranking's November 2025 study of about 300,000 domains found 10.13% adoption and no relationship with citation frequency; OpenAI's and Anthropic's crawler documentation never mention it. Keep one because it is cheap and useful to agents and people; expect zero citation effect. lobod.rocks keeps one and says so. Also: llms.txt (archivo de resumen para modelos de lenguaje).
Content-Signal
A robots.txt extension published by Cloudflare in September 2025 that lets a site declare how its content may be used: search, ai-input and ai-train, each yes or no. It sits on 3.8 million or more Cloudflare-managed domains, is not a standard, and Google has not committed to honouring it. A harmless declaration of intent; not a visibility lever in either direction.
Also: Content-Signal (política de señales de contenido de Cloudflare).
IndexNow
A protocol launched by Microsoft and Yandex in 2021 that pushes new or changed URLs to participating search engines within seconds instead of waiting for a crawl. The Bing index it feeds serves Bing, Copilot, DuckDuckGo, Ecosia and part of ChatGPT's retrieval. Free, about 30 minutes to set up, and not a documented ranking factor anywhere. Do it as hygiene; do not expect it to be "the ChatGPT switch". Also: IndexNow; indexación push.
Structured data / schema.org
Machine-readable markup, usually JSON-LD, drawn from the schema.org vocabulary founded in 2011 by the major search engines, that states what a page is about: Organization, Person, Service, Article, FAQPage, Dataset. Google says no markup is required for AI features and recommends structured data as part of overall SEO (July 2026). Vendor figures such as "2.3x more citations" are correlations from vendor data. Use it for entity disambiguation and eligibility, not as a citation lever. Also: datos estructurados; schema.org; marcado semántico.
FAQPage
The schema.org type for a page that contains a list of questions with their answers. The markup makes the block explicit to machines; the block itself answers the sub-questions that fan-out retrieval issues, in the two-sentence form assistants lift. One vendor analysis reports pages with FAQ blocks cited 1.9 times more often (BrightEdge, 2026, vendor claim). lobod.rocks puts an FAQ block on every content page, this one included. Also: FAQPage (schema de preguntas frecuentes); bloque de FAQ.
Earned media in AI answers
Third-party coverage a brand did not pay for: journalism, trade press, reviews, analyst and directory pages. Muck Rack's May 2026 analysis of more than 25 million links cited by ChatGPT, Claude and Gemini across 17 industries found 84% of citations point to earned media, journalism alone 27%, and paid or advertorial content 0.3%; press releases appear 3.5 times more often in trend answers than in best-of answers. This is why AI visibility work aligns with PR. The outlets each engine trusts in Spain differ (El País, La Vanguardia, El Español, Xataka and others; Ahrefs ES, July 2026). Also: earned media en respuestas de IA; cobertura ganada; medios ganados.
Confidence interval in visibility reporting
The range within which the true rate probably lies, given that every rate is estimated from a sample of runs. At 95%, rate ± 1.96 × √(rate × (1 − rate) ÷ runs). With 30 questions and five runs (150 runs per engine), a 32% mention rate carries about ±7.5 points. A single run has no interval at all; Kevin Indig (June 2026) calls single-run reports astrology, and the AirOps study he reports, of 815,000 prompt-page pairs, found only 2.2% of ChatGPT citations survived three runs. Report per engine, put the bar on every chart, and do not celebrate a three-point move. Also: intervalo de confianza; margen de error.
Prompt set
The fixed list of questions run in every wave of measurement. lobod.rocks audits use 30 questions across four groups (category, problem, comparison, brand), built with the client and frozen after sign-off. Phrasing is short and conversational, because Semrush found short conversational prompts produce 30 to 50 times more brand mentions than long structured ones (June 2026). Any change to the set is logged, so waves stay comparable. Also: conjunto de prompts; conjunto de preguntas; question set.
Run
One execution of one question on one engine in a fresh session: no login, no memory, no earlier conversation. The unit of the dataset. Each run records engine, model version, date, the brands mentioned in order, the brands cited with URLs, the client's own mention and position, sentiment, a ghost-citation flag and the cited domains. Five runs per question per engine is the minimum lobod.rocks uses, following Indig's June 2026 method; more runs tighten the interval. Also: ejecución; pasada; run.
Spain locale
Running the prompt set from a Spanish location, in the buyer's language (Spanish for most clients, English where they sell abroad), on google.es and with the Spain version of AI Mode. Engines localise their sources: for Spanish prompts run in Spain, El País is the fifth most-cited domain on ChatGPT and leads AI Mode, and Perplexity's top domains are YouTube, Spanish Wikipedia, Reddit and TikTok (Ahrefs ES, 2025 to 2026). The engine mix follows Spanish referral shares: ChatGPT 70%, Gemini 12%, Perplexity 12%, Copilot 3%, Claude 2% (SE Ranking, June 2026). A US-located run measures a different market. Also: localización en España; configuración regional de España; Spain locale.
How these fit together
The crawlers (OAI-SearchBot, PerplexityBot, Claude-SearchBot, Googlebot with Google-Extended) decide what enters each engine's index. At answer time, grounding retrieves from that index, usually through query fan-out into sub-questions, and the model writes an answer that names some brands (mentions) and lists some URLs (citations), often without connecting the two (ghost citations). AI visibility is the measurement of that outcome: a prompt set, several runs per question, a Spain locale, and share of voice per engine with a confidence interval. What moves the numbers is the entity being clear (entity home, sameAs, knowledge graph), pages that answer the sub-questions (structured data and FAQPage help the machine read them), and above all earned media, which is where most citations go. What does not move them, on current evidence, is llms.txt, Content-Signal or IndexNow on their own. Everything here is applied in the AI Visibility Audit and Program; the packages are listed here, and agencies can run the same work co-branded or white-label through the partner programme.