AI SearchGEOEntity SEO

AI Search Glossary: 56 Terms Every Marketer Needs

This glossary defines the 56 terms that come up most in AI search and generative engine optimization, in plain English. Each definition is written to stand on its own. Where we have published a full guide on a term, the definition links to it.

Jump to: A · B · C · D · E · F · G · H · I · J · K · L · M · N · O · P · R · S · T · V · W · Z

A

AEO (answer engine optimization)

Optimizing content to be the direct answer to a question, originally for featured snippets and voice assistants and now for AI answers as well. It overlaps heavily with GEO. See AEO vs GEO vs SEO.

AI agent

An AI system that carries out multi-step tasks for a user, such as comparing vendors, filling in forms, or booking, rather than only answering questions. See agentic search.

AI crawler

A bot that fetches web pages for an AI company, either to collect training data, such as GPTBot, CCBot, and ClaudeBot, or to retrieve pages for live answers, such as OAI-SearchBot and PerplexityBot. See robots.txt for AI crawlers.

AI Mode

Google’s conversational search experience, powered by Gemini, where a generated answer with citations replaces the list of results and follow-up questions are expected. See Google AI Mode explained.

AI Overviews

Google’s AI-generated summaries shown above traditional results on a standard search page, with links to their sources. See how to appear in AI Overviews.

AI visibility

How often, how prominently, and how accurately a brand appears in AI-generated answers. Usually measured as presence rate, citations, accuracy, and share of model. See the AI visibility metrics that matter.

Answer engine

Any system that responds to a question with a written answer instead of a list of links: ChatGPT, Perplexity, Claude, Gemini, and Google’s AI Overviews and AI Mode. Optimizing for them is what AEO and GEO describe. See how AI engines choose citations.

Answer-first writing

Structuring content so the first sentence under each heading states the answer, with the supporting detail after it. It makes passages easier for AI engines to extract and quote.

B

Branded mention

Any reference to a brand by name on a third-party site, linked or not. In 2026 correlation analyses, branded mentions tracked AI Overviews visibility far more closely than backlinks did. See digital PR for AI search.

Bingbot

Microsoft’s search crawler. It matters for AI search beyond Bing itself, because several AI answer products draw on Bing’s index, so a page Bing has not indexed is often invisible to them too. See how to rank in ChatGPT search.

C

Chunk

A passage of a page that a retrieval system stores and matches as its own unit. AI engines retrieve and quote chunks, not whole pages. See content chunking for AI.

Citation

A source an AI answer links to or names as the basis for a claim. Being cited is different from being retrieved: only a fraction of retrieved pages end up cited. See retrieval vs citation and what is AI citation tracking.

ClaudeBot

Anthropic’s crawler for collecting public web content. Like GPTBot, it is controlled through its own user-agent line in robots.txt, and blocking it is a separate decision from blocking the bots that fetch pages for live answers. See should you block AI crawlers.

Common Crawl

A nonprofit that publishes a large, free archive of web crawl data, collected by its CCBot crawler. Filtered versions of that archive have been a major component of many language models’ training data.

Corroboration

Independent sources confirming the same facts about an entity. Corroboration is what turns a brand’s claim about itself into a fact an engine is willing to repeat.

D

dateModified

The schema.org property that states when a page’s content last changed. It is only a useful freshness signal when it moves because the content genuinely changed, not on every deploy. See structured data for AI search.

E

E-E-A-T

Experience, Expertise, Authoritativeness, and Trustworthiness: the qualities Google’s search quality rater guidelines use to assess content. Named, credentialed authors and verifiable facts are how sites demonstrate them. See Person schema and author entities.

Embedding

A numerical representation of a passage’s meaning. Retrieval systems compare the embedding of a query with the embeddings of stored passages to find matches, which is why a passage can be retrieved without containing the exact words of the query.

Entity

A distinct, identifiable thing, such as a company, person, product, or place, that search engines and AI models can recognize and attach facts to. See the entity SEO guide.

Entity disambiguation

The work of making engines confident which real-world thing your brand is when its name is shared with other companies, people, or places. See entity disambiguation.

Entity home

The single page, usually the homepage or About page, that states authoritatively what an entity is, and that other signals such as structured data and profiles point back to.

F

Fan-out

Short for query fan-out: the process by which an AI engine breaks one question into many related sub-queries, runs them, and combines the results. In one 2026 analysis, 95% of fan-out queries had no traditional search volume. See Google AI Mode explained.

FAQPage schema

Structured data that marks up a list of questions and answers. Google has limited FAQ rich results to well-known government and health sites since August 2023, but the markup still describes Q&A content to machines. See FAQ schema for AI answers.

G

GEO (generative engine optimization)

The practice of getting a brand understood, trusted, and cited by AI engines such as ChatGPT, Perplexity, Gemini, and Google’s AI features. See what is GEO and the complete GEO guide.

Google-Extended

A robots.txt token that controls whether Google may use a site’s content to train Gemini models and to ground their answers. It does not affect Googlebot or regular search rankings.

GPTBot

OpenAI’s crawler for collecting training data. Blocking it does not stop ChatGPT search from retrieving your pages; that is done by OAI-SearchBot.

Grounding

Tying an AI answer to specific retrieved sources rather than the model’s memory alone. Grounded answers can cite where their facts came from; ungrounded ones cannot.

H

Hallucination

An AI answer that states something false with confidence, often filling a gap where the model lacks reliable information. Thin or inconsistent entity data makes hallucinations about a brand more likely. See when AI gets your brand wrong.

I

IndexNow

An open protocol that lets a site notify participating search engines, including Bing, Yandex, Seznam, and Naver, the moment a URL is added or changed. Google does not participate.

J

JSON-LD

The recommended format for adding structured data to a page: a block of JSON inside a script tag, kept separate from the visible HTML. See structured data for AI search.

K

Knowledge cutoff

The date after which a language model has no information from its training data. Anything newer can only reach the model’s answers through live retrieval. See LLM SEO.

Knowledge Graph

Google’s database of entities and the relationships between them, which powers Knowledge Panels and shapes how search understands brands. See the knowledge graph optimization checklist.

Knowledge Panel

The information box Google shows for an entity it has resolved confidently, drawn from the Knowledge Graph. See what is a Google Knowledge Panel.

L

LLM (large language model)

An AI model trained on very large amounts of text to understand and generate language. ChatGPT, Gemini, Claude, and the models behind Perplexity are all built on LLMs.

LLM SEO

Also called LLM optimization or LLMO: optimizing a brand to be found, understood, and cited by large language models. Largely interchangeable with GEO. See LLM SEO.

llms.txt

A proposed file, placed at the root of a site, that lists its most important content in a clean format for language models. Proposed in September 2024; the major engines have not confirmed they use it. See what is llms.txt.

M

Multimodal

Able to process more than one kind of input, such as text and images together. Multimodal models can read the text inside an image and interpret a chart. See image SEO for AI search.

N

NAP consistency

Keeping a business’s name, address, and phone number identical everywhere they appear online. Inconsistent NAP lowers confidence in both local and entity signals. See the local SEO checklist.

O

OAI-SearchBot

OpenAI’s crawler for ChatGPT search. Allowing it lets ChatGPT retrieve and cite your pages; it is separate from GPTBot, which collects training data.

Organization schema

Structured data that defines a company as an entity, with a stable @id that every other node on the site references. It is the anchor of a connected schema graph. See Organization schema for AI search.

P

Parametric memory

What a language model knows from training, stored in its parameters rather than in any document. It is frozen at the knowledge cutoff and cannot cite sources. See LLM SEO.

PerplexityBot

Perplexity’s crawler for indexing the pages its answers draw on. It respects robots.txt; Perplexity-User, which fetches pages on a user’s live request, generally does not. See how Perplexity chooses sources.

Presence rate

The share of prompts in a tracked set in which an AI engine mentions your brand. See measuring brand visibility in ChatGPT.

Prompt set

A fixed list of real buyer questions run against AI engines on a regular cadence to measure visibility over time. It is the foundation of AI visibility measurement.

R

RAG (retrieval-augmented generation)

A method in which an AI system retrieves relevant documents and uses them to generate its answer, rather than relying on training data alone. It is the mechanism behind cited AI answers.

Reranking

The step after retrieval in which candidate passages are scored again, usually by a more expensive model, and only the top few are passed to the answer. A page can be retrieved and still lose here. See retrieval vs citation.

Retrieval

The step in which an AI engine fetches candidate pages or passages to answer a question. Retrieval is necessary for a citation but not sufficient. See retrieval vs citation.

S

sameAs

A schema.org property that lists other URLs representing the same entity, such as official profiles and a Wikidata item, so engines can connect them. See the sameAs guide.

Schema.org

The shared vocabulary of types and properties used for structured data, used by Google, Bing, and other engines and maintained by an open community.

Server-side rendering

Generating a page’s HTML on the server so the full content is present in the initial response. It matters because many AI crawlers do not run JavaScript. See JavaScript rendering and AI crawlers.

Share of model

How often an AI engine names your brand compared with competitors across a prompt set: the AI equivalent of share of voice. See share of model.

Structured data

Machine-readable information added to a page, usually as JSON-LD using the schema.org vocabulary, that states facts explicitly instead of leaving engines to infer them. See structured data for AI search and Service schema.

T

Topic cluster

A pillar page covering a subject broadly, surrounded by interlinked pages that each answer one sub-question in depth. See topic clusters that AI engines cite.

V

Finding content by meaning rather than by exact words: the query and each passage are turned into embeddings, and the closest matches are returned. It is why a page can be retrieved for a question that shares none of its keywords. See content chunking for AI.

W

Wikidata

A free, structured knowledge base of entities and sourced facts, run by the Wikimedia Foundation and widely drawn on by search engines and AI systems. See Wikidata for brands.

Z

A search that ends without a click to any website, because the answer was shown on the results page or in an AI answer. See AI Overviews and the zero-click problem.

Using this AI search glossary

The vocabulary in this field is still settling, and several of these terms describe the same work under different names. What stays constant is the mechanism underneath: engines retrieve passages, resolve entities, and repeat facts they can corroborate. To put the terms to work, start with the GEO checklist; to understand where AI answers come from, read LLM SEO. If you want these concepts applied to your own brand, request an AI visibility audit.

Frequently asked questions

What is generative engine optimization (GEO)?

Generative engine optimization is the practice of getting a brand understood, trusted, and cited by AI engines such as ChatGPT, Perplexity, Gemini, and Google's AI Overviews and AI Mode. It is also called LLM SEO, LLMO, or AI SEO.

What is query fan-out?

Query fan-out is when an AI engine breaks one question into many related sub-queries, runs them in parallel, and combines the results into a single answer. It means pages can be retrieved through questions the user never typed.

What is RAG in AI search?

RAG, or retrieval-augmented generation, is a method in which an AI system retrieves relevant documents and uses them to write its answer instead of relying only on what it learned in training. It is the mechanism that lets AI answers cite sources.

What does grounding mean in AI search?

Grounding means tying an AI answer to specific retrieved sources rather than the model's memory alone. Grounded answers can cite where their facts came from; ungrounded answers cannot.

What is share of model?

Share of model measures how often an AI engine names your brand compared with competitors across a fixed set of prompts. It is the AI-search equivalent of share of voice.

OK

Olga Kunger

Founder & Lead Strategist, Ambeltek

Olga leads Ambeltek's web development, AI SEO, and GEO work — helping brands rank on Google and get cited by AI engines. More about Olga →

Want AI engines to cite your brand?

Our GEO program makes your site the source ChatGPT, Perplexity and Google AI quote. Growth tier from $9k.