Icarus Works
AI Search Glossary

The AI search glossary.

Every term you'll meet in AI search — 47 of them — defined in plain language, each with a one-line note on why it actually matters. Anchor-linkable, so you can point someone straight to a definition.

TL;DR

AI search has spawned a thicket of acronyms — AEO, GEO, AIO, RAG, E-E-A-T, llms.txt. This glossary cuts through it: crisp definitions plus why each one changes how you get found. Jump to any term below.

About this glossary

Plain definitions, honest about what's still moving.

AI search is young and shifting. Where a term's meaning is settled, we state it plainly; where the field is still forming (like llms.txt adoption), we say so instead of pretending certainty.

All 47 terms · A–Z
AEOAnswer Engine Optimization

The practice of structuring content so AI answer engines quote it directly as the answer to a question, not just rank it as a link.

Why it matters: AEO targets the citation itself — the moment your brand is named in the response.

AI OverviewsGoogle AI Overviews (formerly SGE)

Google's generative answer that appears above traditional results, summarizing multiple sources and linking to them.

Why it matters: It sits at the top of the highest-traffic search engine, so a citation there is highly visible.

AI share of voiceCitation share

The fraction of your target prompts, across engines, that return your brand in the answer — measured against competitors on the same prompt set.

Why it matters: It's the single most useful AI-visibility KPI because it's comparable and trackable over time.

AI user agentGPTBot, PerplexityBot, Google-Extended, ClaudeBot

The named crawlers AI companies use to fetch web content, each identified by its user-agent string.

Why it matters: If robots.txt blocks these agents, your content can't be retrieved or cited by that engine.

AIOAI Optimization

An umbrella term for optimizing a brand's presence across all AI-mediated discovery — AEO and GEO together, plus entity and platform work.

Why it matters: It frames AI visibility as one program rather than a set of disconnected tactics.

Answer engine

A system that returns a synthesized response to a question rather than a page of results. Answer engines quote and cite sources inside the answer.

Why it matters: Being the source an answer engine quotes is the new equivalent of ranking first.

Answer-firstInverted pyramid

Writing that leads with the direct answer, then adds supporting detail below.

Why it matters: Answer-first sections give engines a quotable statement in the first sentence.

Canonical URL

The single preferred URL for a piece of content when duplicates exist, declared with a canonical link tag.

Why it matters: It consolidates signals onto one URL so engines cite the right page, not a duplicate.

Chunking

Splitting a page into smaller passages so a retrieval system can index and return the most relevant part.

Why it matters: Self-contained, well-scoped sections are more likely to be retrieved cleanly as a chunk.

Citation

A reference to a source inside an AI answer — sometimes a linked footnote, sometimes just your brand named in the prose.

Why it matters: Citations are the conversion event of AI search: the buyer sees your name at the moment of the answer.

Citation position

Where in an answer your brand appears — named first, listed among several, or relegated to a caveat.

Why it matters: First-named sources tend to get disproportionate consideration from the buyer.

Corroboration

Independent, off-site confirmation of your facts — directories, reviews, press, and profiles that agree with your site.

Why it matters: Engines trust claims that multiple independent sources repeat.

DefinedTerm / DefinedTermSet

Schema types marking a defined term and the glossary it belongs to.

Why it matters: They tell engines 'this page authoritatively defines this concept' — useful for definitional queries.

Disambiguation

Resolving which specific entity a name refers to when several share it.

Why it matters: Clear category, location, and sameAs signals stop an engine confusing you with someone else.

E-E-A-TExperience, Expertise, Authoritativeness, Trust

Google's framework for content quality: demonstrated first-hand experience, subject expertise, recognized authority, and trustworthiness.

Why it matters: The same signals that satisfy E-E-A-T give AI engines reasons to treat you as a reliable source.

Embedding

A numeric vector representing the meaning of a piece of text, so similar meanings sit near each other in vector space.

Why it matters: Embeddings power semantic retrieval — your content is matched by meaning, not just keywords.

Entity

A distinct thing an engine can recognize and reason about — a business, person, place, or product — identified consistently across the web.

Why it matters: AI engines cite entities they're confident about; a clear, consistent entity is easier to name.

Entity consistencyNAP consistency

Presenting the same name, address, phone, and core facts about your business everywhere it appears online.

Why it matters: Inconsistent details make an engine uncertain which entity you are — and uncertainty suppresses citations.

Extractability

How easily an AI can lift a clean, correct answer from your page without needing surrounding context.

Why it matters: High extractability is the practical difference between being read and being quoted.

FAQPage schema

Structured data marking a list of questions and answers on a page.

Why it matters: It maps your content directly onto the question-answer shape engines look for.

Fine-tuning

Additional training that adapts a base model to a narrower task or style using a smaller, targeted dataset.

Why it matters: It's mostly a model-builder concern, but it explains why engine behaviour differs and shifts over time.

GEOGenerative Engine Optimization

Optimizing content so generative engines retrieve, surface, and cite it when composing an answer. GEO emphasizes being in the pool of sources an engine draws from.

Why it matters: You can't be cited if you're never retrieved; GEO addresses the retrieval layer.

Google Business ProfileGBP

A free Google listing of a local business's name, categories, hours, reviews, photos, and services.

Why it matters: It's a primary, structured, trusted source local AI answers draw on — keep it complete and accurate.

Grounding

Constraining an AI's answer to retrieved, verifiable sources rather than only its trained-in memory, so claims can be attributed.

Why it matters: Grounded answers cite sources — which means grounded engines are the ones you can influence with content.

Hallucination

When an AI states something false or fabricated with confidence.

Why it matters: Clear, structured, corroborated facts about your business reduce the chance an engine gets you wrong.

HowTo schema

Structured data describing a step-by-step process, its steps, tools, and supplies.

Why it matters: It signals a procedure engines can lift as an ordered answer to 'how do I…' prompts.

JSON-LD

The recommended format for embedding schema.org structured data as a JSON script in a page's HTML.

Why it matters: It's the format Google and most tools expect for reliable structured-data parsing.

Knowledge graph

A structured network of entities and the relationships between them (e.g., Google's Knowledge Graph).

Why it matters: Being a well-defined node in a knowledge graph reinforces that you're a real, citable entity.

LLMLarge Language Model

A neural network trained on large text corpora to predict and generate language — the engine behind ChatGPT, Claude, Gemini, and others.

Why it matters: Understanding that LLMs predict likely text explains why clear, conventional, well-structured writing is easier to quote.

llms.txt

A proposed plain-text file at a site's root that summarizes the site and points AI systems to its most important, citable pages.

Why it matters: It's a low-cost way to hand engines a clean map of what you want cited. (Adoption is still emerging.)

Local packMap pack

The map-plus-listings block Google shows for local queries.

Why it matters: The signals that win the local pack (proximity, relevance, prominence) overlap heavily with local AI citations.

LocalBusiness schema

Structured data describing a local business — name, address, phone, hours, area served, and service type.

Why it matters: It's the backbone of local AI answers: it states, in machine terms, who you are and where you serve.

Named recommendation

When an AI answer recommends specific brands by name (“try X or Y”) rather than describing a category generically.

Why it matters: Named recommendations drive action; being one of the 1–3 names is the goal.

Parametric answer

An answer generated purely from the model's trained-in weights, with no live retrieval.

Why it matters: Parametric answers are shaped by your long-run reputation across the web, not a single page edit.

Passage / snippet

A short, self-contained chunk of a page that an engine can retrieve and quote on its own.

Why it matters: Writing in coherent passages makes it easy for an engine to grab exactly the right piece.

Prompt setPrompt volume

The fixed list of buyer questions you test across engines to measure visibility repeatably.

Why it matters: A stable prompt set is what makes share-of-voice comparable month over month.

Query fan-out

When an engine expands one user question into several related sub-queries, retrieves for each, and synthesizes the results.

Why it matters: It means covering the whole cluster of related questions — not one keyword — raises your citation odds.

RAGRetrieval-Augmented Generation

An architecture that retrieves relevant documents at query time and feeds them to the language model so the answer is based on current, external content.

Why it matters: RAG is why fresh, well-structured web content can change what an AI says — even after training is finished.

Retrieval

The step where an engine fetches candidate sources for a query before generating an answer.

Why it matters: If your page isn't retrieved, it can't be cited; retrieval is the first gate.

Schema markupStructured data

Machine-readable annotations (usually via schema.org vocabulary) that label what content means.

Why it matters: Schema removes ambiguity about your facts, making them easier to extract and trust.

Speakable

Schema marking the sections of a page best suited to be read aloud by a voice assistant.

Why it matters: It nominates your cleanest answer blocks for spoken results.

Training data

The corpus of text a model learned from during training. It shapes what the model 'knows' by default.

Why it matters: Consistent, widespread mentions of your brand improve the odds it's represented in training data.

Zero-click

A search that's resolved on the results surface itself, with no click through to a website.

Why it matters: In AI answers, being named IS the win — the citation happens whether or not the user clicks.

Know the words. See where you stand.

A free AI visibility scan shows where the engines name you today.