What Entity SEO Means in the AI Era
Entity SEO is the practice of establishing a brand, person, or product as a uniquely identifiable entity that search and AI engines recognize independent of wording. Classic SEO optimized pages to rank among the ten blue links; entity SEO optimizes the machine's record of who you are, which AI engines consult before deciding whom to mention or cite.
Google set this direction in 2012 with the Knowledge Graph and the phrase that still explains it best: things, not strings. A string is a sequence of characters that may or may not mean anything. A thing is a record with an identifier, a type, and relationships to other records. Entity based SEO is the work of making sure your brand exists as a thing.
Two terms need to be exact before anything else in this post makes sense, because they are used interchangeably in the wild and they measure different behavior.
A mention is your brand named in the text of an AI answer, with no link attached. A citation is a linked source the engine used to build that answer. An engine can mention you without citing you, and it can cite your page without ever naming your brand. The two move independently, which is why they need separate tracking.
Entity work does not replace ranking work. Ranking pages remain some of the strongest citation candidates an engine has, and organic position still earns the clicks that still happen. The goal is additive: keep ranking, and make your entity legible at the same time. If generative engine optimization is new territory, start there. For the markup layer specifically, Structured Data for AI Search is the implementation companion to everything below.
How Knowledge Graphs Ground AI Answers
A knowledge graph is a database of entities and the typed relationships between them. Google's is reported to hold more than 1.6 trillion facts about 54 billion entities, and it grounds AI Overviews, AI Mode, and Gemini. When an AI engine answers, it first resolves the entities in the query, then retrieves and cites sources that match those entities.
Those figures come from industry documentation of Google's numbers as of mid-2024; Google's own last public statement, in May 2020, put it at 500 billion facts about 5 billion entities. Either way, the direction of travel since then is as interesting as the scale. In June 2025 Google pruned more than three billion entities from the graph, explicitly to prioritize quality for AI applications. Google is actively curating the graph for AI grounding, which makes it live infrastructure for the answers being generated today rather than a leftover from 2012 era search.
Grounding, in the technical sense, is what separates a generative answer from a guess. In retrieval-augmented generation, the engine retrieves documents, resolves them against known entities, and then generates an answer that cites what it retrieved. The sequence is the whole argument of this post: entity resolution happens before citation selection. An entity the engine cannot confidently resolve never reaches the stage where sources get chosen. That is the mechanism behind "how does ChatGPT know about my brand," and the answer is usually that it knows a record, not a website.
You can check your own record today. Google's Knowledge Graph Search API returns entities as JSON-LD with machine IDs and a resultScore confidence ranking, which means any brand can query how confidently Google recognizes it, right now, without a tool.
Wikidata is the other registry worth knowing. It holds somewhere in the range of 90 to 115 million items depending on which count you use, and it is open to contribution in a way Wikipedia is not.
Wikipedia itself occupies an odd position in 2026 and deserves a nuanced read rather than a rule. The Wikimedia Foundation reported in October 2025 that human pageviews were down about 8% year over year, attributing the decline to AI answering questions directly, while noting that nearly all major language models train on Wikipedia. Its share of visible citations can be small (Semrush found Google AI Mode citing it in only about 2% of responses) while its role as training substrate and entity registry stays large. Presence there still helps engines resolve you, even when it rarely shows up in a citation list.
Disambiguation and the Entity Home
Entity disambiguation is how an engine decides which entity a name refers to: the word seal can mean an animal, a Navy unit, an emblem, a mechanical part, or a musician. Brands resolve ambiguity by maintaining an entity home, one canonical page whose facts every other profile on the web corroborates.
Two definitions carry this section.
An entity home is the single canonical page, usually your About page or homepage, that engines treat as the source of truth for your entity. SameAs reconciliation is the practice of explicitly linking every representation of that entity, your site markup, Wikidata, LinkedIn, Crunchbase, Google Business Profile, so engines merge the scattered records into one confident one.
The sequencing matters more than any individual tactic. A workable order:
- Check your current status through the Knowledge Graph Search API before doing anything else. You may already have a record with facts you did not write.
- Earn coverage in authoritative publications. Corroboration from sources the engine already trusts is what raises confidence in a record.
- Add Organization schema with sameAs links on your entity home, pointing at every profile that represents you.
- Create a Wikidata entry. It is often more achievable than Wikipedia, and it feeds the same reconciliation process.
- Earn a Wikipedia article where notability genuinely supports one, and never before then.
Consistency runs through all five. Conflicting founding dates, product names, or descriptions across profiles are what keep a record ambiguous. The mechanics of @id and sameAs implementation live in our structured data guide rather than here.
The cost of leaving ambiguity in place is not theoretical. Columbia's Tow Center tested eight AI search engines across 1,600 queries in March 2025 and found incorrect answers, mostly misattributed sources, in more than 60% of them. Weak or conflicting entity signals are how brands inherit someone else's facts, or lose credit for their own.
This work also travels across engines. Microsoft's Fabrice Canel confirmed at SMX in March 2025 that Bing and Copilot's language models use schema markup to understand content, and that fresh content pushed through IndexNow carries weight. One entity strategy serves several engines at once.
A practical method for finding what is missing comes from Search Engine Land's July 2026 entity gap work: build the ideal model of your entity and its relationships, compare it against what your content actually asserts today, then prioritize the gaps by business value. Connected entity graphs, where the organization links to its products and its people, consistently outperform loose collections of facts.
The Citation Layer: Where AI Actually Pulls Sources
The citation layer is the recurring set of third-party sources an AI engine cites for a category: community sites, reference sites, video, reviews, and comparison pages. Semrush analyzed more than 100 million citations and found each engine builds a different layer, and the mix shifts over time, so brands need per-engine visibility.
Semrush's November 2025 study, drawn from more than 230,000 prompts and over 100 million citations, shows how unstable that layer is. Before September 2025, ChatGPT cited Reddit in roughly 60% of responses and Wikipedia in about 55%. Afterward those figures fell to roughly 10% and under 20% respectively, a deliberate rebalancing away from over-cited domains. Google AI Mode cites Wikipedia in only about 2% of responses. LinkedIn is among the most stable sources at around 15%. Perplexity's top sources skew toward Reddit, LinkedIn, NIH, Microsoft, and Google.
AI Mode is worth pausing on. It is a separate conversational search surface rather than a feature inside the results page, and it builds its own citation layer with its own preferences. Google AI Mode Explained covers it properly; here it matters as evidence that per-engine measurement is not optional.
Content type shapes citations as much as domain does. Search Engine Journal's coverage of a 768,000-citation study found product-related content, meaning "best of" lists, comparisons, and product pages, accounting for up to 70% of citations on bottom-of-funnel queries, while blog content ran 3 to 6% and press releases under 2%.
The correlation data points the same direction, and it reframes the problem. Victorious tested 175 brands across five verticals in July 2026 and found that recognition and mention are nearly separate problems: 96% of brands were described accurately when an engine was asked about them directly, yet 89% never appeared in answers when they were not named. Third-party web mentions correlated with appearing at 0.45 and referring domains at 0.49, close enough that neither signal dominates. The blunt number is the useful one: brands with fewer than 2,000 indexed pages mentioning them were named in AI answers just 3% of the time. Off-site presence is the input, whatever form it takes.
Academic work supports the tactical read. The GEO paper presented at KDD 2024 found that adding citations, quotations, and statistics to content lifted visibility in generative responses by up to 40%.
The strategic conclusion is a modest one. In that same research, 99.99% of the 49,391 citations analyzed pointed at third-party sites rather than the brand's own domain. You cannot become the citation layer for your category. What you can do is be accurately represented across it: your entity consistent everywhere the engines look, your products present in the comparison content they lift, your name attached to the claims you want made about you.
Measuring Entity Strength with Gist GEO
Entity work shows up in measurement as two separate columns: mentions, where AI names your brand in answer text, and citations, where it links your pages as sources. Gist GEO tracks both across engines, with Share of Voice for mentions, Share of Citations for sources, and a Top sources view showing Found versus Cited side by side.
Alongside those, Share of Found Links and Earned Media Score track whether your pages are surfacing and whether earned coverage is doing its job. The Top sources view is the closest thing to seeing your category's citation layer directly, with Found and Cited columns side by side.
Two diagnostic patterns turn those numbers into decisions:
- Mentioned but never cited. The entity is recognized and your content is not being retrieved or lifted. That is a content-shape problem, not an identity problem: the record is fine, the pages are not extractable.
- Cited on one engine, invisible on another. The per-engine citation layers have diverged, exactly as the Semrush data predicts. Work the missing engine's layer specifically rather than doing more of what already works elsewhere.
Run a Gist GEO audit to see whether AI engines recognize your brand, and whether they cite you or merely talk about you. For the broader product picture, see What Is Gist GEO?; for metric methodology, How to Measure AI Visibility.



