Research Report

AI Visibility in 2026: How Brands Earn Mentions and Citations

A comprehensive analysis of how generative engines discover, evaluate, and cite sources. Learn why traditional indexability no longer guarantees brand visibility.

SearchCombat Research Team

The transition from traditional search engine results pages to generative answers has fundamentally altered how digital visibility operates in 2026. For two decades, search engine optimization relied on a straightforward contract. You provide crawlable, relevant content, and search engines provide referral traffic via blue links. That contract is being rewritten entirely. Generative artificial intelligence engines such as Google AI Overviews, ChatGPT Search, Perplexity, and Bing Copilot now synthesize information directly. They do not merely rank pages. They read them, extract the factual substance, and generate a cohesive, natural language response natively within the user interface.

This evolution presents a critical challenge for brands and digital marketers alike. While the precise weighting mechanisms of generative models remain opaque, observational studies suggest that ranking well in traditional search environments does not always guarantee visibility within an AI generated answer. Marketers who assume their legacy search engine dominance will automatically translate into generative engine visibility are discovering a more complex reality. Navigating AI visibility requires exploring new measurement frameworks, testing novel content architectures, and observing how different large language models process the open web.

The Mechanics of Discovery and Query Rewriting

Generative engines do not process queries in the same manner as legacy lexical search engines. When a user submits a complex question, platforms like ChatGPT do not simply perform a direct keyword lookup against a static index. Instead, they utilize a sophisticated process known as query rewriting. According to official OpenAI documentation, ChatGPT frequently translates a conversational user prompt into one or more highly targeted sub queries before retrieving information from its search partners.

For example, a user might ask a broad question about the latest developments in targeted cancer treatments. The underlying system will actively rewrite this natural language prompt into a dense, academic keyword string focused specifically on immunotherapy developments in the current year. This rewriting process fundamentally changes how brands must approach keyword targeting. You can no longer optimize a page solely for the exact phrase the user types. You must optimize for the latent entity and the factual substance that the artificial intelligence system will hunt for during its retrieval phase. Maximizing content retrievability means ensuring your structural markup and factual density align perfectly with the reformulated machine query. Furthermore, these systems utilize contextual signals such as approximate location data derived directly from internet protocol addresses. If a user asks for nearby recommendations, the engine automatically appends geographic modifiers to the background search query. Earning visibility requires your content to satisfy these highly specific, machine generated retrieval requests.

The prevalence of these features is vast but remains highly dataset dependent. Ahrefs analyzed 55.8 million AI Overviews across a staggering 590 million searches, representing approximately 9.46 percent of their desktop keyword index. When weighted for search volume, this equates to 23 billion out of 180 billion monthly searches, or 12.8 percent. The rollout and prevalence of these features vary significantly by geography. For example, Ahrefs estimated incidence rates of 12.5 percent in the United Kingdom, 15.5 percent in Brazil, and 16.5 percent in India. These figures represent a snapshot of a specific product index and omit low volume queries entirely. Marketers must understand that AI overview incidence cannot be summarized by a single global percentage. It is a highly volatile landscape.

Crawl Access and the Indexability Prerequisite

Before a brand can earn an AI citation, it must first be technically accessible. AI features rely heavily on the foundational search index. According to Google Search Central documentation, pages must meet all standard technical requirements to be eligible for inclusion in AI Overviews and AI Mode. This means the content must be crawled, indexed, and eligible to display a standard search snippet. There are no secret, hidden technical requirements specific to generative features. If a page is blocked from standard indexing, it is entirely invisible to the generative engine.

Website owners retain absolute control over this access. Google confirms that standard robots directives dictate crawl behaviour exactly as they always have. To limit the information shown from your pages in search, webmasters can utilize standard meta tags such as nosnippet, data-nosnippet, and max-snippet. Additionally, publishers can limit AI training and grounding in Google systems by restricting the Google-Extended crawler. Similarly, OpenAI explicitly states that websites must allow the OAI-Searchbot to crawl their site to be eligible for inclusion in ChatGPT search results. Bing respects all content owner preferences expressed through standard control mechanisms. Technical accessibility is the mandatory first step. Without a clear path for crawlers, all subsequent optimization efforts are entirely futile.

The Ghost Citation Phenomenon

Distribution of 3,981 domain appearances across AI search responses.

Ghost Citations (Cited as source, brand not mentioned)61.7%
Brand Mentions (Brand mentioned, no citation link)25.1%
Both Cited and Mentioned13.2%

Source: Semrush AI Visibility Toolkit Study, June 2026. Based on 115 prompts across 14 countries.

The Ghost Citation Phenomenon

Being cited by an AI search engine is widely considered a powerful authority signal. It serves as a strong indicator that your content was selected by the underlying model for reference. However, the critical realization of 2026 is that being cited is entirely distinct from being mentioned. Industry expert Kevin Indig coined the term ghost citation to describe this exact gap. A ghost citation occurs when an artificial intelligence system uses your website as a source link, but never explicitly mentions your brand name within the generated text. The reader consumes your insights, but walks away without recognizing your brand or attributing the expertise to your organization.

To investigate this phenomenon thoroughly, Semrush analyzed 3,981 domain appearances across 115 prompts in 14 different countries. The findings are stark and highly illuminating. A massive 61.7 percent of AI citations are ghost citations. In these cases, the website receives a source link, but the brand name is completely absent from the answer text. Only 13.2 percent of appearances resulted in both a citation link and a textual brand mention. Interestingly, 25.1 percent of appearances were brand mentions without any accompanying citation link. This means that 74.9 percent of all brand appearances included a citation, but only 38.3 percent included a textual brand mention. The citation rate is nearly double the mention rate.

This data illustrates that appearing as a source does not automatically make your brand visible to users. A citation offers technical attribution, but your brand identity remains hidden from the actual narrative. Publishers and news aggregators may be satisfied with accumulating citation links. However, consumer brands and business to business service providers require textual mentions to build trust and awareness within their target market.

Platform Divergence: Gemini versus ChatGPT

Brand mention rates versus source citation rates across leading artificial intelligence models.

Google Gemini

Brand Mention Rate83.7%
Source Citation Rate21.4%

Behaves conversationally. Generates high textual recall but low explicit linking.

ChatGPT

Source Citation Rate87.0%
Brand Mention Rate20.7%

Behaves academically. Generates high explicit linking but low textual brand inclusion.

Divergent Behaviours Across AI Engines

The Semrush study also revealed that different AI engines treat citations and mentions in fundamentally different ways. You cannot assume that visibility in one platform translates to visibility in another. For instance, when a brand appears in a Google Gemini answer, it is named in the text 83.7 percent of the time. However, Gemini only generates a citation link 21.4 percent of the time. Gemini behaves much like a conversationalist drawing on internal knowledge and established brand entities.

ChatGPT exhibits the exact opposite behavior. It cites brands as sources 87 percent of the time, but mentions those brands in the text only 20.7 percent of the time. ChatGPT constructs answers that resemble academic papers laden with detailed footnotes. Google AI Overviews sit somewhere in the middle, though they lean heavily toward citations over mentions. Google AI Mode, meanwhile, mentions brands at nearly twice the rate of ChatGPT, but still acts more like a footnoted research piece than Gemini does.

Because these engines reward different signals, formats, and sources, there is almost no overlap between the brands ChatGPT cites and the brands Gemini names for the exact same prompt. Aggregator and academic websites tend to accumulate citations, while consumer brands with strong public identities tend to accumulate textual mentions. For example, major technology brands were named in AI answers far more often than they were cited as sources. Conversely, publishing platforms were cited frequently but almost never mentioned in the text. Accurate measurement requires tracking each engine independently.

Query Intent and the Content Type Connection

Query phrasing and intent dramatically alter brand mention outcomes. In the Semrush dataset, short and conversational queries produced mention rates of nearly 100 percent. Long, highly structured prompts resulted in mention rates of just two to three percent. That represents a thirty to fifty times difference on the exact same topic. A short query naturally leads to brand mentions, whereas a long query with extensive context triggers more technical citations but almost zero brand mentions. Brand visibility is highly sensitive to exact phrasing and user intent.

Content type also correlates strongly with these rates. Informational queries, such as those asking for definitions or explanations, earned an 89.3 percent citation rate but a mere 18 percent mention rate. Comparative queries, which ask the AI to evaluate options or recommend the best solution, resulted in a 43.3 percent mention rate. This is 2.4 times higher than the mention rate for purely informational content. How to queries achieved a 42.8 percent mention rate, while commercial queries with buying intent saw a 35.6 percent mention rate alongside an 84.4 percent citation rate.

These observed associations suggest a testable recommendation: because comparative content often involves evaluating named entities, it may provide more natural opportunities for an AI to include brand names in its response. Marketers seeking to increase brand mentions could test investing in comparative and evaluative formats, such as comprehensive competitor analyses, while tracking whether these formats yield a higher rate of textual mentions compared to basic glossaries.

Topic Ownership in ChatGPT

Analysis of 1,094 United States subject categories, measuring brand consistency across five related buyer prompts per category.

53.7%
Unsettled Categories

No single brand appears consistently across even three of the five prompts. The topic remains completely wide open to competitors.

31.2%
Emerging Leaders

A brand appears in at least three prompts but lacks a definitive five percentage point lead over the immediate runner up.

15.2%
Clear Category Owners

A single brand dominates four or more prompts with a commanding lead. Ninety percent of these brands retain their top position month over month.

Topic Ownership and the New Competitive Moat

Winning a single prompt in ChatGPT does not mean you own the topic. Brand visibility in AI search shifts rapidly from prompt to prompt. Two related questions in the same buyer research journey can surface completely different brands. To understand how brands win and lose visibility, Semrush analyzed 1,094 subject categories in ChatGPT. A category was defined as a cluster of five representative prompts covering definition, comparison, alternatives, use cases, and purchasing questions.

The results show that most topics are still highly contestable. Only 15.2 percent of the 1,094 categories had a clear brand owner. A clear owner was defined as a brand appearing in at least four out of the five prompts with a minimum five percentage point lead over the immediate runner up. Another 31.2 percent of categories had an emerging leader, meaning a brand appeared in at least three prompts but lacked the required margin of dominance. The remaining 53.7 percent of categories were completely unsettled, with no single brand appearing consistently. The largest topics by search volume were actually the least likely to have a clear winner, presenting massive opportunities for focused optimization.

Interestingly, broad search engine optimization metrics do not consistently predict topic ownership in ChatGPT. When comparing the category owner against the runner up, traditional metrics like Authority Score and organic traffic correlated with ownership only about half the time. Branded search volume offered a modest edge, aligning with ownership in 55.7 percent of pairs. While strong traditional optimization helps a brand enter the consideration set, building deep, comprehensive coverage of the specific topic is recommended to foster category consistency. Furthermore, leadership changes hands frequently when margins are narrow. However, owners with a five percentage point lead retained their first place position in 90 percent of month over month comparisons. In the realm of AI visibility, margin matters significantly more than mere momentum.

The Economics of Crawling Versus Referral

The fundamental economics of content publishing are facing a severe stress test. Historically, search engine crawlers scanned content a few times and sent substantial visitor traffic in return. The AI delivery model disrupts this equation entirely. According to data published by Cloudflare on its AI Insights Radar, large language model crawlers consume vast amounts of content while returning very little referral traffic to publishers. Cloudflare calculates a specific ratio by dividing the total number of page requests from a platform crawler by the total number of client requests referred back by that exact platform.

The data from June 2025 reveals staggering disparities. Anthropic Claude made nearly 71,000 page requests for every single page referral it provided. In stark contrast, Mistral sent ten times as many referrals as crawl requests, resulting in a fractional ratio. While referral traffic from native applications often lacks referral headers, which might overstate the deficit slightly, the underlying trend is undeniable. These models continually consume more content, at higher frequencies, without proportionally increasing the traffic they send back to the source. The era of receiving equitable visitor traffic in exchange for crawl access is ending rapidly. Content providers must now audit their server logs, review their crawl to refer ratios, and make highly calculated decisions about which bots provide sufficient value to justify access.

The Pew Research Center conducted an observational study of 900 United States adults using tracked devices throughout March 2025. This research revealed striking shifts in user behaviour when an AI summary is present on the screen. In their sample, 18 percent of searches produced an AI summary. When a summary appeared, only 8 percent of visits resulted in a click on a traditional search result link. This is a significant drop compared to the 15 percent click rate observed when no summary was present. Even more telling is the engagement with the AI citations themselves. Only 1 percent of users clicked on a link embedded within the AI summary. Furthermore, 26 percent of sessions ended without any click whatsoever when a summary was present, compared to 16 percent for traditional results. This data highlights a massive and growing trend of zero click resolution.

Stop guessing about your AI visibility.

Run a complete diagnostic on your brand. Discover exactly how ChatGPT, Perplexity, and Google AI view your domain, and uncover the ghost citations you are currently missing.

Get My Free AI Visibility Audit

Structured Data and Entity Clarity

Google Search Central explicitly states that structured data helps their systems understand the exact meaning of page elements. Implementing comprehensive JSON-LD markup offers search systems a clear structure for parsing the core subject, the author, the organization, and the claims being made on a given page.

By clearly defining entities, products, organizations, and authors, you provide search systems with structured, machine readable facts. Official documentation states that structured data helps search engines understand the exact meaning of page elements. While it is not a proven mechanism for earning AI citations or forcing a language model to associate facts with your brand, clearly identifying your brand via organizational schema is a highly recommended practice for clarifying brand details and page meaning across the internet.

Local SEO and Aggregate Sentiment

Local business visibility requires a distinctly localized approach within AI models. The BrightLocal Local Consumer Review Survey from early 2025 revealed that a mere four percent of respondents never read online business reviews. Reviews remain a critical trust signal for both consumers and generative models alike. When an AI evaluates local entities, it relies heavily on aggregated sentiment across the web. The survey indicated that 40 percent of consumers preferred email requests for reviews, while 27 percent favoured in person requests. Furthermore, 48 percent expected to be asked for a food or drink review by the very next day.

These statistics represent consumer expectations rather than direct ranking factors. Active reputation management through consistent review generation remains a core practice for building local prominence. Prompting a language model to find the best local service provider frequently surfaces results associated with aggregate review scores and localized citations, although Google simply notes that positive ratings can help local ranking, rather than guaranteeing it as an absolute weighting mechanism.

Measurement and Attribution in a Zero Click World

Measurement remains one of the greatest challenges in AI visibility. Fortunately, platforms are beginning to offer native telemetry to address this deficit. Bing Webmaster Tools recently introduced an AI Performance dashboard in public preview. This vital tool provides publishers with crucial insights into how their content appears across Microsoft Copilot and AI generated summaries in Bing. Furthermore, prompt panel measurement methodologies allow businesses to manually benchmark their brand inclusion within specific interface environments. These tools mark an important step toward greater transparency between AI systems and the open web.

The dashboard tracks several vital metrics. It records Total Citations, which shows exactly how often your content is referenced by AI systems over a selected time frame. It tracks Average Cited Pages to highlight overall citation patterns. Crucially, it reveals Grounding Queries. Grounding queries are the precise key phrases the AI used when retrieving your content to formulate its answer. This data allows webmasters to validate which pages are serving as AI references and to spot immediate opportunities for structural improvements on pages that are indexed but routinely ignored. Bing also strongly advocates for the use of IndexNow, an initiative that notifies participating search engines immediately when content is added or updated. Freshness appears to be a highly beneficial factor for AI inclusion.

The Implementation Framework for Brand Visibility

How do brands operationalize this vast amount of data? Earning visibility in 2026 requires a deliberate transition from isolated keyword optimization to comprehensive entity and topic optimization. The following implementation framework translates our research findings into an actionable strategic roadmap.

First, secure crawl access and verify indexability relentlessly. AI features rely entirely on the foundational search index. Pages must be technically accessible, fully rendered, and eligible for rich snippets. Ensure your robots directives allow the necessary crawlers, including Googlebot and OAI-Searchbot, unless you have made a highly strategic decision to block them based on poor crawl to referral ratios.

Second, consider testing comparative and evaluative formats. As the Semrush data indicates, informational content is associated with invisible ghost citations, while comparative content correlates with textual brand mentions. Building detailed product comparisons, alternative guides, and robust use case analyses provides a natural context for your brand to be evaluated alongside competitors.

Third, pursue total topic ownership rather than isolated prompt rankings. Identify the core categories your buyers research extensively. Build comprehensive content clusters that answer every variant of the defining questions within that category. Your goal is not to win one single query, but to appear in four out of five related prompts, establishing a definitive and unassailable lead over your competitors.

Fourth, implement comprehensive structured data markup. Use robust JSON-LD formats to clearly define your business entities, providing search engines with explicit, machine readable facts about your organization. Finally, consider testing strategies that increase your brand mentions in third party environments. Because AI models synthesize answers from across the open web, observational data suggests that earning mentions on authoritative forums, review platforms, and industry publications correlates with a higher likelihood of an engine including your brand in its generated response.

Methodology and Data Limitations

Transparency in data sourcing is vital for accurate strategic planning. The insights presented in this guide are derived from multiple distinct industry studies, each with specific methodologies and inherent limitations. The Pew Research Center data is based on an observational study of 900 United States adults using tracked devices in March 2025. It provides a behavioural snapshot of specific users and should not be interpreted as an absolute universal traffic decline metric. The Ahrefs data relies on 55.8 million AI Overviews across 590 million searches within their desktop keyword index. It omits low volume queries and logged out mobile sessions. Ahrefs explicitly notes that rollout varies heavily by country and their methodology likely undercounts total prevalence.

The Semrush Ghost Citations and Topic Authority studies are based on prompt tracking across specific domains and categories in multiple countries. Visibility in AI search is highly sensitive to exact prompt phrasing, and results vary significantly between ChatGPT, Gemini, and Google AI Overviews. Finally, the Cloudflare crawl and referral ratios are aggregated from network telemetry, and referral counts may omit native application traffic that strips referrer headers. Readers should apply these findings as directional trends to inform strategy rather than absolute universal guarantees.

Source Footnotes