Search Optimization

No Platform Shows You the Actual Customer Prompts: What AI Visibility Reporting Can Measure

By Omega Function 34 min read
Published by Omega Function · Reviewed by Omega Function Technical Review · Updated July 2026 · Review policy

At a Glance

No major answer engine currently provides site owners with a verified, complete list of the exact prompts individual users entered before a business appeared. Google provides impression reporting for supported generative Search features. Microsoft provides citation data and sampled grounding-query phrases through Bing Webmaster Tools. Both reports are useful and worth opening, but neither should be read as a complete record of actual customer prompts or of total audience exposure. Beyond those two, prompt-level figures generally come from a vendor running its own test prompts on a schedule, or from a third-party browsing panel scaled up with a model. Those methods can support a trend line. They do not produce a count. This article separates the tiers of AI measurement so you can tell which one you are looking at, and lays out what is genuinely trackable and genuinely improvable right now.

Who this is for Business owners and marketing leads being pitched AI visibility tracking, prompt-level rankings, or a share of voice score for answer engines
The claim to examine carefully Any report presenting the specific prompts real people used, a count of how many people saw your brand inside an AI answer, or a precise share of voice figure
Where the gap is Platforms publish impressions and citations, not verbatim prompts. Other figures are typically modeled from test prompts or browsing panels. Model answers also vary run to run, even at identical settings
What is actually logged Search Console generative AI impressions, Bing Webmaster Tools AI citations, AI assistant referral sessions in Analytics, AI crawler hits in your server logs, branded search trend, and what buyers tell you on your own form
What is actually improvable Crawler access, page structure, entity consistency, extractable answers, original data worth citing, and speed. The fundamentals, aimed at a new kind of reader
Honest setup time About a week to instrument, then a 90-day baseline before you draw a conclusion

Scope and methodology: This article evaluates publicly documented reporting capabilities and common measurement approaches available as of July 28, 2026. Platform features, classifications, crawler behavior, and vendor methodologies can change. Statements about third-party products describe published documentation or general measurement limitations, and are not assertions about any company’s intent, integrity, or contractual performance. Examples are illustrative unless otherwise stated.

There is a specific sales conversation happening in a lot of offices right now. A dashboard comes up on screen. It shows a list of prompts a customer might type into ChatGPT. Next to each one is a rank, or a percentage, or a little colored badge showing whether your brand was mentioned. There is a competitor column. There is a number labeled share of voice, usually carried out to one decimal place. It looks a great deal like the rank tracking reports the industry has been producing for fifteen years, which is part of why it feels familiar and therefore credible.

That resemblance deserves a second look. Rank tracking worked because a search engine returns a stable, orderable list of ten results, which can be checked from a known location and reproduced by anyone who checks it the same way. Answer engines behave differently, and the underlying data that would support prompt-level reporting at that resolution is not released to anyone outside the companies operating the systems.

We want to be clear about the spirit of this piece. There is nothing wrong with wanting to measure AI visibility. It is the right instinct. Traffic patterns are genuinely shifting and it is reasonable to want a number for it. The difficulty is that demand for a number arrived ahead of the infrastructure that would produce one, and a gap like that gets filled. Some of what fills it is careful, clearly labeled estimation. Some of it is a chart with a confident decimal point and no published methodology. Telling those apart is a skill, and it is learnable in about ten minutes. That is what the rest of this is for.

Start with the only question that matters: where would the data come from?

Before evaluating any AI visibility product, ask one question. Where does this number physically come from? Most AI visibility metrics currently trace back to one or more of four broad data sources.

Source one: the platform publishes it. Google’s generative AI performance report provides impression data for supported Google Search features. Microsoft’s AI Performance report in Bing Webmaster Tools provides citation activity, cited URLs, trends, and sampled grounding-query phrases across supported Microsoft AI experiences. These are first-party platform reports and both are worth using. Neither provides a complete list of verbatim user prompts or a verified count of every person exposed to a brand.

Source two: your own servers. Your logs and your analytics record what actually reached your property. This is first-party data, and it is badly underused.

Source three: the vendor asks the model itself. The vendor writes a list of prompts, runs them against ChatGPT or Gemini on a schedule from its own accounts, and records whether you were mentioned. This is a controlled test, not an observation of your customers.

Source four: a purchased browsing panel. The vendor licenses clickstream data from browser extensions and apps that some users have opted into, then scales that sample up to model broader behavior.

Many commercially available AI visibility products rely partly or primarily on synthetic prompt testing, third-party panel data, or a combination of the two. Their results can be useful when the methodology and limitations are disclosed. They should not automatically be read as first-party counts of actual customer behavior. Everything below is detail on why sources three and four resist precision, and how to get more out of one and two.

47%

of Search Console clicks come from queries Google withholds for privacy. Ahrefs measured 46.77 percent across 22 billion clicks and 887,534 properties. This is on the platform that shares the most.

Ahrefs, April 2025

8%

of Google visits with an AI summary ended in a click on a traditional result, against 15 percent without one. From a browsing panel of 900 US adults tracked through March 2025, not a global rate.

Pew Research Center

80

different completions came back from 1,000 identical requests to one model at temperature zero, the setting meant to make output repeatable. One model and one inference setup, but the mechanism is general.

Thinking Machines Lab

What the platforms actually give you, in full

It is worth being precise here, because the gap between what people assume is available and what is documented runs in both directions. More exists than it did a year ago. Less exists than a prompt-tracking dashboard implies.

Google Search Console

On June 3, 2026, Google began rolling out Search Generative AI performance reports, giving site owners a dedicated view of impressions inside AI Overviews and AI Mode. Note the phrasing: Google states it is rolling the report out to a subset of website owners to allow for testing before expanding it. Not every property has it, and a property can also come up empty because it does not have enough impressions in generative AI features. If you do not see the report, that is documented behavior rather than a problem with your site.

Where it is available, the report supports four dimensions: pages, countries, dates, and devices. The metric is impressions, defined as how many times links to your site were shown to a user in a generative AI feature. It does not currently include click data, click-through rate, or queries. The usual limitations of the Search performance report apply, including the 1,000 row limit.

One caveat about how to read it. The generative AI report draws on data from the Web search type in the main Performance report, so those impressions are part of your overall Search Console numbers rather than a separate pool sitting alongside them. The dedicated view isolates them for inspection. It does not turn your main performance report into a non-AI control group, and comparing the two as though one excluded the other will support conclusions the data does not.

Bing Webmaster Tools

In February 2026, Microsoft released the AI Performance report in Bing Webmaster Tools as a public preview, covering Microsoft Copilot, AI-generated summaries in Bing, and select partner integrations. This is first-party, verified-property reporting and it is currently the most detailed AI visibility data available to a site owner. It includes total citations displayed as sources in AI-generated answers, average cited pages per day, citation counts for specific URLs, and grounding queries, which Microsoft describes as the key phrases the AI used when retrieving content that was referenced in answers. Microsoft has since added mapping between grounding queries and the pages they cite.

If you have a Bing Webmaster Tools property and have not opened this report, that is worth doing this week. It is free, it is first-party, and it is the closest thing to an answer engine analytics product currently available.

The objection this raises, and the documented answer

Anyone who knows about the Bing report will reasonably ask whether it undercuts the premise of this article. It does not, and the reason is worth understanding precisely, because it is one of the most commonly confused points in this category.

A grounding query is not a prompt. When someone asks Copilot a question in their own words, the system reformulates that request into one or more retrieval phrases used to find source material. The grounding query is that retrieval phrase. It is generated by the system, not typed by the person. Differently worded questions can produce the same grounding query, and one conversational question can produce several. So the report indicates what the retrieval layer was searching for when it cited you, which is genuinely valuable, and it does not establish what your customer said.

Microsoft also states that grounding query data represents a sample of overall citation activity rather than a complete accounting. So the most detailed first-party AI reporting available today is, by the platform’s own documentation, sampled. That disclosure is exactly what good measurement documentation looks like, and it is a useful reference point when comparing against reports that present the same class of data to one decimal place without describing their sampling.

Everyone else

OpenAI, Anthropic, and Perplexity do not currently provide website owners with comparable verified-property reporting that reveals prompt-level audience behavior. There is no impression count, citation count, or query list available to you as a site owner. Referral traffic and independently conducted prompt testing can still supply useful directional evidence, subject to the limitations described below.

So the accurate summary is this. Two platforms publish first-party numbers, and neither publishes verbatim prompts. When a report includes prompt-level results, determine whether those prompts came from a fixed test panel, sampled grounding-query data, a third-party behavioral panel, or another documented source. They should not be labeled as verbatim customer prompts unless the provider can substantiate that characterization.

How the numbers are typically produced

Here is the part that tends to get compressed in a sales meeting. Both substitute methods rely on sampling, and sampling has properties that a single dashboard cell does not communicate.

Synthetic prompt panels

The most common approach: the vendor builds a list of prompts it believes your customers might use, runs them against several models on a schedule, and parses the answers for brand mentions. The output resembles rank tracking. The resemblance is presentational rather than structural.

A single scraped answer represents one written prompt, from one account, in one location, against one model version, at one moment. Real people arrive with a distribution of phrasings the vendor did not write, from accounts with their own history, in their own geography, on whichever model version the provider routed them to at the time. The vendor is sampling from a population whose shape is not observable, using a sampling frame it constructed. That is a structural property of the method rather than a flaw in any particular vendor’s execution.

It can still be useful. Running fifty consistent prompts monthly and watching your mention frequency climb from 12 percent to 30 percent over two quarters is meaningful directional signal. The methodological problem arises when the output is presented as a stable rank, a verified market share, or a count of actual people without sufficient qualification.

Purchased browsing panels

The other method comes from the traffic estimation tools the industry already uses. Semrush and Similarweb both publicly describe the use of panel, clickstream, partnership, or modeled data in portions of their traffic-estimation methodologies. Semrush describes clickstream partnerships covering a panel of over 200 million anonymized users. Similarweb runs its own browser extensions and integrates partner app data, then extrapolates.

Panel-based traffic estimates are legitimate modeled measurements. Their uncertainty can increase when the underlying number of observed visits is small, because fewer panelists happen to visit a given site. For low-volume or highly localized websites, estimates should therefore be compared against first-party analytics and presented as ranges or directional indicators rather than exact totals. The concern is not that modeled data lacks value. It is that modeled estimates should remain visibly distinct from a website’s observed first-party data.

Even the tools you already trust are sampling

This is not unique to AI. It is worth internalizing that the reports business owners already treat as ground truth do the same thing, with disclosure.

In Google Analytics, sampling occurs when a query exceeds the property quota. For standard properties that limit is 10 million events. Standard reports are generally unsampled, and the Explorations most agencies build custom reporting in will sample above that threshold. Google also documents narrower cases where standard reports can still sample, including filtering large datasets by country, which activates different processing on datasets above 100 million events. When sampling happens, Analytics indicates it with a data quality icon showing what percentage of your data the result was built from. The documentation is clear. The number of people who check that icon is small.

In Search Console, the effect is larger. Google’s documentation states that some queries are omitted from the report to protect user privacy, that these are called anonymized queries, and that they are included in chart totals unless a query filter is applied. It also states that due to internal limitations, Search Console stores and shows only the most important data rows, and that not all queries beyond anonymized queries are shown in the table. Ahrefs quantified the effect across 22 billion clicks and 887,534 properties in April 2025: 46.77 percent of clicks came from queries not visible in the report. The most common range for an individual site was between 45 and 80 percent.

Sit with that. On the most reliable, first-party, logged, platform-published data source in search, roughly half the query detail is withheld, and has been for years. That is the baseline environment for any claim of complete prompt-level visibility into platforms that publish less.

The second constraint: the same answer is not the same answer twice

Suppose a vendor obtained a perfectly representative set of real user prompts tomorrow. The numbers would still move, because the models do not return a stable answer.

Most people assume that setting temperature to zero makes a language model deterministic. It does not, and the reason is instructive. Researchers at Thinking Machines Lab traced production non-determinism to batch invariance: inference kernels produce slightly different numerical results depending on how many other requests are being processed alongside yours. Server load varies, so batch size varies, so output varies. As they put it, from an individual user’s perspective the other concurrent users are not an input to the system but a non-deterministic property of it.

Their experiment is worth citing because it is concrete. One thousand identical requests to a single model at temperature zero produced 80 unique completions. The first 102 tokens were identical every time. At token 103, 992 runs said one thing and 8 said another, and from there they diverged. That test used one specific model and inference configuration, so the exact figures are not a universal constant. The mechanism behind them is general to hosted inference.

On top of that baseline, consumer answer engines add layers designed to make answers differ between people:

  • Memory and personalization. ChatGPT can carry details across conversations, both memories a user saved deliberately and insights drawn from prior chats, and it uses them to shape responses. Two people asking the identical question can receive answers shaped by different histories.
  • Geography and account state. Location, language, subscription tier, and which product surface the question was asked in can all influence retrieval and output.
  • Model routing and version changes. Providers move users between model versions and ship updates without notice. A tracked prompt set can shift for reasons unrelated to your website.
  • Retrieval variance. When a model searches the live web mid-answer, it is subject to whatever the index returned in that instant.

Put these together and a prompt-level rank is better understood as one draw from a distribution than as the distribution itself. A movement in the score alone generally cannot establish why the result changed. Website work may be one cause. Sampling variance, personalization, model routing, geography, retrieval changes, or provider updates may also contribute. That is why a single metric should not carry an attribution argument by itself.

The third constraint: consent, referrers, and unattributed traffic

The measurement problem does not stop at the answer engine. It follows the visitor to your site, and this is where cookie consent enters, along with several factors people incorrectly attribute to consent.

What consent actually costs you

If you serve a consent banner, a share of visitors decline analytics storage. Consented visits are reported normally. Declined visits are not reported as ordinary observed users and sessions. Under advanced Consent Mode the Google tag can still send cookieless, non-identifying pings that feed modeling. Under basic Consent Mode the tag can remain blocked until consent is granted, in which case nothing is transmitted.

Google’s answer to the gap is behavioral modeling, which estimates the behavior of users who declined based on users who accepted. The eligibility thresholds are published and they matter most for smaller properties. A property needs at least 1,000 events per day with analytics storage denied for at least 7 days, plus at least 1,000 daily users with analytics storage granted for at least 7 of the previous 28 days. Google states that if there is not enough consented traffic to inform the model, events from users who decline consent are not reported.

Read that in the context of a business doing a few hundred sessions a day. Such a property is below the modeling threshold. Consented traffic is measured, declined traffic is not reported as observed sessions, and no modeled estimate arrives to fill the gap. AI referral numbers in that situation are not merely sampled. A portion is absent, and the proportion tends to be larger for smaller sites.

Factors that may cost you more than consent

Consent banners are unlikely to be the largest source of unattributed AI traffic. These typically matter more:

  • Missing referrer headers. AI assistants do not all pass a referrer reliably. Traffic that would otherwise read as ChatGPT can arrive as direct.
  • In-app browsers. A link tapped inside a mobile assistant app may open in an embedded browser that does not pass a referrer.
  • Copy and paste. A user who reads about you in an answer and then types your name into a browser bar is an AI-influenced visit with no technical trace linking it back. Its size cannot be measured from your own data, which is part of the point.
  • Changing classifications and incomplete historical attribution. Google Analytics introduced its native AI Assistant channel on May 13, 2026. Earlier traffic may not carry that default-channel label.

That last point needs care, because it is easy to overstate. Historical visits can sometimes be analyzed retroactively when GA4 captured recognizable source, medium, campaign, UTM, or referrer data, and Google’s documentation states that custom channel groups can be applied to your reports retroactively. Visits that arrived without usable identifying information cannot be reliably reconstructed, and any modeled reconstruction should be labeled as an estimate.

So the practical question for a proposal is sharper than a blanket objection. If a report includes AI referral history extending into 2025, ask which historical fields were captured, which classification rules were applied, and which portions are modeled. The absence of a native channel at the time does not mean the source data was absent. Neither does it guarantee that every AI-influenced visit can now be identified.

Sorting any claim into one of four buckets

This is the practical skill. Nearly every AI visibility claim falls into one of four categories, and once you can sort them, the conversation becomes straightforward.

Measurability Check

Which bucket does the claim fall into?

Pick the claim in front of you, or the number you want. The tool sorts it into logged data, sampled platform data, a modeled estimate, or something no platform currently publishes, and suggests what to ask for.

Select a claim below

Each option maps to how that number would have to be produced.

See your result ↑

What can actually be measured, in order of how much weight it carries

Now the constructive half. None of this is exotic. All of it is available to a business willing to instrument properly, and most of it is free.

Tier one: logged and first-party

Generative AI impressions in Search Console. Where the report has rolled out to your property, it gives impressions by page, country, device, and date. Treat it as a visibility index rather than a traffic source. The useful question it addresses is which of your pages Google is surfacing in AI features, and whether that set is growing. If a page you invested heavily in never appears, that is a content and structure signal worth acting on. Remember that these impressions are also part of your overall Web search data, so the main performance report is not a non-AI comparison group.

AI citations in Bing Webmaster Tools. Total citations, average cited pages per day, page-level citation counts, and grounding queries, across Copilot and Bing AI summaries. This is currently the richest first-party AI data available, and the grounding query view is the closest documented signal into retrieval intent. Use it to identify which pages already earn citations and what the retrieval layer was searching for when they did, while holding it as the sample Microsoft describes it to be.

AI assistant sessions in Analytics. With the native AI Assistants channel live, sessions arriving with a recognized assistant referrer are grouped automatically. This will undercount, for the reasons above, and an undercount is workable when you know the direction of the error. Watch the trend, and compare assistant referrals against organic search on conversion rate rather than volume alone. Some sites observe that assistant traffic converts at a different rate, possibly because the visitor arrives further along in their decision. Treat that as a hypothesis to test against your own conversion data rather than a general rule. Getting the GA4 configuration right is what makes any of it interpretable.

AI crawler activity in your server logs. Properly configured server, CDN, or edge logs can record requests from identifiable agents such as OAI-SearchBot, GPTBot, ClaudeBot, and PerplexityBot. You can see which URLs were requested during the retained logging period, how often they were requested, and which responses produced errors, blocks, redirects, or timeouts.

Several qualifications keep this accurate. Logs are first-party and generally unsampled when logging is enabled at all relevant layers, so completeness depends on your origin, CDN, firewall, retention policy, and configuration. The absence of a request does not by itself prove that a crawler intentionally ignored a page; it may not have been discovered, or not recrawled in your window, or handled at a layer you are not reading. User agents can also be spoofed, so verify by published IP ranges where the conclusion matters.

Two distinctions are worth getting right. OpenAI documents OAI-SearchBot as the agent used to surface websites in ChatGPT's search features, while GPTBot is documented as crawling content that may be used in training its foundation models. A GPTBot request is therefore not evidence that a page was retrieved to answer a particular question, and bot activity generally should not be treated as proof of answer-engine citation. Separately, Google-Extended is a robots.txt control token and does not appear as a distinct HTTP user agent in access logs; Google-related requests must be interpreted using the actual Google user agent and verified request information.

Within those limits this is still among the fastest problems in the category to identify. If a bot is receiving a 403 from your firewall or a timeout on a slow template, content work will not compensate for it, and an external dashboard will not surface it. We wrote a full AI crawler governance checklist covering how to audit and control this.

Branded search demand. When more people encounter your name in AI answers, some fraction may later search for you directly. Branded query volume and branded impressions in Search Console are logged and first-party. This is an indirect measurement, which leads some people to dismiss it. It is still one of the few available signals that reflects real human behavior at scale rather than a constructed prompt list.

Self-reported attribution. Add one optional field to your intake form asking how the person found you. Self-reported data is imprecise and skews toward whatever the form makes easy to select. It is also the only instrument positioned to capture the copy-and-paste visitor, the person who encountered your name in an AI answer weeks earlier, and the referral that arrived with no header. Given enough submissions the pattern becomes worth reading alongside your other indicators.

Tier two: modeled, directional, worth doing carefully

Running your own prompt panel is defensible when you run it as an experiment rather than a scoreboard. That means: a fixed prompt list written from real customer language rather than keyword tools, multiple runs of each prompt rather than one, results recorded as a mention rate across runs rather than a rank, model versions and locations documented, and a review cadence long enough that ordinary variance is not read as a trend. Report it as a percentage of runs with variance shown rather than as a position.

Run this way it addresses a real question: whether the corpus of information about your business is improving in a way models can find and reuse. Our generative engine optimization work is built around this framing.

Tier three: not currently published to site owners

The verbatim words real people typed. A verified count of how many people saw you. Precise share of voice. A stable rank for a prompt. Guaranteed placement in an answer. If a proposal presents these as observed measurements, ask for documentation of the source and method before relying on them.

What can actually be optimized

The encouraging part is that the levers are not mysterious. They are the same technical and editorial fundamentals that determine whether a machine can find, parse, and reuse your content. What changed is the reader, not the requirements.

Access

Can it reach you at all

  • AI user agents allowed or blocked deliberately, not by accident
  • No firewall or CDN rule unintentionally returning 403 to assistant crawlers
  • Server-rendered content, not text that only appears after JavaScript runs
  • Fast responses under crawl load, which is a page speed problem
  • Clean status codes and a sitemap that matches reality

Structure

Can it extract an answer

  • A direct answer near the top of the page, before the buildup
  • Headings that state the question a person would ask
  • Self-contained paragraphs that survive being quoted alone
  • Accurate structured data that agrees with the visible page
  • Tables and lists for anything comparative

Substance

Is there a reason to cite you

  • Consistent entity details everywhere your business is listed
  • First-hand data, pricing, or process nobody else has published
  • Specific service and location coverage rather than generic pages
  • Named review and update dates so freshness is verifiable
  • Mentions on sources models already draw from

Notice what is absent from that list: any tactic aimed at gaming a specific prompt. Optimizing for an individual prompt targets one sample from a distribution you cannot observe, on a system that may return a different answer to the next person who asks. Improving the corpus of accurate, structured, distinctive information about your business addresses the whole distribution instead. That approach is slower to show up in a chart, which is part of why the narrower one presents better in a demo.

Reading a proposal before you sign it

Here is the part we would want a family member to have before signing something. Work through the list below against any AI visibility proposal in front of you. Each item you check is a claim that should be substantiated before you rely on it.

Proposal Review

Count the claims that need substantiation

Check every statement that appears in the proposal, the dashboard demo, or the conversation. This is not an assessment of the provider. It is a checklist for which figures should come with documentation before you act on them.

0
Check the boxes that apply

A proposal with none of these checked is not necessarily the right one, but it is transparent about its evidence and limitations.

On the data itself

On methodology

On commitments

See your result ↑

If you are going to try one anyway

We know how this goes. The category is loud, the fear of being left behind is real, and a chart with your competitor's name on it is a powerful thing to look at. A fair number of people reading this will sign up for a prompt tracking tool in the next few months regardless, and that is a defensible way to learn something firsthand. So rather than argue, here is how to run the experiment so that it produces information.

Before you start, write down four numbers from sources you control: generative AI impressions in Search Console if the report has reached your property, citations in Bing Webmaster Tools, AI assistant sessions in analytics, and leads that mention AI on your intake form. Then set a calendar reminder for 90 days out. On that date, put the tool's headline score next to your own logged numbers and ask whether they moved together.

If the tool's score changes without corresponding movement in the available first-party indicators, treat that as evidence that the two measurements may not be closely aligned during the period. It does not prove that either measure is invalid, but it is a reason to examine the methodology before increasing investment. If they move together, you have learned something useful about your market.

A 90-day comparison provides an initial validation period. Longer observation may be necessary for low-volume businesses, seasonal markets, or metrics with delayed effects. Either way it costs nothing beyond the subscription you were going to buy, and it puts the question on evidence you control rather than on anyone's word, including ours.

The 90-day version of doing this properly

Days 1-30

Instrument

  • Check whether your property has the generative AI report yet. Record a baseline if so, note that it has not arrived if not
  • Verify your Bing Webmaster Tools property and open the AI Performance report
  • Confirm the AI Assistants channel is appearing in analytics
  • Turn on server log retention and start filtering identifiable AI user agents
  • Add a how did you find us field to the intake form

Days 31-60

Fix access and structure

  • Audit logs for AI crawlers hitting errors, timeouts, or blocks
  • Check that key content renders without JavaScript
  • Move the direct answer to the top of your top twenty pages
  • Reconcile structured data against visible page content
  • Fix entity inconsistencies across your listings and profiles

Days 61-90

Evaluate against the baseline

  • Compare every logged number against its documented starting value
  • Review which pages earn Bing citations and what grounding queries reached them
  • Segment assistant referrals by conversion rate, not volume alone
  • Publish something with first-hand data nobody else has
  • Decide the next quarter from logged movement, not from a single score

What we tell clients

The evidence-supported position is narrower than the market wants and more durable than it sounds. You can know whether machines can reach your content, whether they are requesting it, whether Google is surfacing it in AI features, whether Copilot is citing it and for which retrieval phrases, whether assistant referrals are arriving and converting, and whether people are searching for you by name more than they used to. That is a real dashboard, and most businesses are not looking at any of it.

What you cannot currently know is what your customer actually said. Not the retrieval phrase the system generated, but the words the person typed. The platforms may continue withholding much of this because of privacy, product design, competitive, or data governance considerations. Unless a platform documents its reasoning, privacy is best treated as a plausible factor rather than the established explanation.

What makes this workable is that the actions do not change once you accept the limitation. A business with accessible pages, accurate structured data, consistent entity information, and clear, distinctive content is generally better positioned to be discovered and reused by search and answer systems. These improvements can strengthen eligibility and machine readability. They do not guarantee inclusion, citations, rankings, traffic, leads, or revenue. Much of the technical work overlaps with established SEO and information-quality practice, although platform behavior and final results remain outside any website owner's control. The technical work is largely the same work, applied with a new kind of reader in mind.

What changes is what you agree to pay for. Ask for precision to be documented. Fund access, structure, substance, and measurement you can verify yourself.

If you have a proposal in front of you and you want a second opinion on which numbers in it can be sourced, send a message or a Loom and we will read it with you. No obligation.

Omega Function’s proposal reviews are technical and methodological assessments, not legal advice or determinations regarding contractual rights, enforceability, or breach.

Important limitations: AI visibility measurements, prompt-panel results, crawler activity, citations, impressions, referrals, branded-search changes, and self-reported attribution each measure different parts of the customer journey. They should not be interpreted individually as complete audience counts, verified market share, proof of causation, or guaranteed predictors of business performance. Omega Function does not guarantee rankings, inclusion in AI-generated answers, citations, traffic, conversions, revenue, or results from any third-party platform.

Frequently Asked Questions

Not verbatim prompts, and this is the most commonly confused point in the category right now. The report shows grounding queries, which Microsoft describes as the key phrases the AI used when retrieving content that was referenced in answers. Those phrases are generated by the retrieval system rather than typed by a person. A single conversational question can produce several of them, and differently worded questions can produce the same one. Microsoft also states that grounding query data represents a sample of citation activity rather than a complete record. It is genuinely useful for understanding what the retrieval layer was searching for when it cited you, and it should not be presented as a transcript of what your customer asked.

Some are, for the right purpose. A tool that runs a consistent prompt set on a schedule, discloses its methodology, reports mention rates across multiple runs rather than ranks, and labels its output as an estimate is doing legitimate directional research. That has value for tracking whether your presence in model outputs is changing over time. Tools that surface and organize the first-party Google and Microsoft data are also doing something useful. The issue to watch is the presentation layer, where sampled estimates can be rendered with the same visual weight as logged data. Ask which figures are observed and which are modeled, and how each was produced.

Google has not published a detailed rationale, so treat what follows as a plausible reading rather than an established explanation. Search queries are short and repeated by many people, which allows aggregation without identifying anyone. Even so, Google omits queries it considers too rare to display, a threshold commonly described as queries not issued by more than a few dozen users over a two to three month window, and Ahrefs found anonymized queries account for roughly 47 percent of all Search Console clicks. AI prompts are longer, more conversational, and more likely to be unique to one person, often carrying personal context. The share that could survive that kind of aggregation would be smaller, which is consistent with the generative AI report launching with impressions only. Privacy, product design, competitive, and data governance factors may all contribute.

It contributes, and it is probably not the largest factor. When visitors decline analytics storage, those visits are not reported as ordinary observed sessions. Under advanced Consent Mode the tag may still send cookieless pings for modeling; under basic Consent Mode it can remain blocked entirely. Google's behavioral modeling fills the gap only for properties above published thresholds: at least 1,000 events per day with analytics storage denied for seven days, plus at least 1,000 daily users with storage granted for seven of the previous 28 days. Below that, Google states the declined events are not reported. For many local businesses that leaves a gap with no estimate covering it. Missing referrer headers, in-app browsers, and users who read about you and then type your name in directly may each account for more unattributed traffic than the banner does.

Because the target is a distribution rather than a single outcome. The aim is to raise the probability that any given run mentions you, not to lock in one result. Non-determinism means a single check carries little information, and it does not mean the underlying odds cannot be influenced. A business whose information is accurate, consistent across sources, structurally easy to extract, and distinctive enough to be worth citing may improve its likelihood of appearing across repeated tests, compared with one whose information is thin or contradictory. That relationship cannot be verified by asking once, and results still depend on platform behavior outside your control.

Two things, and both are free. The first is your Bing Webmaster Tools AI Performance report, which many site owners have never opened because they set Bing aside years ago on market share grounds. The second is your server logs. They are first-party and generally unsampled where logging is enabled at every relevant layer, and each request from an identifiable AI crawler is recorded with a URL, a timestamp, and a response code. Sites sometimes discover through logs that a security rule, a bot filter, or a slow template is returning errors to systems they are actively trying to be visible in. Two limits apply: logs only cover what your configuration and retention actually capture, and a page that never appears was not necessarily refused, since it may simply never have been requested.

A guarantee of organic inclusion or citation should be treated as unsupported unless it is tied to a clearly documented paid placement product or another verifiable mechanism. There is no general submission process for organic citation, outputs vary run to run even with settings held constant, and providers ship model changes without notice. A responsible service commitment focuses instead on improving technical eligibility, content quality, and measurable conditions, while being explicit that the platform's final output is not within the provider's control.

That framing has the relationship backwards. The systems that answer questions in AI interfaces retrieve from indexes and live crawls that depend on the same crawlability, speed, structure, and quality signals traditional search depends on. A site that cannot be crawled efficiently is unlikely to be cited. A page that buries its answer under six paragraphs of preamble is generally harder to extract a clean quote from. Structured data that contradicts the visible page creates problems for both kinds of reader. Treating AI visibility as a replacement for technical fundamentals rather than an application of them tends to skip the layer everything else depends on.

Reporting has expanded faster in this category than many people realize. Microsoft went from no publisher-facing AI reporting to citations and grounding queries in early 2026 and has continued adding to it. Google has said it plans to add more metrics to the generative AI report over time based on site owner feedback. Click reporting would be a useful potential addition, but Google has not committed publicly to a specific metric or release timeline. Query-level data is a harder problem for the reasons above and should not be assumed. Building your reporting on what is documented today, and treating anything further as an upgrade, is the position least likely to require rework.

Sources

Platform reporting in this category changes quickly. This article was reviewed on July 28, 2026 against the sources listed above. Verify the current capabilities of the Search Console generative AI report, the Bing Webmaster Tools AI Performance report, and the Google Analytics AI Assistants channel before making a budget decision based on them. See our editorial review policy for how we source and update this material.

Corrections and updates: Platform capabilities in this category change quickly. Readers may report a suspected factual error through our contact page. Material corrections will be reflected in the article's updated date and source list.

Want this kind of insight applied to your stack?

Send a Message or Loom walking through your current setup and we'll come back with a scoped plan, not a sales pitch.

Get Started →