Bing Webmaster Tools Is Quietly Getting Better at Measuring AI Search

Bing Webmaster Tools Is Quietly Getting Better at Measuring AI Search
Sponsored

For much of the generative-search boom, publishers have been asked to optimize for something they could barely measure. A page might be cited by an AI assistant, used to ground an answer or disappear entirely from the generated response, but website owners often had to reconstruct what was happening through manual testing, third-party monitoring and inference.

Bing Webmaster Tools is beginning to change that. Microsoft’s AI Performance report, introduced in public preview earlier this year and expanded with additional analytical capabilities, gives verified site owners a first-party view of how their content participates in supported AI-generated answers. The result is not yet the equivalent of mature web analytics for AI, but it represents something the industry has badly needed: a measurement layer designed specifically for generative search.

From rankings and clicks to citations and grounding

Traditional search analytics are built around a familiar sequence. A user enters a query, a page receives an impression, the user may click, and the publisher can measure the resulting visit. Generative search breaks that sequence because content can influence an answer without producing a conventional search-result impression or a click.

Microsoft’s AI Performance report starts from that new reality. According to Bing’s official documentation, it measures visible citations across Microsoft Copilot, AI-generated summaries in Bing and select partner AI integrations. Publishers can see total citations, the average number of unique cited pages, page-level citation activity, trends over time and the grounding queries associated with their content.

Grounding queries are particularly important because they offer a window into the retrieval process behind an AI answer. They are not the user’s complete prompt. Bing describes them as grouped phrases representing the concepts used when retrieving content that was subsequently cited. That makes them a different kind of keyword data: less a transcript of what a person typed and more an indication of what the AI system believed it needed to retrieve.

Bing is adding context, not just counts

The newer capabilities make the report more interesting than a citation counter. Bing is previewing Intents, Topics, Citation Share and Compare, each designed to answer a different question about AI visibility.

Intents classify grounding queries according to broader purposes such as informational, commercial, navigational, research, creation and local intent. Topics group related queries into larger themes, allowing publishers to examine AI visibility around subject areas rather than individual phrases. Compare overlays performance from different periods so changes can be examined over time.

Citation Share may be the most strategically useful addition. Instead of showing only how many times a site was cited, it estimates the percentage of citations attributed to that site for a particular grounding query. A publisher can therefore begin to distinguish between being present occasionally and occupying a meaningful share of the citation space around a topic.

Microsoft is careful about what this number means. Citation Share is not a ranking, traffic metric or quality score, and it does not reveal which competing domains hold the remaining share. Changes can reflect shifts in user demand, content across the web, model behavior, freshness and other factors. That caution is important because AI-search metrics are still easy to overinterpret.

For the first time, GEO can become more empirical

The rise of Generative Engine Optimization, Answer Engine Optimization and similar labels has produced an industry full of recommendations that are often difficult to validate. Publishers are told to structure content clearly, demonstrate expertise, answer questions directly, keep information fresh and make entities unambiguous. Much of that is sensible content practice regardless of AI, but without first-party visibility data it has been difficult to determine whether specific changes correspond with greater participation in generated answers.

Bing’s reporting does not solve causality. Its own documentation explicitly warns that a rise or fall in citations cannot be attributed automatically to a content update, model change or any single event. The data is aggregated and representative rather than a complete log of every AI citation. Even so, measurement changes the conversation.

A publisher can now identify which pages are repeatedly cited, which grounding themes are associated with those pages and whether citation share changes over time. It can compare editorial sections, observe whether fresh content begins to appear around relevant topics and investigate why pages that perform strongly in conventional search may have limited visibility in AI answers. GEO starts to look less like a collection of anecdotes and more like a field in which hypotheses can at least be tested against first-party signals.

The Bing-Google contrast is philosophical

The most interesting part of Microsoft’s approach may not be any individual metric. It is the philosophy implied by exposing the data.

Google remains vastly more important to most publishers in terms of traditional search volume, and its Search Console is foundational to modern SEO. But as Google integrates AI Overviews and AI Mode into Search, publishers have repeatedly wanted a cleaner way to isolate and understand their performance inside those generative experiences. The industry’s frustration is not simply that AI exists; it is that a major new discovery surface can change visibility and traffic while remaining difficult to analyze as a distinct channel.

Bing is taking a more explicit route. Microsoft describes AI Performance as an early step toward tooling for Generative Engine Optimization and says it wants publishers to understand how their content participates in AI-driven experiences. That framing treats citations and grounding as first-class webmaster signals rather than an invisible layer inside a broader search metric.

This does not automatically make Bing more transparent in every respect, nor does it make the AI Performance report complete. Microsoft says the dashboard represents aggregated citation activity and may not contain every instance in which content was referenced. It also emphasizes that citations are not clicks: a page can become highly visible to an AI system without receiving equivalent human traffic.

But that limitation is itself useful. It forces publishers to separate two concepts that traditional SEO often allowed them to merge: visibility and visitation.

AI visibility is becoming its own performance layer

In classic search, a high-ranking result was valuable largely because it created an opportunity for a click. In AI search, a publisher can create value upstream by grounding an answer even if the user never leaves the interface. That means the industry needs metrics capable of describing participation before traffic occurs.

Bing’s citation data offers one such layer. Search analytics can tell a publisher whether people reached its site. AI Performance can help show whether the site’s information was used visibly in generated answers. The two datasets describe different outcomes and should not be confused.

That distinction also has commercial consequences. Publishers debating the value they provide to AI platforms need evidence about how frequently their work appears in those experiences. Brands trying to understand AI reputation need to know whether their own pages are being cited around important topics. SEO teams need to know whether conventional ranking success translates into generative visibility. Without platform data, all three groups are forced to approximate the answer externally.

What publishers can actually do with the data

The most productive use of AI Performance is unlikely to be chasing every citation fluctuation. Bing warns that changes can result from multiple external factors, and the dashboard should be treated as a trend and comparative tool rather than a precise ledger of every generated answer.

Instead, publishers can use it to build a baseline. Which subject areas produce citations? Which URLs repeatedly ground answers? Do particular content formats or editorial sections appear more often? Does a site have high citation volume but a low share around strategically important queries? Are citations concentrated in a small number of pages, creating dependence on a narrow slice of the content library?

Those questions lead to better editorial experiments. Teams can improve clarity, deepen supporting evidence, refresh outdated information, strengthen coverage around relevant topics and then observe whether the broader pattern changes. The data cannot prove that an edit caused an increase, but it can make the process more disciplined than repeatedly prompting an AI assistant and recording screenshots.

The next search console needs to measure machines as well as humans

Bing Webmaster Tools is not suddenly replacing Google Search Console as the center of SEO operations. Google’s scale ensures that its data remains essential for publishers, and Microsoft’s AI report is still evolving. Yet Bing may be showing what webmaster tooling needs to become in an AI-mediated web.

The first generation of search analytics measured how humans found websites. The next generation also needs to show how machines find, retrieve, cite and synthesize those websites before a human ever decides whether to click.

That is why Bing’s quiet progress matters. AI visibility has spent too long as something publishers are told to optimize without being able to observe directly. By exposing citations, grounding queries, topics, intent and relative citation presence, Microsoft is beginning to turn an opaque new discovery system into something webmasters can interrogate with data. The metrics are imperfect, but imperfect measurement is a meaningful improvement over having to guess.

0%