Google's John Mueller has confirmed what many practitioners already suspected: the way Search Console reports on AI Mode and AI Overviews will change over time, because the surfaces themselves are still changing. That is not a criticism of Google. It is an honest statement about the stage of development these products are at. But it has real implications for anyone who is trying to build a credible, repeatable measurement framework around AI search visibility.
The practical problem is not that the data will change. Data always changes as products mature. The problem is that teams are being asked to justify investment in AEO and GEO programmes right now, using metrics that Google itself has flagged as provisional. That tension needs to be managed carefully.
What Mueller Actually Said
Mueller's comments were direct: position and impression tracking for AI Mode, AI Overviews and standard Google Search will evolve as those products evolve. This is not a minor caveat. AI Mode and AI Overviews do not behave like traditional search results. They do not have ten blue links with stable positions. A citation inside an AI Overview is not the same as a rank-one organic result, and the way impressions are counted does not yet have a settled, agreed definition.
Search Console has started surfacing some AI Overview data, which is genuinely useful. But Mueller's statement is a reminder that the methodology behind that data is still being worked out. If you are pulling GSC reports on AI Overviews and treating position numbers with the same confidence you would treat a page-one ranking from three years ago, you are working with a false sense of precision.
Why This Creates a Measurement Problem Right Now
Most AEO programmes are at an early stage of proving their value internally. Teams need to show that appearing in AI-generated answers actually drives traffic, leads and - ultimately - revenue. That case is hard to make when the primary reporting tool for Google's AI surfaces has openly provisional data.
There is also a fragmentation issue. AI search is not a single surface. Google AI Overviews and AI Mode sit inside Search Console's reporting scope. ChatGPT, Perplexity and Gemini do not. Each of those platforms handles citations and referral traffic differently, and none of them offer anything close to the reporting depth that Google Ads or GA4 provide. Waiting for GSC to mature before building a measurement approach is not a viable position when queries are already being answered by these systems at scale.
Building a Measurement Framework That Does Not Depend on One Signal
The right response to provisional platform data is to triangulate. GSC AI Overview data is one signal. GA4 referral source analysis - identifying sessions attributed to ChatGPT, Perplexity, Gemini and similar - is another. A third is CRM-level attribution: tagging inbound leads and enquiries by source so that AI-referred contacts can be tracked through to pipeline and revenue. No single layer gives you the full picture. All three together start to build a defensible view.
Within GA4, AI referral traffic requires explicit configuration. Perplexity, ChatGPT and Gemini do not always pass referral data cleanly. Some sessions arrive as direct traffic. Channel groupings need to be updated to capture these sources correctly, and that configuration work needs to happen before you can trust any trend data. If that work has not been done, your AI traffic figures are almost certainly understated.
On the CRM side, the question to ask is simple: do you know which of your current customers first found you through an AI-generated answer? If the answer is no, you have a gap in your attribution that no amount of Search Console data will fill. UTM parameters on AI-referred links, combined with a clear lead source field in your CRM, are the minimum requirement for making that connection.
What Provisional GSC Data Is Still Useful For
None of this means you should ignore what Search Console does report. Even with the caveats Mueller has flagged, GSC AI Overview data gives you a directional view of which queries are surfacing your content in AI-generated answers, and which pages are earning those appearances. That directional signal is valuable for content strategy, even if the position numbers are not yet precise enough for competitive benchmarking.
Use GSC AI data to identify patterns - which topic clusters, which content formats, which pages are appearing. Use that to prioritise where you optimise content structure, add structured data, or build supporting citation coverage through digital PR. The signal is real. The precision is not yet there. Work with the former, acknowledge the latter.
The Broader Point About AEO Measurement Maturity
Mueller's confirmation that Search Console reporting will evolve is a useful reminder that AEO and GEO measurement is still in an early phase across the whole industry. Google is working out how to report on AI surfaces. Practitioners are working out how to attribute value from them. Both are happening simultaneously, which is uncomfortable but not unusual for a channel in its first few years.
The mistake would be to treat that uncertainty as a reason to deprioritise measurement investment. The teams that build solid attribution infrastructure now - before GSC reporting stabilises, before GA4 channel groupings are standardised for AI sources, before CRM lead sources are properly tagged - will be in a significantly stronger position when the reporting does mature. They will have months of clean data. Everyone else will be starting from scratch.
Reporting tools catch up with reality eventually. The question is whether your measurement foundations will be ready when they do.