Audit: Baidu, Google AI overviews barely overlap on sources
A cross-lingual audit compared AI-generated overviews on Baidu and Google using English queries from the MS MARCO benchmark and their Chinese translations, in a paper accepted at WAC@EMNLP 2026. Overview availability and which host domains appeared varied substantially by platform and language, with little overlap between the sets of domains visible in each setting’s overviews — even as answers matched to the same query intent scored a median cosine similarity of 0.701 to 0.813 across settings. The authors say source-exposure concentration and answer-level similarity capture distinct dimensions of AI-search behavior, so judging an AI overview on its answer text alone misses how differently platforms select and expose sources.
Why it matters: It shows that answer content and source exposure are separate things to track when optimizing across engines and languages — matching a competitor's answer text says nothing about matching its cited sources.
The record: Google AI OverviewsGoogle
Glossary: AI Overviews
Via arXiv AI-search citations ↗
Posted to the wire September 28, 2026. Edited by Joe Balewski.