Answer Bubbles: Information Exposure in AI-Mediated Search
cs.IR, cs.CL
Submitted: 2026-03-17
Updated: 2026-08-28
Comments: EMNLP 2026: 16 pages, 3 figures, 9 tables
Code: https://github.com/sloria/TextBlob
License: http://creativecommons.org/licenses/by/4.0/
The gist: Generative search systems are increasingly replacing link-based retrieval with AI-generated summaries, yet little is known about how these systems differ in sources, language, and fidelity to cited
Terminology
Abstract
Generative search systems are increasingly replacing link-based retrieval with AI-generated summaries, yet little is known about how these systems differ in sources, language, and fidelity to cited material. We examine responses to 11,000 real search queries across five systems---vanilla GPT, Search GPT, Perplexity Search with Grok, Google AI Overviews, and traditional Google Search---at three levels: source diversity, linguistic characterization of the generated summary, and source-summary fidelity. We find that generative search systems exhibit significant source-selection biases in their citations, favoring certain sources over others. Incorporating search also selectively attenuates epistemic markers, reducing hedging by up to 60% while preserving confidence language in the AI-generated summaries. At the same time, AI summaries further compound the citation biases: Wikipedia and longer sources are disproportionately overrepresented, whereas cited social media content and negatively framed sources are substantially underrepresented. Our findings highlight the potential for answer bubbles, in which identical queries yield structurally different information realities across systems, with implications for user trust, source visibility, and the transparency of AI-mediated information access.
Sources
- Large Language Model Agent for Fake News Detection
- A Survey of Generative Search and Recommendation in the Era of Large Language Models
- Retrieval-Augmented Generation for Large Language Models: A Survey
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- BERTopic: Neural topic modeling with a class-based TF-IDF procedure
- Do We Know What They Know We Know? Calibrating Student Trust in AI and Human Responses Through Mutual Theory of Mind
- Transparency, Privacy, and Fairness in Recommender Systems
- ZeroSearch: Incentivize the Search Capability of LLMs without Searching
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG