WildSEEK: Evaluating Language Models for Information-Seeking
cs.CL, cs.CY
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: 9 pages, accepted at EMNLP Main 2026
Project page: https://www.usa.gov
License: http://creativecommons.org/licenses/by/4.0/
The gist: Language models are increasingly mediating information access to end users, urging a systematic evaluation of their responses for a fair and reliable information ecosystem.
Terminology
Abstract
Language models are increasingly mediating information access to end users, urging a systematic evaluation of their responses for a fair and reliable information ecosystem. Existing evaluations, however, are often topic-specific or synthetic, limiting their ability to capture the complexity of "in the wild" information-seeking queries and the risks present in model responses. To address this gap, we introduce WildSEEK, a manually annotated dataset of 3k information-seeking queries from real user interactions, and an evaluation framework for LLM-generated responses. WildSEEK includes annotations for risk-sensitive domains (e.g. health and financial information), and distinguishes factoid queries from analytical queries which seek responses beyond facts. We train classifiers on WildSEEK to analyze more than 1.8M realistic user queries. We find that over a third of information-seeking queries are high-risk and more often analytical. Our findings show that LLM responses fail more often in four criteria: sycophantic behavior, overreliance, a default US-centric perspective, and poor handling of vulnerable populations -- with failure rates being mostly higher for analytical queries. By providing methods to monitor the reliability, safety, and fairness of LLM behavior, our dataset and evaluation framework offer an empirical foundation for the broader question of how these systems should behave as they take on a growing role in information access.
Sources
- Large Language Model Hacking: Quantifying the Hidden Risks of Using LLMs for Text Annotation
- What Is The Political Content in LLMs' Pre- and Post-Training Data?
- A Survey on Large Language Models for Critical Societal Domains: Finance, Healthcare, and Law
- It's About Time: The Temporal and Modal Dynamics of Copilot Usage
- Structuring the Space of Perspectives
- BERTopic: Neural topic modeling with a class-based TF-IDF procedure
- Auditing Google's AI Overviews and Featured Snippets: A Case Study on Baby Care and Pregnancy
- TRIDENT: Benchmarking LLM Safety in Finance, Medicine, and Law
- Conversational AI increases political knowledge as effectively as self-directed internet search
- Safe for Whom? Rethinking How We Evaluate the Safety of LLMs for Real Users
- WildChat: 1M ChatGPT Interaction Logs in the Wild
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering