How You Ask Shapes What You Get: A Theory-Seeded Measurement of Articulation in Advice-Seeking LLM Conversations
cs.CL
Submitted: 2026-08-30
Updated: 2026-08-30
Code: https://github.com/Juneha-Baek/articulation-segmentation-emnlp2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Users articulate the same advice-seeking request in different ways: some specify detailed constraints, others gesture at a vague need.
Terminology
Abstract
Users articulate the same advice-seeking request in different ways: some specify detailed constraints, others gesture at a vague need. Prior work treats this variation as noise to be averaged away; we instead treat it as a stable, measurable structure in the input distribution. We ask whether articulation (how people ask) forms latent dimensions separable from topic (what they ask about), and whether it is associated with how language models respond. We extract interpretable features from 16,447 advice-seeking prompts pooled from public chat corpora (WildChat, LMSYS, and ShareChat) and recover a small set of latent articulation factors that replicate across train/test splits and across corpora. Because this structure is largely separable from topic, the populations it defines cut across topics and stay invisible to topic- or task-based evaluation. The factors define a handful of recurring articulation styles, one of which stands out: a long-form but information-poor style, roughly one in six prompts in the largest corpus, where models return shorter, vaguer answers and do not ask for clarification even though under-specification is exactly the condition that warrants it. The contrast holds within every topic group and length quintile, and is not under-specification alone -- a second, equally under-specified style does draw clarifying questions. Two independent human annotators reproduce this contrast. We argue that benchmarks should stratify on articulation, and we offer the extracted structure as a measurement instrument for doing so.
Sources
- Eliciting Human Preferences with Language Models
- Native Design Bias: Studying the Impact of English Nativeness on Language Model Performance
- Conversational User-AI Intervention: A Study on Prompt Rewriting for Improved LLM Response Generation
- Clio: Privacy-Preserving Insights into Real-World AI Use
- Text Embeddings by Weakly-Supervised Contrastive Pre-training
- ShareChat: A Dataset of Chatbot Conversations in the Wild
- Disparities in LLM Reasoning Accuracy and Explanations: A Case Study on African American English
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering