Rating the Pitch, Not the Product: User Evaluations of LLMs Reflect Expectations More Than Performance
cs.CL, cs.CY, cs.HC
Submitted: 2026-07-06
Updated: 2026-08-31
Comments: Accepted to EMNLP 2026 Main Conference
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- From Prompting to Partnering: Personalization Features for Human-LLM Interactions
- Making Sense of AI Limitations: How Individual Perceptions Shape Organizational Readiness for AI Adoption
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering