User Feedback Provides a Unique Signal that LLMs Can not Detect
cs.CL
Submitted: 2026-09-02
Updated: 2026-09-02
License: http://creativecommons.org/licenses/by/4.0/
The gist: Harnessing naturally occurring feedback from user interactions offers a promising learning signal for Large Language Models (LLMs).
Terminology
Abstract
Harnessing naturally occurring feedback from user interactions offers a promising learning signal for Large Language Models (LLMs). However, recent studies suggest this feedback is inherently noisy and difficult to leverage effectively. We challenge this conception by demonstrating that user feedback is a highly actionable signal for improvement, and that its perceived ineffectiveness stems from a systematic bias in current evaluation paradigms. To isolate the usefulness of feedback, we construct synthetic data with a definitive ground truth, alongside naturalistic data to validate that our findings hold in real-world scenarios. By comparing model revisions generated with and without access to feedback across both settings, we show that feedback-informed revisions resolve targeted issues at significantly higher rates than baseline revisions. Finally, we expose the root of the evaluation bias: when a model successfully fixes an issue exclusively due to feedback, LLM judges frequently fail to identify the genuinely corrected response, systematically preferring inferior baseline outputs instead.
Sources
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference
- Naturally Occurring Feedback is Common, Extractable and Useful
- gpt-oss-120b & gpt-oss-20b Model Card
- Aligning Language Models from User Interactions
- The Era of Real-World Human Interaction: RL from User Conversations
- From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
- Gemini: A Family of Highly Capable Multimodal Models
- Qwen3 Technical Report
- Verbosity Bias in Preference Labeling by Large Language Models
- JudgeLM: Fine-tuned Large Language Models are Scalable Judges
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering