Pigeonholing: how bad prompts hurt models, causing collapse and mistakes
cs.CL, cs.AI
Submitted: 2026-06-23
Updated: 2026-08-26
Comments: EMNLP 2026 (Findings)
Code: https://github.com/TIGER-AI-Lab/MMLU-Pro
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
The gist: While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing."
Terminology
Abstract
While in-context learning is generally shown to be effective in Large Language Models (LLMs), bad contexts can cause performance degradation and mode collapse, a phenomenon we call "pigeonholing." **Unintentionally bad** contexts can happen without malicious jailbreaking intents: For example, a user asks the model to justify an incorrect math theorem or fails to correct the model's buggy code. Specifically, we investigate ``pigeonholing" in two scenarios: (1) when the user suggests a solution, and (2) when the conversation context includes the assistant's previous (incorrect) responses. Our experiments across 10 verifiable and open-ended tasks with 10 different models show that pigeonholing manifests in several ways: (1) repeating the incorrect answers from context (leading to 38-40% performance drop), (2) converging on a narrow set of answers in coding and text generation without exploring alternatives, and (3) flipping stance on controversial topics to align with the user or the assistant's previous claims. We find that pigeonholing worsens almost monotonically with the number of conversation turns (performance drops by additional 14+% as repeated mistakes increase from 1 to 5), and pigeonholing-induced mode collapse can happen even when the provided example is correct. As a step toward mitigation, we propose RLVR with synthetic errors which improves models by 43-60% under bad contexts compared to vanilla RLVR baselines.
Sources
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- How LLMs Distort Our Written Language
- Language Models are Few-Shot Learners
- The Capacity for Moral Self-Correction in Large Language Models
- Sycophantic Chatbots Cause Delusional Spiraling, Even in Ideal Bayesians
- ChatInject: Abusing Chat Templates for Prompt Injection in LLM Agents
- The Llama 3 Herd of Models
- From Yes-Men to Truth-Tellers: Addressing Sycophancy in Large Language Models with Pinpoint Tuning
- ToMBench: Benchmarking Theory of Mind in Large Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- OpenRLHF: An Easy-to-use, Scalable and High-performance RLHF Framework
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Large Language Models Cannot Self-Correct Reasoning Yet
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- Way Off-Policy Batch Deep Reinforcement Learning of Implicit Human Preferences in Dialog
- Artificial Hivemind: The Open-Ended Homogeneity of Language Models (and Beyond)
- A Simple "Try Again" Can Elicit Multi-Turn LLM Reasoning
- The PRISM Alignment Dataset: What Participatory, Representative and Individualised Human Feedback Reveals About the Subjective and Multicultural Alignment of Large Language Models
- Improving Interactive In-Context Learning from Natural Language Feedback
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering