Do as I Say, Not as I Do: Instruction-Induction Conflict in LLMs
cs.CL, cs.AI
Submitted: 2026-05-19
Updated: 2026-09-01
Terminology
Sources
- Emergent Misalignment via In-Context Learning: Narrow in-context examples can produce broadly misaligned LLMs
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- A Survey on In-context Learning
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
- Rethinking the Role of Demonstrations: What Makes In-Context Learning Work?
- Can LLMs Follow Simple Rules?
- Olmo 3
- Generalizing Verifiable Instruction Following
- InFoBench: Evaluating Instruction Following Ability in Large Language Models
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Next-token pretraining implies in-context learning
- The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions
- Adversarial Demonstration Attacks on Large Language Models
- Jailbroken: How Does LLM Safety Training Fail?
- Larger language models do in-context learning differently
- Benchmarking Complex Instruction-Following with Multiple Constraints Composition
- Instruction-Following Evaluation for Large Language Models
- Hijacking Large Language Models via Adversarial In-Context Learning
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering