Spontaneous Context Restoration: How Language Models Recover from Corrupted Inputs
cs.CL, cs.AI
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- Cutting Through the Noise: Boosting LLM Performance on Math Word Problems
- Eliciting Latent Predictions from Transformers with the Tuned Lens
- Exploring LLM Reasoning Through Controlled Prompt Variations
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Not All Language Model Features Are One-Dimensionally Linear
- Localizing Model Behavior with Path Patching
- How to use and interpret activation patching
- The Remarkable Robustness of LLMs: Stages of Inference?
- Improving the Robustness of Large Language Models for Code Tasks via Fine-tuning with Perturbed Data
- The Geometry of Truth: Emergent Linear Structure in Large Language Model Representations of True/False Datasets
- Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference
- The Hydra Effect: Emergent Self-repair in Language Model Computations
- How Context Affects Language Models' Factual Predictions
- SQuAD: 100,000+ Questions for Machine Comprehension of Text
- Explorations of Self-Repair in Language Models
- Robustness of Large Language Models to Perturbations in Text
- A Mechanistic Interpretation of Arithmetic Reasoning in Language Models using Causal Mediation Analysis
- Interpretability in the Wild: a Circuit for Indirect Object Identification in GPT-2 small
- Interpreting Arithmetic Mechanism in Large Language Models through Comparative Neuron Analysis
- Towards Best Practices of Activation Patching in Language Models: Metrics and Methods
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering