FOCUS & RePAIR: Mitigating Text Degeneration via Token-Level Guidance for Pruned Large Language Models
cs.CL, cs.AI, cs.LG
Submitted: 2026-08-27
Updated: 2026-08-27
Comments: Accepted to ICML 2026 as a Spotlight
Code: https://github.com/tatsu-lab/stanford_alpaca
License: http://creativecommons.org/licenses/by/4.0/
The gist: Pruning is a practical approach to compress large language models (LLMs), but it can amplify text degeneration, especially repetition loops, even when perplexity and task accuracy remain largely
Terminology
Abstract
Pruning is a practical approach to compress large language models (LLMs), but it can amplify text degeneration, especially repetition loops, even when perplexity and task accuracy remain largely unchanged. In this work, we present a token-level analysis of this failure mode by viewing decoding as a dynamical process that enters and persists in a small set of recurrent contexts. Our analysis decomposes degeneration into loop entry risk and loop persistence, and shows that persistence is controlled by the escape mass assigned to plausible alternatives within the token sampling set. Motivated by these findings, we propose two token-level guidance objectives for post-pruning fine-tuning. FOCUS reweights distillation toward high-confidence teacher regions to suppress leakage, while RePAIR uses onset-centered positive/negative continuation pairs with a margin loss to promote plausible alternatives and prevent early commitment to repetition loops. Experiments on open-ended continuation and instruction-based generation show that both methods consistently reduce repetition and improve generation quality.
Sources
- A Survey on Large Language Models for Code Generation
- Shortened LLaMA: Depth Pruning for Large Language Models with Comparison of Retraining Methods
- DistiLLM: Towards Streamlined Distillation for Large Language Models
- The Stable Entropy Hypothesis and Entropy-Aware Decoding: An Analysis and Algorithm for Robust Natural Language Generation
- Towards Robust Pruning: An Adaptive Knowledge-Retention Pruning Strategy for Language Models
- Lost in Pruning: The Effects of Pruning Neural Networks beyond Test Accuracy
- Straight to the Gradient: Learning to Use Novel Tokens for Neural Text Generation
- SparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot
- Rethinking and Refining the Distinct Metric
- Undistillable: Making A Nasty Teacher That CANNOT teach students
- The Curious Case of Neural Text Degeneration
- LLM-Pruner: On the Structural Pruning of Large Language Models
- Large Language Models: A Survey
- A Survey on Knowledge Distillation of Large Language Models
- Revisiting Weak-to-Strong Generalization in Theory and Practice: Reverse KL vs. Forward KL
- Can Students Beyond The Teacher? Distilling Knowledge from Teacher's Bias
- Efficiently Scaling Transformer Inference
- BERTScore: Evaluating Text Generation with BERT
- Direct Preference Optimization: Your Language Model is Secretly a Reward Model
- Membership and Memorization in LLM Knowledge Distillation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering