Learning When Not to Listen: Selective Anti-Interference Pretraining for Language Models
cs.CL
Submitted: 2026-08-22
Updated: 2026-08-22
Terminology
Sources
- LongBench: A Bilingual, Multitask Benchmark for Long Context Understanding
- Longformer: The Long-Document Transformer
- Extending Context Window of Large Language Models via Positional Interpolation
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
- What is Wrong with Perplexity for Long-context Language Modeling?
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Reformer: The Efficient Transformer
- LooGLE: Can Long-Context Language Models Understand Long Contexts?
- NeedleBench: Evaluating LLM Retrieval and Reasoning Across Varying Information Densities
- Long-context LLMs Struggle with Long In-context Learning
- Lost in the Middle: How Language Models Use Long Contexts
- NoLiMa: Long-Context Evaluation Beyond Literal Matching
- YaRN: Efficient Context Window Extension of Large Language Models
- Train Short, Test Long: Attention with Linear Biases Enables Input Length Extrapolation
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Long Context Alignment with Short Instructions and Synthesized Positions
- HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering