Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing
cs.CL
Submitted: 2026-05-11
Updated: 2026-08-29
License: http://creativecommons.org/licenses/by/4.0/
The gist: Long-context adaptation is often viewed as window scaling, but this misses a token-level supervision mismatch: in packed training with document masking, each target token's effective context remains
Terminology
Abstract
Long-context adaptation is often viewed as window scaling, but this misses a token-level supervision mismatch: in packed training with document masking, each target token's effective context remains short. We introduce EXACT, a supervision-allocation objective that assigns extra weight to long effective-context targets by inverse frequency within the long tail. Across seven Qwen/LLaMA CPT configurations, EXACT improves all 28 trained/extrapolated NoLiMa and RULER comparisons. On Qwen2.5-0.5B, NoLiMa improves by +10.09 (trained) and +5.34 (extrapolated); RULER by +10.69 and +5.55. On LLaMA-3.2-3B, RULER improves by +17.91 and +16.11. Standard QA/reasoning are preserved (+0.24 macro change across six benchmarks). A distance-resolved probe shows gains arise when evidence is thousands of tokens away, while short cases remain unchanged. Results support a supervision-centric thesis: long-context adaptation depends on how strongly training supervises long-context predictions.
Sources
- LongAlign: A Recipe for Long Context Alignment of Large Language Models
- Longformer: The Long-Document Transformer
- Extending Context Window of Large Language Models via Positional Interpolation
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
- Generating Long Sequences with Sparse Transformers
- Rethinking Attention with Performers
- FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
- LongNet: Scaling Transformers to 1,000,000,000 Tokens
- LongRoPE: Extending LLM Context Window Beyond 2 Million Tokens
- What is Wrong with Perplexity for Long-context Language Modeling?
- Data Engineering for Scaling Language Models to 128K Context
- The Llama 3 Herd of Models
- Token Weighting for Long-Range Language Modeling
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- LLM Maybe LongLM: Self-Extend LLM Context Window Without Tuning
- One Thousand and One Pairs: A "novel" challenge for long-context language models
- Ring Attention with Blockwise Transformers for Near-Infinite Context
- NoLiMa: Long-Context Evaluation Beyond Literal Matching
- Leave No Context Behind: Efficient Infinite Context Transformers with Infini-attention
- OpenWebMath: An Open Dataset of High-Quality Mathematical Web Text
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering