Lost-in-the-Middle in Long-Text Generation: Synthetic Dataset, Evaluation Framework, and Mitigation
cs.CL, cs.AI
Submitted: 2025-03-10
Updated: 2026-09-13
Code: https://github.com/OnlyAR/RAL-Writer
License: http://creativecommons.org/licenses/by/4.0/
The gist: Existing long-text generation methods produce lengthy outputs from short inputs, leaving long-input-to-long-output generation underexplored.
Terminology
Abstract
Existing long-text generation methods produce lengthy outputs from short inputs, leaving long-input-to-long-output generation underexplored. As input length increases, LLMs increasingly overlook information in the middle of the context--a limitation known as the "lost-in-the-middle" phenomenon--leading to inconsistent and incoherent outputs. To address this problem, we propose Retrieval-Augmented Long-Text Writer (RAL-Writer), a training-free framework consisting of a Planner that generates writing steps and a Writer that produces content according to these steps. The Writer jointly models semantic relevance and positional bias to compute importance scores, dynamically retrieve critical input segments, and strategically restate them. We also construct a benchmark dataset for long-input-to-long-output generation and introduce three evaluation metrics covering length, consistency, and quality. We evaluate RAL-Writer against comparable baselines on this dataset. The results demonstrate the effectiveness of our approach. Code is available at https://github.com/OnlyAR/RAL-Writer.
Sources
- Make Your LLM Fully Utilize the Context
- LongWriter: Unleashing 10,000+ Word Generation from Long Context LLMs
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language Models
- YaRN: Efficient Context Window Extension of Large Language Models
- LongDPO: Unlock Better Long-form Generation Abilities for LLMs via Critique-augmented Stepwise Information
- Language Models can Self-Lengthen to Generate Long Texts
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- HelloBench: Evaluating Long Text Generation Capabilities of Large Language Models
- PoSE: Efficient Context Window Extension of LLMs via Positional Skip-wise Training
- LongGenBench: Benchmarking Long-Form Generation in Long Context LLMs
- C-Pack: Packed Resources For General Chinese Embeddings
- Qwen2.5 Technical Report
- LongQLoRA: Efficient and Effective Method to Extend Context Length of Large Language Models
- Found in the Middle: How Language Models Use Long Contexts Better via Plug-and-Play Positional Encoding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering