Summarization is Not Dead Yet
cs.CL, cs.AI
Submitted: 2026-06-06
Updated: 2026-08-31
Comments: EMNLP 2026 Main & Long Conference Paper
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an
Terminology
Abstract
The progress of large language models (LLMs) has fueled claims that model-generated summaries rival or even surpass human-written references, raising questions about whether summarization remains an open research problem. We re-examine this narrative through a multi-track evaluation covering diverse datasets and state-of-the-art LLMs, combining controlled human assessment, bias-mitigated LLM-as-Judge protocols, factuality verification against external knowledge, and corpus-level linguistic analysis. Our findings reveal a more nuanced landscape in which human references continue to demonstrate advantages in informativeness and faithfulness, whereas LLM outputs are preferred mainly for surface-level coherence and fluency. Factuality verification indicates that human references remain more reliable, particularly for claims involving reasoning or synthesis, and linguistic analysis uncovers a pattern of stylistic homogeneity across different models. These observations suggest that current LLMs have raised the floor of summarization quality, but the ceiling of their performance remains below human capabilities.
Sources
- From Local to Global: A Graph RAG Approach to Query-Focused Summarization
- News Summarization and Evaluation in the Era of GPT-3
- Summarization is (Almost) Dead
- ArchRAG: Attributed Community-based Hierarchical Retrieval-Augmented Generation
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
- Fairness or Fluency? An Investigation into Language Bias of Pairwise LLM-as-a-Judge
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering