Are Finer Citations Always Better? Rethinking Granularity for Attributed Generation
cs.CL
Submitted: 2026-04-01
Updated: 2026-09-20
Code: https://github.com/JHU-CLSP/citation-granularity
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Citation-Grounded Code Comprehension: Preventing LLM Hallucination Through Hybrid Retrieval and Graph-Augmented Context
- Hidden in the Haystack: Smaller Needles are More Difficult for LLMs to Find
- Attributed Question Answering: Evaluation and Modeling for Attributed Large Language Models
- Long-context LLMs Struggle with Long In-context Learning
- A Comprehensive Survey on Long Context Language Modeling
- Concise and Sufficient Sub-Sentence Citations for Retrieval-Augmented Generation
- Lost in the Middle: How Language Models Use Long Contexts
- A Survey of Scaling in Large Language Model Reasoning
- The Llama 3 Herd of Models
- Teaching language models to support answers with verified quotes
- WebGPT: Browser-assisted question-answering with human feedback
- gpt-oss-120b & gpt-oss-20b Model Card
- Attribution, Citation, and Quotation: A Survey of Evidence-based Text Generation with Large Language Models
- LongRAG: Enhancing Retrieval-Augmented Generation with Long-context LLMs
- More Documents, Same Length: Isolating the Challenge of Multiple Documents in RAG
- LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA
- Finding the Minimal Parameter Budget for Implicit Reasoning: A Data Complexity Driven Scaling Law for Language Models
- Retrieval Head Mechanistically Explains Long-Context Factuality
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering