The Copy Ceiling: An Input-Exposure Control for Ontology-Grounded Generation over Curated Corpora
cs.CL, cs.CY
Submitted: 2026-09-21
Updated: 2026-09-29
Comments: 28 pages, 6 figures
Code: https://github.com/DreamLab-AI/loom
License: http://creativecommons.org/licenses/by/4.0/
The gist: When a language model answers from a curated corpus via graph-based retrieval, a large grounding uplift does not establish reasoning over the retrieved structure: the context may already expose the
Terminology
Abstract
When a language model answers from a curated corpus via graph-based retrieval, a large grounding uplift does not establish reasoning over the retrieved structure: the context may already expose the gold answers. We propose exposure accounting, which classifies each gold item by whether the shown context exposes it and whether the answer recovers it. Its scalar reference is the copy ceiling, the recall a verbatim copy of the context achieves; signed gain over copy measures the model's recall relative to this deterministic, judge-free baseline. Across ten models, unaided recall averages 0.26 and grounded recall 0.92, yet gain over copy is uniformly negative (-0.067 to-0.022). Of 11,360 gold-item observations, representing 1,136 target instances evaluated under ten models, only three unexposed items receive lexical credit. A stratified model-judged audit of 423 observations, with a symmetric quotation-verification policy, estimates that 97.1% of credited items assert the requested relation; all three unexposed credits fail relational adjudication. On targets the scaffold does not expose, lexical recovery falls from 0.121 unaided to 0.004 grounded; adjudication validates 71 of the 92 unaided credits and none of the three grounded credits, without establishing full-frame relational recovery rates. Rephrasing questions outside the graph's title vocabulary reduces exposure from 0.964 to 0.328, while an absence-triggered fallback activates on only 2 of 506 questions. A paired production study improves judged quality by +0.27 pooled, but negative controls do not establish content specificity beyond a well-formed on-corpus block. These results support exposure accounting as a standing control for corpus-derived evaluations. The accounting distinguishes exposed-item omissions from beyond-exposure recoveries; it does not determine whether reasoning occurred.
Sources
- Query Performance Prediction for Neural IR: Are We There Yet?
- Evaluating Correctness and Faithfulness of Instruction-Following Models for Question Answering
- Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection
- Enabling Large Language Models to Generate Text with Citations
- Retrieval-Augmented Generation for Large Language Models: A Survey
- Neurosymbolic AI: The 3rd Wave
- The Power of Noise: Redefining Retrieval for RAG Systems
- LightRAG: Simple and Fast Retrieval-Augmented Generation
- Annotation Artifacts in Natural Language Inference Data
- G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering
- Adaptive-RAG: Learning to Adapt Retrieval-Augmented Large Language Models through Question Complexity
- Ragas: Automated Evaluation of Retrieval Augmented Generation
- Sufficient Context: A New Lens on Retrieval Augmented Generation Systems
- How Much Reading Does Reading Comprehension Require? A Critical Investigation of Popular Benchmarks
- On Robustness and Reliability of Benchmark-Based Evaluation of LLMs
- When Not to Trust Language Models: Investigating Effectiveness of Parametric and Non-Parametric Memories
- Stronger Baselines for Retrieval-Augmented Generation with Long-Context Language Models
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Answer Presence Drives RAG Rewriting Gains
- Towards Practical GraphRAG: Efficient Knowledge Graph Construction and Hybrid Retrieval at Scale
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering