Learning Evidence Sufficiency Boundaries for Selective Answering in Grounded Multi-Hop QA
cs.CL
Submitted: 2026-09-01
Updated: 2026-09-01
Code: https://github.com/huggingface/peft
License: http://creativecommons.org/licenses/by/4.0/
The gist: Grounded question answering systems should answer only when the supplied evidence supports the answer.
Terminology
Abstract
Grounded question answering systems should answer only when the supplied evidence supports the answer. In multi-hop QA, this requirement is difficult because partial evidence can make an unsupported answer appear plausible. We study selective answering through evidence sufficiency boundaries: for the same question, a model should abstain under unsupported or partially supported context, answer when the context first becomes sufficient, and keep the answer stable when redundant evidence is added. We introduce Evidence Sufficiency Boundary Training, a generation-native training framework that constructs ordered evidence chains and supervises the abstain-to-answer transition directly. The method combines level supervision, a boundary flip margin, post-boundary stability, and answer recall protection. We build evidence chains from HotpotQA, 2WikiMultiHopQA, and MuSiQue, then evaluate models with chain metrics, raw QA utility, and unsupported-answer rates on external non-answerable sets. With Qwen2.5-3B-Instruct and LoRA adaptation, Evidence Sufficiency Boundary Training gives the strongest boundary localization among the tested systems, with flip accuracy of 0.807 compared with 0.781 for a token-level abstention baseline. It also achieves the lowest overall unsupported-answer rate on external non-answerable evaluation, 0.095 compared with 0.101 for the same baseline, while retaining competitive raw QA F1. The results show that grounded selective answering improves when training marks the evidence level where refusal should give way to answering.
Sources
- Sufficient Context: A New Lens on Retrieval Augmented Generation Systems
- Language Models (Mostly) Know What They Know
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- FaithEval: Can Your Language Model Stay Faithful to Context, Even If "The Moon is Made of Marshmallows"
- RAGTruth: A Hallucination Corpus for Developing Trustworthy Retrieval-Augmented Language Models
- GPT-4 Technical Report
- Qwen2.5 Technical Report
- Do Large Language Models Know What They Don't Know?
- R-Tuning: Instructing Large Language Models to Say `I Don't Know'
- Learning to Look Again: Loss-Gap Supervision for Free-form Crop Routing in Vision-Language Models
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Learning Less Is More: Premature Upper-Layer Attention Specialization Hurts Language Model Pretraining
- HeLa-Mem: Hebbian Learning and Associative Memory for LLM Agents
- Where Does Long-Context Supervision Actually Go? Effective-Context Exposure Balancing
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering