The Missing Complement: State-Conditioned Minimal Sufficient Evidence for Coding Agents
cs.IR, cs.AI, cs.CL
Submitted: 2026-09-17
Updated: 2026-09-26
Comments: 32 pages, 3 figures. Benchmark and evaluation resources: https://github.com/LordTARN1SHED/SERBench
Code: https://github.com/LordTARN1SHED/SERBench
License: http://creativecommons.org/licenses/by/4.0/
The gist: A coding agent halfway through an issue has already read much of what a retriever ranks highest.
Terminology
Abstract
A coding agent halfway through an issue has already read much of what a retriever ranks highest. Relevance is scored per passage, but sufficiency belongs to the set: a ranker can fill its budget with variants of one required fact and leave the decision unsupported. We formulate state-conditioned minimal sufficient evidence recovery: given a captured agent state, recover a compact evidence combination that supplies the support its next decision still lacks. SERBench measures this on 500 held-out states from 45 repositories, recording what the agent has seen and crediting only sets that cover every fact the current decision was annotated to require. MSS-Complement treats acquisition as set construction, not ranking. Three semantic calls propose a jointly sufficient set, search for what it lacks, and return 4-8 intact source units within 6,144 tokens. One configuration, fixed on calibration data, recovers a complete set for 73.0% of those states at five items and 80.6% at eight, against 61.4% and 72.4% for Qwen3 embedding with reranking. A matched control ranking by similarity alone reaches 66.6%, placing the gain in the set-level policy, not the computation. From frozen repository source with no gold-derived pool, the lead is 5.0 points. On AMA-Bench it answers from a 76.2% smaller answer prompt, with accuracy 2.08 points above that benchmark's own memory agent. Removing one required group from an otherwise complete set costs 12.3 and 11.1 points of repair-localization precision under two executors. Retrieval for agents is better posed as recovering what a decision lacks than re-ranking what an issue resembles.
Sources
- Less can be More: Relieving RAG Bottlenecks via Evidence Frontloading and Pressure-Adaptive Budgeting
- AgentIR: Reasoning-Aware Retrieval for Deep Research Agents
- From Distillation to Hard Negative Sampling: Making Sparse Neural IR Models More Effective
- RULER: What's the Real Context Size of Your Long-Context Language Models?
- Beyond RAG for Agent Memory: Retrieval by Decoupling and Aggregation
- ContextBench: A Benchmark for Context Retrieval in Coding Agents
- S2G-RAG: Structured Sufficiency and Gap Judging for Iterative Retrieval-Augmented QA
- Stop When Memory Suffices: Evidence-Conditioned Progressive Execution for LLM Agents
- Semantic XPath: Structured Agentic Memory Access for Conversational AI
- MemGPT: Towards LLMs as Operating Systems
- Agent Retrieval Bench: Evaluating Repository Context Retrieval for Coding Agents
- $\tau$-Knowledge: Evaluating Conversational Agents over Unstructured Knowledge
- AutoSearch: Adaptive Search Depth for Efficient Agentic RAG via Reinforcement Learning
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- Agentless: Demystifying LLM-based Software Engineering Agents
- C-Pack: Packed Resources For General Chinese Embeddings
- REVEAL: A Rubric-Guided Agent for Explicit Evidence Sufficiency Verificationin Long-Video Question Answering
- CORE-Bench: A Comprehensive Benchmark for Code Retrieval in the Era of Agentic Coding
- SWE-Explore: Benchmarking How Coding Agents Explore Repositories
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG