FineVerify: Scaling Test-Time Compute with Fine-Grained Self-Verification for Agentic Search
cs.CL
Submitted: 2026-05-30
Updated: 2026-09-01
Code: https://github.com/XuZhao0/fineverify
Terminology
Sources
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- Training Verifiers to Solve Math Word Problems
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- RE-Searcher: Robust Agentic Search with Goal-oriented Planning and Self-reflection
- DeepSearchQA: Bridging the Comprehensiveness Gap for Deep Research Agents
- xbench: Tracking Agents Productivity Scaling with Profession-Aligned Real-World Evaluations
- Winning Gold at IMO 2025 with a Model-Agnostic Verification-and-Refinement Pipeline
- Agentic Aggregation for Parallel Scaling of Long-Horizon Agentic Tasks
- WebGPT: Browser-assisted question-answering with human feedback
- ParallelMuse: Agentic Parallel Thinking for Deep Information Seeking
- Learning to Reason Across Parallel Samples for LLM Reasoning
- WebResearcher: Unleashing unbounded reasoning capability in Long-Horizon Agents
- DeepSeekMath-V2: Towards Self-Verifiable Mathematical Reasoning
- OpenAI GPT-5 System Card
- Scaling Long-Horizon LLM Agent via Context-Folding
- Tongyi DeepResearch Technical Report
- Evaluating Stochasticity in Deep Research Agents
- Recursive Language Models
- BrowseComp: A Simple Yet Challenging Benchmark for Browsing Agents
- A Survey on Test-Time Scaling in Large Language Models: What, How, Where, and How Well?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering