BudgetVerify: Budget-Tiered Verification for Financial QA
cs.AI, cs.LG
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/vllm-project/semantic-router
Terminology
Sources
- Large Language Monkeys: Scaling Inference Compute with Repeated Sampling
- Why Do Multi-Agent LLM Systems Fail?
- FrugalGPT: How to Use Large Language Models While Reducing Cost and Improving Performance
- RM-R1: Reward Modeling as Reasoning
- FinanceBench: A New Benchmark for Financial Question Answering
- Let's Verify Step by Step
- RouteLLM: Learning to Route LLMs with Preference Data
- Hard2Verify: A Step-Level Verification Benchmark for Open-Ended Frontier Math
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- CARROT: A Cost Aware Rate Optimal Router
- S2J: Bridging the Gap Between Solving and Judging Ability in Generative Reward Models
- Self-Consistency Improves Chain of Thought Reasoning in Language Models
- RL Tango: Reinforcing Generator and Verifier Together for Language Reasoning
- Adaptive Test-Time Compute Allocation for Reasoning LLMs via Constrained Policy Optimization
- Generative Verifiers: Reward Modeling as Next-Token Prediction
- Variation in Verification: Understanding Verification Dynamics in Large Language Models
- Evaluating Judges as Evaluators: The JETTS Benchmark of LLM-as-Judges as Test-Time Scaling Evaluators
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection