RubricRefine: Improving Tool-Use Agent Reliability with Training-Free Pre-Execution Refinement
cs.LG, cs.SE
Submitted: 2026-05-10
Updated: 2026-09-09
Code: https://github.com/huggingface/smolagents
Project page: https://llm-as-a-verifier.notion.site
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Self-Distillation Zero: Self-Revision Turns Binary Rewards into Dense Supervision
- Executable Code Actions Elicit Better LLM Agents
- PreFlect: From Retrospective to Prospective Reflection in Large Language Model Agents
- ToolACE: Winning the Points of LLM Function Calling
- Language Models (Mostly) Know What They Know
- FunReason: Enhancing Large Language Models' Function Calling via Self-Refinement Multiscale Loss and Automated Data Refinement
- Nemotron-Research-Tool-N1: Exploring Tool-Using Language Models with Reinforced Reasoning
- ReTool: Reinforcement Learning for Strategic Tool Use in LLMs
- Chain-of-Tools: Utilizing Massive Unseen Tools in the CoT Reasoning of Frozen Language Models
- Graph RAG-Tool Fusion
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Constitutional AI: Harmlessness from AI Feedback
- Teaching Large Language Models to Self-Debug
- CodeT: Code Generation with Generated Tests
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
- Code Generation with AlphaCodium: From Prompt Engineering to Flow Engineering
- Rubrics as Rewards: Reinforcement Learning Beyond Verifiable Domains
- Agentic Rubrics as Contextual Verifiers for SWE Agents
- Enabling Calibration In The Zero-Shot Inference of Large Vision-Language Models
- Chasing the Tail: Effective Rubric-based Reward Modeling for Large Language Model Post-Training
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks