ExecRubrics: Executable Tool-Augmented Rubrics for Verifiable and Efficient Long-Form Evaluation
cs.AI, cs.CL, cs.IR
Submitted: 2026-08-23
Updated: 2026-08-31
Comments: Accepted to EMNLP 2026 Findings
License: http://creativecommons.org/licenses/by/4.0/
The gist: Rubrics aim to make language-model evaluation transparent by decomposing response quality into interpretable criteria.
Terminology
Abstract
Rubrics aim to make language-model evaluation transparent by decomposing response quality into interpretable criteria. However, natural-language rubrics are often ambiguous, require LLM judges, and typically assume criteria aggregated through linear weighted sums, limiting their ability to capture dependencies, alternatives, penalties, and override conditions. We propose ExecRubrics, a framework for representing rubrics as compact executable programs. ExecRubrics encodes evaluation logic as verifiable Python scoring functions, giving natural-language rubric intent an operational semantics: a fixed decision procedure that can be inspected, executed, and edited. On three long-form response benchmarks -- HealthBench, HelpSteer, and ArgQuality -- we show that ExecRubrics can recover substantial preference signal without an LLM judge at evaluation time. On ArgQuality and HelpSteer, the strongest executable variants are within 1.1 and 4 percentage points, respectively, of the direct GPT-5.5 agentic baseline. Executable rubrics are also considerably faster, achieving a 192x average speedup. We show that incorporating external logic and resources from text processing libraries such as NLTK and spaCy can further improve preference accuracy. Our results suggest a novel way of approaching automated evaluation, by offering a faster, more explainable, and less ambiguous alternative to black-box rubric evals, particularly in high-stakes domains such as healthcare and banking where precision and auditability are critical.
Sources
- PRBench: Large-Scale Expert Rubrics for Evaluating High-Stakes Professional Reasoning
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
- TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
- How Fine-Grained Should a RAG Benchmark Be? A Hierarchical Framework for Synthetic Question Generation
- Open-World Evaluations for Measuring Frontier AI Capabilities
- DoGMaTiQ: Automated Generation of Question-and-Answer Nuggets for Report Evaluation
- Humanity's Last Exam
- Data-driven Circuit Discovery for Interpretability of Language Models
- RUBEN: Rule-Based Explanations for Retrieval-Augmented LLM Systems
- ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and Judge
- Qwen3 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection