Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?
cs.CL
Submitted: 2026-06-29
Updated: 2026-09-02
Code: https://github.com/THU-KEG/RuVerBench
Terminology
Sources
- OctoBench: Benchmarking Scaffold-Aware Instruction Following in Repository-Grounded Agentic Coding
- TRAJECT-Bench:A Trajectory-Aware Benchmark for Evaluating Agentic Tool Use
- Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
- AlpaGasus: Training A Better Alpaca with Fewer Data
- BrowseComp-Plus: A More Fair and Transparent Evaluation Benchmark of Deep-Research Agent
- RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback
- DeepResearch Bench II: Diagnosing Deep Research Agents via Rubrics from Expert Reports
- ResearchRubrics: A Benchmark of Prompts and Rubrics For Evaluating Deep Research Agents
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- JudgeBench: A Benchmark for Evaluating LLM-based Judges
- Large Language Models are not Fair Evaluators
- Self-Instruct: Aligning Language Models with Self-Generated Instructions
- Reframing Instructional Prompts to GPTk's Language
- Long-form factuality in large language models
- AgentSCOPE: Evaluating Contextual Privacy Across Agentic Workflows
- A Comprehensive Survey of Deep Research: Systems, Methodologies, and Applications
- gpt-oss-120b & gpt-oss-20b Model Card
- ResearcherBench: Evaluating Deep AI Research Systems on the Frontiers of Scientific Inquiry
- RubricEval: A Rubric-Level Meta-Evaluation Benchmark for LLM Judges in Instruction Following
- BatchEval: Towards Human-like Text Evaluation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering