Can LLM-as-a-Judge Reliably Verify Rubrics in Agentic Scenarios?

arXiv:2606.29920 · cs.CL · Submitted 2026-06-29 · Read on arXiv

cs.CL

Submitted: 2026-06-29

Updated: 2026-09-02

Code: https://github.com/THU-KEG/RuVerBench

Terminology

Sources

Related papers