Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation

arXiv:2606.01629 · cs.CL · Submitted 2026-06-01 · Read on arXiv

cs.CL

Submitted: 2026-06-01

Updated: 2026-08-28

Code: https://github.com/cjj826/LongJudgeBench

Terminology

Sources

Related papers