Benchmarking LLM-as-a-Judge for Long-Form Output Evaluation
cs.CL
Submitted: 2026-06-01
Updated: 2026-08-28
Code: https://github.com/cjj826/LongJudgeBench
Terminology
Sources
- DeepResearch Bench: A Comprehensive Benchmark for Deep Research Agents
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods
- RewardBench 2: Advancing Reward Model Evaluation
- OpenAI GPT-5 System Card
- Skill Retrieval Augmentation for Agentic AI
- Towards Unification of Hallucination Detection and Fact Verification for Large Language Models
- SurGE: A Benchmark and Evaluation Framework for Scientific Survey Generation
- Kimi K2: Open Agentic Intelligence
- MetaSyn: A Benchmark for LLM Agents on Meta-Analysis Articles from Nature Portfolio
- Qwen3 Technical Report
- Beyond Correctness: Evaluating Subjective Writing Preferences Across Cultures
- GLM-5: from Vibe Coding to Agentic Engineering
- BERTScore: Evaluating Text Generation with BERT
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering