Designing Reliable LLM-as-a-Judge Measurement Systems for Multi-Turn Business Agents
cs.AI
Submitted: 2026-09-27
Updated: 2026-09-27
Terminology
Sources
- HealthBench: Evaluating Large Language Models Towards Improved Human Health
- EvalAssist: A Human-Centered Tool for LLM-as-a-Judge
- $\tau^2$-Bench: Evaluating Conversational Agents in a Dual-Control Environment
- Can Large Language Models Be an Alternative to Human Evaluations?
- Ask, Don't Judge: Binary Questions for Interpretable LLM Evaluation and Self-Improvement
- Humans Hallucinate Too: Language Models Identify and Correct Subjective Annotation Errors With Label-in-a-Haystack Prompts
- Detecting Label Errors by using Pre-Trained Language Models
- TICKing All the Boxes: Generated Checklists Improve LLM Evaluation and Generation
- Shopping Reasoning Bench: An Expert-Authored Benchmark for Multi-Turn Conversational Shopping Assistants
- Building Customer Support AI Agents at 100M-User Scale: An Evaluation-Driven Framework
- LLM-Rubric: A Multidimensional, Calibrated Approach to Automated Evaluation of Natural Language Texts
- CRMArena: Understanding the Capacity of LLM Agents to Perform Professional CRM Tasks in Realistic Environments
- Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
- MT-Eval: A Multi-Turn Capabilities Evaluation Benchmark for Large Language Models
- How to Correctly Report LLM-as-a-Judge Evaluations
- CheckEval: A reliable LLM-as-a-Judge framework for evaluating text generation using checklists
- From Generation to Judgment: Opportunities and Challenges of LLM-as-a-judge
- DnA-Eval: Enhancing Large Language Model Evaluation through Decomposition and Aggregation
- G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment
- HD-Eval: Aligning Large Language Model Evaluators Through Hierarchical Criteria Decomposition
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection