Evaluating the Evaluator: Summarization Metrics and LLM-Judges beyond English
cs.CL, cs.AI
Submitted: 2025-03-21
Updated: 2026-09-02
Code: https://github.com/hitz-zentroa/summarization
Terminology
Sources
- GPT-4 Technical Report
- Atla Selene Mini: A General Purpose Evaluation Model
- The Llama 3 Herd of Models
- A Survey on LLM-as-a-Judge
- An Empirical Study of LLM-as-a-Judge for LLM Evaluation: Fine-tuned Judge Model is not a General Substitute for GPT-4
- Exploring the Limit of Outcome Reward for Learning Mathematical Reasoning
- Qwen2.5 Technical Report
- Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models
- Language Models are Multilingual Chain-of-Thought Reasoners
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering