Large Language Models for Machine Translation Quality Annotation: Humans and Models Are Both Challenged
cs.CL
Submitted: 2026-10-07
Updated: 2026-10-07
Terminology
Sources
- The Alternative Annotator Test for LLM-as-a-Judge: How to Statistically Justify Replacing Human Annotators with LLMs
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- DeepSeek-V3 Technical Report
- xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection
- LLMs as Span Annotators: A Comparative Study of LLMs and Humans
- gpt-oss-120b & gpt-oss-20b Model Card
- MQM Re-Annotation: A Technique for Collaborative Evaluation of Machine Translation
- Enhancing Human Evaluation in Machine Translation with Comparative Judgment
- Are Large Reasoning Models Good Translation Evaluators? Analysis and Performance Boost
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering