Last Translation Benchmark
cs.CL
Submitted: 2026-09-03
Updated: 2026-09-29
Code: https://github.com/zouharvi/last-translation-benchmark
Terminology
Sources
- Achieving Human Parity on Automatic Chinese to English News Translation
- Beyond Accuracy: Community Perspectives on Machine Translation
- Searching the Internet for Challenging Benchmarks at Scale
- Pearmut: Human Evaluation of Translation Made Trivial
- Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness
- Play Favorites: A Statistical Method to Measure Self-Bias in LLM-as-a-Judge
- Contrastive ESA: Human Evaluation of Multiple Translations at Once
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- Gemma 4 Technical Report
- Hy-MT2: A Family of Fast, Efficient and Powerful Multilingual Translation Models in the Wild
- Command A: An Enterprise-Ready Large Language Model
- Seed-X: Building Strong Multilingual Translation LLM with 7B Parameters
- Nemotron 3 Ultra: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- The Llama 3 Herd of Models
- When LLMs Benchmark Themselves: Deconstructing Self-Bias in Automated Evaluation
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering