Can We Triage LLM Translation Errors in Classical Texts Without Human References? Source Novelty, GEMBA Scoring, and Budgeted Review through Pali-to-English Translation
cs.CL
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: Preprint. This article has not yet been peer reviewed
Code: https://github.com/suttacentral/bilara-data
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Latin BERT: A Contextual Language Model for Classical Philology
- Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
- Large Language Models Are State-of-the-Art Evaluators of Translation Quality
- GEMBA-MQM: Detecting Translation Quality Error Spans with GPT-4
- Gemini Embedding: Generalizable Embeddings from Gemini
- The Multi-Range Theory of Translation Quality Measurement: MQM scoring models and Statistical Quality Control
- PaliBench: A Multi-Reference Blueprint for Classical Language Translation Benchmarks
- Participatory Research for Low-resourced Machine Translation: A Case Study in African Languages
- Learning from Many Voices: Literary MT Using Multi-Reference Human and Synthetic Data
- Evaluating LLM-Based Translation of a Low-Resource Technical Language: The Medical and Philosophical Greek of Galen
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering