Machine Translation Hallucination Detection for Low and High Resource Languages using Large Language Models
Kenza Benkirane, Laura Gongas, Shahar Pelles, Naomi Fuchs, Joshua Darmon, Pontus Stenetorp, David Ifeoluwa Adelani, Eduardo Sánchez
cs.CL, cs.AI
Submitted: 2024-10-20
Updated: 2026-08-19
Comments: Authors Kenza Benkirane and Laura Gongas contributed equally to this work
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Recent advancements in massively multilingual machine translation systems have significantly enhanced translation accuracy; however, even the best performing systems still generate hallucinations,
Terminology
Abstract
Recent advancements in massively multilingual machine translation systems have significantly enhanced translation accuracy; however, even the best performing systems still generate hallucinations, severely impacting user trust. Detecting hallucinations in Machine Translation (MT) remains a critical challenge, particularly since existing methods excel with High-Resource Languages (HRLs) but exhibit substantial limitations when applied to Low-Resource Languages (LRLs). This paper evaluates sentence-level hallucination detection approaches using Large Language Models (LLMs) and semantic similarity within massively multilingual embeddings. Our study spans 16 language directions, covering HRLs, LRLs, with diverse scripts. We find that the choice of model is essential for performance. On average, for HRLs, Llama3-70B outperforms the previous state of the art by as much as 0.16 MCC (Matthews Correlation Coefficient). However, for LRLs we observe that Claude Sonnet outperforms other LLMs on average by 0.03 MCC. The key takeaway from our study is that LLMs can achieve performance comparable or even better than previously proposed models, despite not being explicitly trained for any machine translation task. However, their advantage is less significant for LRLs.
Sources
- SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
- Beyond English-Centric Multilingual Machine Translation
- xCOMET: Transparent Machine Translation Evaluation through Fine-grained Error Detection
- Better Zero-Shot Reasoning with Role-Play Prompting
- MADLAD-400: A Multilingual And Document-Level Large Audited Dataset
- Unsupervised Real-Time Hallucination Detection based on the Internal States of Large Language Models
- No Language Left Behind: Scaling Human-Centered Machine Translation
- Enhancing Uncertainty-Based Hallucination Detection with Stronger Focus
- Multilingual Machine Translation with Large Language Models: Empirical Results and Analysis
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering