Reference-Free Post-Training of Open Large Language Models for Multilingual Machine Translation

arXiv:2608.10812 · cs.CL, cs.AI · Submitted 2026-08-12 · Read on arXiv

Chris Han, Pengzhi Gao, Pei Fu, Jian Luan

Xiaomi Inc.

cs.CL, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/xiaomi-research/gemmax

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 85/100

The gist: The paper studies reference-free post-training for multilingual machine translation with open large language models.

Terminology

Summary

The paper studies reference-free post-training for multilingual machine translation with open large language models. Starting from the supervised fine-tuned MiLMMT-46-v0.1 models, the authors apply Group Relative Policy Optimization (GRPO) with a reward that averages two reference-free quality estimation models (XCOMET and COMETKiwi) and is gated by language identification using OpenLID-v3. The reward is defined as the average of the two quality scores if the predicted language of the translation matches the intended target language, and zero otherwise. The RL dataset is derived from the MiLMMT SFT data by discarding reference translations, yielding 263,982 instances across 192 translation directions, which is then filtered to 31,572 instances based on reward mean and standard deviation thresholds. After RL training, the authors linearly interpolate the SFT and RL checkpoints to obtain MiLMMT-46-v1.0, with an interpolation coefficient α = 0.5 at all three model scales (1B, 4B, and 12B).

Across 46 languages, the resulting models consistently improve translation quality over their SFT counterparts. Averaged over the three model scales, XCOMET and COMETKiwi scores on WMT24++ improve by 2.75 and 2.44 points, respectively. On FLORES+, reference-based XCOMET, reference-free XCOMET, and COMETKiwi improve by 1.17, 1.41, and 1.17 points, while spBLEU decreases by 1.21 points. The models outperform strong recent open baselines including Seed-X, HY-MT2, and TranslateGemma, and achieve leading reference-free scores against evaluated proprietary systems such as Google Translate, Gemini 3 Pro, and GPT-5. MiLMMT-46-12B-v1.0 achieves the highest reference-free XCOMET and COMETKiwi scores on WMT24++ and across all four FLORES+ direction groups, and MiLMMT-46-1B-v1.0 outperforms TranslateGemma-4B on every reported metric despite using only one quarter as many parameters.

The interpolation analysis shows a consistent trade-off between spBLEU and reference-based XCOMET across all three model scales. As the SFT weight α increases, spBLEU increases monotonically while reference-based XCOMET decreases. The selected α = 0.5 checkpoints recover 2.79, 3.70, and 4.21 spBLEU points at the 1B, 4B, and 12B scales, respectively, while reducing reference-based XCOMET by only 0.53, 0.56, and 0.54 points relative to the RL endpoints.

The paper further investigates on-policy distillation (OPD) as an alternative post-training approach, using MiLMMT-46-12B-v1.0 as a teacher to train 1B and 4B students. The authors adopt the policy-gradient form of OPD (PG-OPD) and compare three training variants: OPD alone, RL+OPD combining GRPO and distillation losses, and OPD initialized from v1.0. The results show that OPD closely matches the 4B baseline on FLORES+, achieving 34.17 versus 33.96 spBLEU and 90.82 versus 90.91 reference-based XCOMET, while the 1B student remains slightly behind in reference-based XCOMET (85.22 versus 85.94). Varying the distillation weight λ in RL+OPD produces a trade-off similar to checkpoint interpolation. Overall, OPD reaches but does not surpass the quality frontier achieved by RL with checkpoint interpolation. The authors release the models and code to facilitate future research.

Improvements for AI systems

Improvements to AI Systems:

  1. Reference-Free Quality-Gated Reward for RLHF/GRPO: Implement a reward function that averages multiple reference-free quality estimators (e.g., XCOMET + COMETKiwi) and gates the reward to zero when the output language mismatches the target (using a language ID classifier). This eliminates the need for human references during post-training, reduces reward hacking, and enforces language fidelity—improving multilingual MT without costly parallel data.

  2. Checkpoint Interpolation for Trade-Off Control: After RL, linearly interpolate the SFT and RL checkpoints (α=0.5) to recover lexical accuracy (spBLEU) while retaining most semantic quality (XCOMET). This provides a simple, compute-free method to balance fluency vs. faithfulness, applicable to any seq2seq model—enabling fine-grained control over output style without retraining.

  3. Reward-Filtered RL Dataset Curation: Filter the RL training set by keeping only instances whose reward mean and standard deviation fall within thresholds (e.g., discard low-quality or high-variance examples). This reduces noise and improves training stability, making RL more sample-efficient for low-resource directions.

  4. On-Policy Distillation (PG-OPD) as a Lightweight Alternative: Use policy-gradient on-policy distillation from a larger RL-tuned teacher (12B) to train smaller students (1B/4B), matching or approaching teacher quality without full RL. This enables efficient deployment of high-quality MT on edge devices with 4x fewer parameters, preserving semantic scores while sacrificing minimal spBLEU.

  5. Language-Gated Zero-Reward Penalty for Multilingual Generalization: Integrate the OpenLID-v3 gating into any multilingual generation task (e.g., summarization, dialogue) to prevent off-target language outputs—improving reliability in zero-shot or low-resource settings.

What the Improved AI System Can Do:

  • Translate across 46 languages with higher semantic accuracy (XCOMET/COMETKiwi) than SFT baselines, while maintaining controllable lexical fidelity via interpolation.

  • Operate without reference translations during training, enabling post-training on any unlabeled multilingual corpus.

  • Deploy a 1B-parameter model that outperforms a 4B proprietary baseline (TranslateGemma) on all reported metrics, reducing inference cost by 75%.

  • Automatically reject translations that switch to the wrong language, ensuring output consistency for multilingual chatbots or document translation.

  • Distill large-model quality into smaller models for on-device use, achieving near-teacher performance without the need for expensive RL compute.

Sources

Related papers