FACET at WMT 2026 Automated Translation Quality Evaluation Task

arXiv:2610.00096 · cs.CL · Submitted 2026-09-08 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "FACET at WMT 2026 Automated Translation Quality Evaluation Task".

Jane: FACET introduces a reference-free method for evaluating machine translation quality by decomposing it into three distinct passes—Fluency, Accuracy, and Consistency—each tailored to specific error types.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: So, to recap, FACET at WMT two thousand twenty-six Automated Translation Quality Evaluation Task is proposing a method that breaks translation evaluation down into Fluency, Accuracy, and Consistency passes to get a better picture of where the translation fails. We're also looking at how the authors are structuring their inputs for each of these specialized checks.

Jane: Exactly, Tom; it’s about giving each check exactly what it needs to judge its specific type of error without overwhelming the model with everything simultaneously. The focus is on making sure we capture different kinds of mistakes more clearly than before.

Lu: The structure they use for input construction is particularly interesting, especially how they handle segment re-segmentation and coordinate mapping to keep track of where things came from in the original text, which is crucial for reliable error tracking.

Meng: I see the focus on stability there; if those segments aren't mapped back correctly to their original positions, you can't accurately pinpoint an error location in a massive translation. That seems like a practical hurdle they had to clear.

Lalam: From my perspective, this technical refinement means that future AI systems will be able to catch errors that are deeply contextual, not just surface-level mistakes in one sentence. That’s where the real improvement for general AI quality lies for us.

The paper's summary: Tom: Looking at the summary of FACET at WMT two thousand twenty-six Automated Translation Quality Evaluation Task, it boils down to this: they use a single model but prompt it three times with different inputs to get error spans, quality scores, and error-free labels for Fluency, Accuracy, and Consistency.

Jane: So the core idea is that the system yields these three shared outputs from just one fixed model acting under asymmetric prompting conditions. It’s a clever way to isolate the different types of errors being measured in each pass.

Lu: The paper explicitly follows the MQM error taxonomy, where Accuracy and Fluency are top levels, and Consistency is treated as a subtype of Fluency because it can't be judged within just one segment; it requires the whole document view.

Meng: That distinction between segment-level checks for fluency and document-level checks for consistency is really important because it highlights that certain errors, like inconsistent naming conventions across multiple segments, simply require a larger context to spot.

Lalam: This decomposition means we move away from a single judgment score and toward a richer set of data points that tell us exactly *why* something failed—whether it was grammar, meaning drift, or just inconsistency. That’s much more useful for iterative improvement.

The paper's improvements: Tom: The improvements the authors suggest focus on how to handle these different input scopes and how to score the results. They detail specific methods for re-segmentation to stabilize iteration and how they manage input lengths for each pass.

Jane: They also lay out a specific scoring mechanism where the final segment score is calculated as the negated sum of the severities of all surviving errors across every pass. That means a segment gets penalized for multiple issues if it has both an accuracy problem and a fluency issue.

Lu: The Consistency pass specifically checks for things like whether a single referent appears under two or more names or if shortened references conflict with their original forms, which is a very targeted way to catch document-level discrepancies.

Meng: I’m looking at the ablation study where they test FACET without the Consistency pass; it shows that while removing it leaves system rankings nearly unchanged, suggesting segment-level errors are generally more influential on the final ranking than document-level ones right now.

Lalam: That finding is insightful because it tells us that for most applications, fixing those segment-level fluency and accuracy issues will have a bigger immediate impact on overall quality than solving the harder document-wide consistency problems.

Conclusion: Tom: So, to wrap up, the FACET at WMT two thousand twenty-six Automated Translation Quality Evaluation Task paper introduces this three-pass approach—Fluency, Accuracy, and Consistency—to provide context-specific error characterization for translation quality assessment.

Jane: We learned that by using asymmetric input and a merged error span aggregation scoring system, we can get detailed metrics on meaning preservation versus linguistic correctness versus internal document consistency.

Lu: The implication is that we gain a reference-free way to characterize predictions without needing gold labels, which is a significant step toward more automated, self-evaluating translation tools.

Meng: Practically speaking, the paper shows us where the current limitations are; they note that because models don't track recurring items well across long inputs yet, document-level consistency errors aren't as impactful on the overall score as segment-level ones.

Lalam: I think the future work points toward using stronger long context models or a scoring scheme that credits document-level consistency directly to solve that reliability gap we discussed.

Tom: Fascinating stuff, Jane; FACET at WMT two thousand twenty-six Automated Translation Quality Evaluation Task gives us a much clearer roadmap for building more robust AI evaluation pipelines moving forward.

Ahrii Kim, Chanjun Park

AI-Bio Convergence Research Inst. School of Software Dept. of Intelligent Semiconductors Soongsil University

cs.CL

Submitted: 2026-09-08

Updated: 2026-09-08

Comments: Accepted at WMT 2026 (shared task system paper)

Code: https://github.com/trotacodigos/facet

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: FACET introduces a reference-free method for evaluating machine translation quality by decomposing it into three distinct passes—Fluency, Accuracy, and Consistency—each tailored to specific error

Key concepts

Fluency
This pass checks if a translated segment sounds natural in the target language, focusing on grammar, spelling, and register. It evaluates the segment as if it were an isolated piece of text to ensure it is well-formed.
Accuracy
The Accuracy pass judges whether the meaning of a translated segment is preserved when comparing the source and target texts. It assesses whether the core information has been correctly transferred between languages.
Consistency
This pass checks for internal consistency across an entire document, looking for recurring issues like inconsistent naming or terminology. It examines how entities or terms are handled throughout the whole text, not just in isolated sentences.

Terminology

Summary

FACET introduces a reference-free method for evaluating machine translation quality by decomposing it into three distinct passes—Fluency, Accuracy, and Consistency—each tailored to specific error types. This approach is significant because it addresses the limitation of traditional LLM evaluators that often prompt models with all evaluation criteria simultaneously, allowing for more precise error characterization based on context-specific evidence.

The gist

FACET decomposes translation quality evaluation into Fluency, Accuracy, and Consistency passes, providing each pass only the context its specific error type requires.

How it works

The system employs a single fixed model prompted three times under asymmetric input to yield the three shared-task outputs: error spans, quality scores, and error-free labels. The passes follow the MQM error taxonomy (Lommel et al., 2013), where Accuracy and Fluency are top-level dimensions, and Consistency is a subtype of Fluency promoted to its own pass because it cannot be judged within a single segment.

The three passes operate under different input scopes:

  1. The Accuracy pass reads both the source and the target to judge whether meaning is preserved.

  2. The Fluency pass reads only the target, judging whether a segment is well-formed in the target language, covering properties like spelling, grammar, and register.

  3. The Consistency pass reads only the whole document to judge internal consistency regarding recurring items like named entities, domain terms, and fixed expressions.

Input Construction

The paper details how inputs are managed for each pass to fix their scope:

. For the Fluency and Accuracy passes, segments are re-segmented into sentences using a rule-based multilingual segmenter. Newlines are replaced with spaces in a length-preserving substitution so that the character span maps back directly onto the original text coordinates. This provides a stable structure to iterate over, which reduces its tendency to skip material in long inputs.

. For the Consistency pass, segments sharing a document identifier are grouped and ordered to reconstruct the full document. The system allows for different input lengths based on the pass: 1.5× input plus a floor of 1,000 tokens for Fluency and Accuracy, capped at 8,000 tokens; and 2× input plus a floor of 2,000 tokens for the document-level Consistency pass, capped at 16,000 tokens.

Error Characterization and Scoring

Each pass rates errors on a shared four-point scale (Scale 1 to Scale 4), with Scale 3 being the cut point for Major errors. The final segment score is calculated as the negated sum of the severities of its surviving errors across all passes. A segment scores exactly zero when no error survives, meaning it is labeled error-free.

The Consistency pass specifically identifies divergences in recurring items. It checks if a single referent appears under two or more names, and whether shortened or pronominal references agree with the form they refer back to. The paper notes that Consistency errors at Scale 1 are discarded on morphologically rich languages, focusing instead on higher-severity divergences that reflect cross-segment errors.

System Analysis and Findings

The authors characterize the predictions of FACET without gold labels through system rankings. Their primary submission, FACET, places post-edited human translation first in every language pair. While the Consistency pass changes about a tenth of segment scores, it leaves the overall system ranking nearly unchanged, suggesting that document-level consistency errors are rare relative to segment-level scoring. The authors attribute this effect to the limited reliability with which current models track recurring items across long inputs and the mismatch between a document-level signal and segment-level scoring.

Ablation Study

To measure the contribution of the Consistency pass, FACET−C is submitted, omitting the Consistency pass. Comparing FACET and FACET−C shows that while the Consistency pass changes 10.1% of segment scores, it leaves system rankings almost untouched (no system moves by more than 0.2 in mean rank). This indicates that document-level consistency errors are not as influential on overall quality ranking as segment-level Fluency and Accuracy errors. The paper concludes that stronger long-context models or a scoring scheme that credits document-level Consistency directly would be needed to address the current limitations regarding long context reliability.

Qualitative Example

The authors present a full Czech–German document evaluation where Figure 2 illustrates how the three passes operate. Fluency and Accuracy act within a segment, while Consistency acts across the document, demonstrating how it catches an error of a different kind—a divergence in naming conventions between segments that cannot be detected by segment-level scoring alone. For instance, one facility named Bürstenfabrik becomes Schleiferei in the next segment, an error class only detectable by the Consistency pass.

Improvements for AI systems

Based on the FACET paper, here are specific improvements that can be made to existing AI translation systems:

  1. Improve Translation Quality Evaluation by Decomposing Error Analysis into Distinct Passes: Instead of using a single prompt that asks a model to judge fluency, accuracy, and consistency simultaneously (which leads to all-or-nothing judgment), systems should be redesigned to use the FACET decomposition.

  2. Implement Asymmetric Input for Specialized Evaluation: Design evaluation pipelines where different error types are assessed using the context they require:

  3. Apply Context-Specific Evaluation Passes:

  4. Accuracy Pass Improvement (Source/Target Comparison): Improve systems by explicitly separating meaning preservation checks from linguistic correctness checks. This allows developers to pinpoint whether a failure is due to semantic drift (Accuracy) or grammatical error (Fluency).

  5. Consistency Check for Long Documents: For document-level tasks, implement a dedicated Consistency pass that reads the entire target text to track recurring entities and domain terms across segments. This system will catch errors where one entity is rendered under multiple names throughout a long text, even if each individual segment passes local checks.

  6. Error Span Localization with Original Coordinates: Enhance error reporting by requiring models to return exact character offsets within the original source/target text, rather than relying on abstract model-reported offsets. This allows downstream systems to precisely locate and correct errors in the original document format.

  7. Refine Scoring Mechanism via Merged, Segment-Level Error Aggregation: Instead of letting one pass overwrite another's results, implement a scoring mechanism that merges predictions from the three passes for each segment. The final quality score should be derived from the negated sum of severities of surviving errors across all passes (e.g., Score = -Sum(Severities)). This ensures that a segment is penalized for multiple types of errors simultaneously.

  8. Contextual Filtering and Sensitivity Tuning: Incorporate rules to filter out non-error phenomena (like fillers, hesitations, or common grammatical variations) before passing text to the strict passes (Fluency and Accuracy). This prevents false positives from preference rather than true error flags.

  9. Ablation Studies for Feature Contribution Analysis: When developing new evaluation metrics or models, use ablation studies (like FACET vs. FACET-C) to quantify exactly how much each component (e.g., Consistency pass) contributes to the final quality score and system ranking, helping researchers understand the value of different error types in practice.

  10. System Ranking Benchmarking: Use a multi-faceted evaluation approach to rank translation systems, incorporating mean scores across multiple language pairs rather than averaging results globally. This provides a more robust measure of a system's performance within specific linguistic contexts.

Abstract

Different error types in machine translation require different evidence. Whether meaning is preserved can be judged only against the source, while whether the target is well-formed, or whether it names one entity consistently, can be judged from the target alone. We present FACET, our reference-free submission to the WMT26 Automated Translation Quality Evaluation Task, which decomposes evaluation into Fluency, Accuracy, and Consistency passes and gives each pass only the context its error type requires. A single fixed model is prompted three times, and the merged error spans yield the three task outputs, error spans, quality scores, and error-free labels, with no trained components. We also submit FACET-C, which omits the Consistency pass. Without gold labels, we characterize the predictions of FACET. Its system rankings place post-edited human translation first, and the Consistency pass changes about a tenth of segment scores while leaving the ranking nearly unchanged.

Sources

Related papers