TACTICS: Taxonomy-Aware Intelligent Corpus Sampling for Machine Translation
cs.CL, cs.LG
Submitted: 2026-09-16
Updated: 2026-10-06
Comments: Accepted at The Eleventh Conference in Machine Translation 2026 (WMT2026)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Large-scale machine-translation (MT) systems are typically evaluated on random samples from a corpus whose distributional composition is an artifact of how it was assembled.
Terminology
Abstract
Large-scale machine-translation (MT) systems are typically evaluated on random samples from a corpus whose distributional composition is an artifact of how it was assembled. Such a sample inherits the phenomena the collection happens to contain rather than the full space a system must handle, spanning rule-governed conventions (terminology, punctuation, currency formatting) and context-dependent phenomena (tone, honorifics, document-level coherence), and thus provides no coverage guarantee for assessing robustness. We propose TACTICS (Taxonomy-Aware Coverage-opTimized Intelligent Corpus Sampling), which recasts coverage as an explicit objective. TACTICS induces a hierarchical taxonomy from a locale style guide, classifies segments against it, and selects a fixed-budget subset jointly optimizing coverage of rare categories, document-level coherence, and distributional fidelity to the full corpus. Applied to MT evaluation across four translation directions, TACTICS improves coverage of rare categories over lexical and embedding-based selection. By targeting the phenomena that separate systems, TACTICS makes a fixed evaluation budget go further, recovering the true system ranking from far fewer segments than random sampling wherever a real quality gap exists and never signaling a difference where none exists.
Sources
- Language Models are Few-Shot Learners
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- When Less is More: Investigating Data Pruning for Pretraining LLMs at Scale
- No Language Left Behind: Scaling Human-Centered Machine Translation
- GPT-4 Technical Report
- COMET: A Neural Framework for MT Evaluation
- BLEURT: Learning Robust Metrics for Text Generation
- Balanced Data Sampling for Language Model Training with Clustering
- D4: Improving LLM Pretraining via Document De-Duplication and Diversification
- LLaMA: Open and Efficient Foundation Language Models
- TnT-LLM: Text Mining at Scale with Large Language Models
- Data Selection for Language Models via Importance Resampling
- Chain-of-Layer: Iteratively Prompting Large Language Models for Taxonomy Induction from Limited Examples
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering