When the Knowledge Base Becomes the Gold Standard: Measuring Resource-Shared Evaluation Loops in Entity-Level Machine Translation
Jinhyung Bae, Dain Kil, Seongmin Oh, Seungmin Lee
Hankuk University of Foreign Studies
cs.CL, cs.DL
Submitted: 2026-08-12
Updated: 2026-08-13
Comments: 21 pages, 3 figures. Code and model outputs: https://github.com/nepersoned/malmoi-sjw-eval
Code: https://github.com/nepersoned/malmoi-sjw-eval
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: The paper investigates a resource-shared evaluation loop in entity-level machine translation of the Seungjeongwon Ilgi, a UNESCO Memory of the World record that is only 37.4% translated.
Terminology
Summary
The paper investigates a resource-shared evaluation loop in entity-level machine translation of the Seungjeongwon Ilgi, a UNESCO Memory of the World record that is only 37.4% translated. The central issue: "Low-resource historical domains have no expert gold standard for entity translation, so practitioners substitute a knowledge base (KB) for the gold. That KB is the same resource injected into the system: scoring becomes self-referential and the metric measures instruction compliance rather than translation quality."
The paper demonstrates this with a worked example: the hanja library renders 沈 as chim and 金 as geum, but as surnames the correct readings are Sim and Kim. When the system is injected with the wrong reading 침열 (Chim-yeol)
for 沈悅, and the gold standard is built from the same library, the output is counted correct: The system is ordered to emit a wrong name, and a gold standard built from the same library certifies it.
Gold independence measurement: Of 527 expert-annotated mentions, only 31.1% lie outside the injection pipeline.
For NER Recall, independence is 0%. The residual loop is not uniform: in the overlapping segment the injected reading agrees with the human translation 97.8% of the time against 70.1% in the independent one, so the segment that looks healthiest is the one the loop is holding up.
Only 4.3% of the independent segment is KB-registered, meaning mentions an NER model finds and experts tag are well-attested figures, while those only experts tag are obscure ones absent from the KB.
Distortion bounding: Disagreement between injected/scored readings and human translations is 2.2% in the KB-registered segment but 31.2% in the hanja-transliteration fallback segment, totaling 10.8% overall. Linear projection gives 10.8% at 70% coverage, 16.7% at 50%, 22.5% at 30% and 28.3% at 10%.
Three-level decline under changing gold provenance (Qwen3-8B, NERinject condition): with the complete loop (NER Recall), ETS = 0.944; with partial loop (expert tags), ETS = 0.765; with no loop (independent segment), ETS = 0.348. The baseline independent segment scores 0.342, so under an independent gold NERinject is effectively indistinguishable from baseline. The reportable figure is 2.7× the honest one.
Difference-in-differences results (240 common documents): The loop contribution ranges from +0.722 (Qwen3-4B) to +0.294 (gemma-4-26b), with Δ independent at or below zero for all models (−0.026 for three models, −0.068 for gemma-4-26b). The confinement of the gain is thus established by paired testing rather than by a correlation coefficient.
McNemar's exact test shows the overlap segment is overwhelmingly significant for all four models (improvement-to-regression counts of 171:1, 160:0, 98:3, 56:1; p from 5.8e-50 to 8.1e-16), while in the independent segment three models are non-significant (3:6, p=0.508) and gemma-4-26b is significantly worse (0:8, p=0.008). No model shows an improvement there.
Correlation between baseline ETS and loop contribution: r = −0.938.
Ceiling clustering: Post-injection preservation clusters in a narrow 0.910–0.996 band (width 0.086) while pre-injection performance varies from 0.213–0.770 (width 0.557). Pre-injection performance varies 6.5× more than the ceiling. The reported gain is therefore governed by prior performance rather than by the ceiling, and weaker models appear to improve more dramatically.
Spillover is negative, not zero: "Injection does not generalise beyond the injected list and, in the stronger model, is mildly harmful: attention to the injected names appears to come at the expense of the remaining ones. That the spillover is negative rather than zero was not predicted."
Reading provenance does not distort the gain: Holding the denominator fixed and varying only reading provenance, all four models fall within 0.02 of zero difference (inflation −0.006, −0.017, −0.003, −0.006). The inflation arises entirely on the set axis.
However, KB-derived golds penalize more capable models more heavily: ETS-KB minus ETS-REF widens from −0.017 (Qwen3-4B) to −0.068 (gemma-4-26b).
The design separates injection and scoring pipelines: injection uses SillokBERT-NER over the source with KB readings; scoring uses NIKH expert idx person tags with KB readings (else hanja transliteration). The evaluation set comprises 300 documents drawn from 62,476 parallel pairs of the Injo reign via proportional stratified sampling over length buckets × style patterns, seed 42. A BLEU ≥ 20 filter was applied to remove merge errors from date-identifier joining, not for model selection.
Four conditions are tested: baseline (minimal prompt), few-shot (five fixed style exemplars), NERfix (rule-based post-hoc name correction), and NERinject (Hanja→Hangul name block inserted before the source). The few-shot exemplars share only one incidental substring match with the evaluation gold, making few-shot a valid control.
An independent 300-document sample removing only the BLEU filter (zero id intersection with the original) shows the independence rate is stable (31.1% vs 31.8%). The filter cleaned data: misalignment candidates drop from 14/300 (4.7%) to 0/300. Removing the 14 misaligned pairs raises Qwen3-8B from 11.29 to 13.59, so 2.30 of the 2.79 gap (83%) disappears.
The difference-in-differences replicates within model while discriminating between models: "Within-model reliability — for a given model the two samples differ by 0.003–0.015 and the confidence intervals overlap... Between-model discrimination — within a given sample the two models differ by 0.21–0.23 and the confidence intervals do not overlap. The paper concludes:
The measurement reflects a property of the model, not of the sample."
Classical Chinese pretraining does not transfer: Qwen3-8B, which carries a Classical Chinese (文言文) corpus, scores below a 2B-class model without one
(baseline ETS 0.396 vs 0.499 for gemma-3n-E2B). Sharing the Han script is not sufficient: the Korean readings of Joseon person names conflict with Chinese ones.
Style guidance and entity accuracy are not a trade-off: Only one of the four measurements is significant (Qwen3-8B, original set, p=0.003), it fails to replicate on the independent sample (p=0.386), and the remaining three are non-significant with inconsistent sign.
BLEU improves consistently (+2.29 to +2.71) across all four model × sample combinations.
A misread name is not a surface error: The error does not merely assert a falsehood; it removes the entry from the net of archival search.
Two cases illustrate this: 沈悅 rendered as Chim-yeol (no such person exists in the Joseon biographical record) instead of Sim-yeol, and 金瑬 rendered as Geum-ryu instead of Kim Ryu (a principal figure in the 1623 Injo Restoration), causing the entry to drop out of any search tracing the movements of the restoration's meritorious subjects.
The served behavior of the alias gemma-4-26b-a4b-it changed during the study: Under identical code, parameters and prompts, mean output length fell from 244 to 34 characters and 185 of 300 outputs became empty: the model emits reasoning whose tokens consume the output budget.
Reasoning cannot be disabled on this endpoint; raising the token budget to 8,192 resolved most cases but 34 documents did not complete even at 32,768 tokens. An unversioned commercial API alias does not preserve reproducibility even when the code is archived.
-
Build the evaluation key from a resource disjoint from the injected one.
-
Where separation is impossible, report KB coverage and projected distortion.
-
Do not report an entity metric as a single number; separate overlapping from independent segments.
-
When a weak model shows a large gain, check for a ceiling effect first.
-
Do not compare absolute performance across models with a KB-derived gold.
The paper acknowledges: prior-art coverage relies on web indices with residual risk of unindexed precedent; absolute BLEU is not comparable to prior work; the injection gold is a reconstruction (the original was overwritten by a same-named file); false positives are defined over Injo-period KB names only (2,533 entries); gemma-4-26b covers only 240/300 documents; the unfiltered replication covers only two mid-sized models; the style comparison is blind but author-conducted and preliminary.
Improvements for AI systems
Based on this paper, I can make the following specific improvements to AI systems:
The improved AI system can automatically detect when its training data or knowledge base is used as the evaluation gold standard, flagging scores as potentially inflated. It can quantify the overlap between injected knowledge and evaluation data, and report separate metrics for overlapping vs. independent segments rather than a single aggregate score.
The system can distinguish between transliteration (character-by-character conversion) and semantic reading (correct name pronunciation in context), using historical/biographical context to resolve ambiguities like 沈 as Sim (surname) vs. chim (general reading). It can cross-reference against biographical databases to verify that a translated name corresponds to a real historical figure.
When a weak model shows large improvements from knowledge injection, the system can automatically check whether the gain is confined to a narrow performance band (0.910–0.996 in this study) and report the gain relative to pre-injection variance, preventing overstatement of model improvements.
The system can track whether knowledge injection improves only the injected items while degrading performance on non-injected ones (negative spillover observed in stronger models), and report both targeted and generalizable improvements separately.
The system can detect when a commercial API alias changes behavior mid-study (e.g., output length dropping from 244 to 34 characters), automatically flag reproducibility failures, and require versioned endpoints or local model snapshots for scientific claims.
The system can automatically check whether evaluation gold standards are built from resources disjoint from those used for training or prompting, and when separation is impossible, compute and report knowledge-base coverage rates and projected distortion bounds (e.g., 10.8% at 70% coverage, scaling to 28.3% at 10% coverage).
For historical document translation, the system can verify that generated names exist in period-appropriate biographical records, flagging names that match no known historical figure (like Chim-yeol
instead of Sim-yeol
) and warning that such errors remove entries from archival search networks.
The system can use paired testing (McNemar's exact test) across overlapping and independent segments to determine whether improvements are statistically significant in both segments or confined to the injected one, rather than relying on correlation coefficients that can mask segment-specific effects.
Sources
- On the Evaluation of Machine Translation for Terminology Consistency
- LLM-Evaluation Tropes: Perspectives on the Validity of LLM-Evaluations
- Insider Knowledge: How Much Can RAG Systems Gain from Evaluation Secrets?
- HERITAGE: An End-to-End Web Platform for Processing Korean Historical Documents in Hanja
- Shared Heritage, Distinct Writing: Rethinking Resource Selection for East Asian Historical Documents
- HUE: Pretrained Model and Dataset for Understanding Hanja Documents of Ancient Korea
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering