Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection

arXiv:2604.13232 · cs.CL · Submitted 2026-05-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection".

Jane: The paper was written by Bach Phan-Tat, Kris Heylen, Dirk Geeraerts, Stefano De Pascale and Dirk Speelmana from Institute for Dutch Language, Leiden, The Netherlands and Department of Linguistics and Literary Studies, Vrije Universiteit Brussel, Brussels, Belgium and Department of Linguistics, KU Leuven, Leuven, Belgium.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary of Core Arguments: Tom: So, building on that critique of the title, let’s break down what the paper actually argues in its core findings. They focus on three main areas where they see issues with SemEval-two thousand twenty Task one.

Jane: First, they operationalize semantic change in a very narrow way—they only look at whether a word has gained or lost a discrete sense inventory. But the paper suggests this is too rigid for gradual change.

Meng: And that’s tied to their second major point about data quality, which I found really concerning. They found significant issues like OCR noise and truncated sentences in the corpora used for testing.

Lu: That noise issue completely undermines the notion of "discrete senses" because if the input data is corrupted, how can our models reliably detect subtle linguistic shifts?

Lalam: It’s a huge challenge to get a clean signal when your foundational data, as described in this study, is so messy. We are trying to measure subtle shifts but the raw materials are flawed.

Tom: The authors also pointed out limitations in the design itself, suggesting that these small target sets don't really represent real-world language use.

Jane: That’s right, they are arguing that using a small, curated list of words doesn' not give us a fair picture of how AI systems perform across the entire vocabulary.

Meng: They also mention issues with inconsistent lemmatization and POS tagging errors in the data, which makes it incredibly difficult to trust the resulting gold scores for any model.

Lu: So, we are seeing a convergence of these problems—a narrow theory, noisy data, and then weak experimental design—that is making the benchmark unreliable.

Lalam: We need to recognize that if the evaluation framework itself has these fundamental blind spots, we cannot use it to claim definitive progress.

Suggested Improvements: Tom: Given these serious critiques, what does the paper suggest as a path forward? They aren't just complaining; they' are proposing concrete ways to improve the field.

Jane: The authors really push for moving away from that rigid, discrete sense model towards continuous modeling of semantic change over time.

Meng: But how do we handle the practical complexity and cost of analyzing a continuous trajectory instead of just looking at two time slices?

Lu: We should be looking at broader linguistic theories too, like focusing on collocational or constructional changes rather than just the sense inventory.

Lalam: It's about giving AI models more explanatory targets—not just asking *if* a word changed, but *how* and why it might have evolved.

Tom: The paper also suggests that we must document our preprocessing steps with complete transparency, which is a huge operational shift for the data scientists.

Jane: That means researchers have to release their normalisation scripts so that when someone else runs the test, they know exactly how the original corpus was handled.

Meng: Transparency is key to reproducibility, and if you can’t replicate your results because of hidden preprocessing steps, the performance comparison is meaningless.

Lu: I think this requires a rethinking of what we consider a "valid" input, moving beyond just the raw text to include the full linguistic context.

Lalam: This shift allows us to build models that don' not just recognize a change but understand the mechanism of change itself, which will be transformative for cultural understanding.

Conclusion: Tom: As we wrap up our discussion on "Evaluating the Evaluator," it’s clear this paper is a necessary call for rethinking how we measure linguistic change.

Jane: The goal isn't to abandon benchmarks, but to make them far more comprehensive and statistically robust than what SemEval-two thousand twenty offered.

Meng: We need larger target inventories and a much broader scope of languages, not just the four European ones currently in the task.

Lu: I think we should also be looking at how these new, richer benchmarks will allow AI to detect more complex linguistic patterns that current models simply miss.

Lalam: By demanding greater rigor, this paper forces us toward a more holistic understanding of language evolution, ensuring our future work is truly representative of human history.

Tom: It’s a powerful reminder that the quality of our evaluation dictates the progress we make in AI semantic detection. We hope this has given you listeners a good overview of what the authors are suggesting for better computational linguistics.

Conclusion: Tom: So, we’ve spent a lot of time looking at the core weaknesses in SemEval-two thousand twenty Task one but it is absolutely critical to remember that this entire discussion about "Evaluating the Evaluator" isn't just an academic exercise.

Jane: It's a call for pushing boundaries, Tom, because we’ve seen how much these small target sets and noisy data can limit our ability to truly measure progress in the world of lexical semantic change detection.

Lu: I think what’s most exciting is that this failure to a paradigm shift—the realization that the old way wasn't working—is unlocking so many creative possibilities for how we model language evolution now.

Meng: For me, it means that as engineers, we are finally being given clear instructions on what to prioritize: building systems with robust pre-processing and a much greater attention to data quality rather than just relying on the flawed gold scores.

Lalam: We have learned that if we want our AI to understand language culturally, we must first understand its history accurately, and acknowledging these errors is how we begin building a more nuanced understanding of human thought.

Tom: It’s a humbling thing to see how much the limitations of the evaluation method can dictate what kind of models get built in the field.

Jane: We're hopeful that this critique will pave the way for future datasets that allow us to track those gradual, continuous shifts in meaning we've been missing.

Lu: I really love thinking about how much more complex language can be when we move beyond just semantic change and look at things like grammatical profiling.

Meng: It makes sense that if we’ are aiming for real-world applicability, the system needs to be able to handle the messy reality of historical data rather than idealized inputs.

Lalam: We want our future AI to reflect the full richness of human history, and this study is how we start making sure it doesn' that.

Tom: Well, we hope these insights into "Evaluating the Evaluator" give you a clear picture of what’s possible for the next time we talk about this topic.

Bach Phan-Tat, Kris Heylen, Dirk Geeraerts, Stefano De Pascale, Dirk Speelmana

Institute for Dutch Language, Leiden, The Netherlands · Department of Linguistics and Literary Studies, Vrije Universiteit Brussel, Brussels, Belgium · Department of Linguistics, KU Leuven, Leuven, Belgium

cs.CL

Submitted: 2026-05-27

Updated: 2026-08-25

Code: https://github.com/phantatbach/LChange26-Dep

Importance score: 77/100

The gist: The paper re-examines SemEval-2020 Task 1, a highly influential shared benchmark for lexical semantic change detection, through a three-part evaluative framework: operationalisation, data quality,

Key concepts

Lexical Semantic Change Detection
This field studies how the meaning of words evolves over time. The paper examines current methods for identifying these shifts, noting that existing benchmarks often fail to capture subtle or gradual linguistic changes accurately.
SemEval-2020 Task 1 Issues
The authors critique this benchmark for being flawed. Problems include operationalizing change too narrowly (only discrete senses), using small target sets, and suffering from data quality issues like OCR noise and inconsistent tagging.
Continuous Modeling
This suggests moving away from looking at a word having one or two specific meanings (discrete senses). Instead, it involves modeling how the meaning changes gradually over time, allowing for a more nuanced understanding of linguistic evolution.
Data Transparency
Researchers must fully document their preprocessing steps and release normalization scripts. This ensures that other scientists can replicate results, which is crucial since hidden steps make performance comparisons meaningless.

Terminology

Summary

The paper re-examines SemEval-2020 Task 1, a highly influential shared benchmark for lexical semantic change detection, through a three-part evaluative framework: operationalisation, data quality, and benchmark design.

Operationalisation Limitations

At the level of operationalization, the benchmark models semantic change mainly as gain, loss, or redistribution of discrete senses. The authors argue that this framing is too narrow to capture gradual, constructional, collocational, and discourse-level change. Furthermore, they note that the gold labels are not objective truths but are outcomes of annotation decisions, clustering procedures, and threshold settings, which could potentially limit the validity of the task.

Data Quality Issues

The benchmark is significantly affected by substantial corpus and preprocessing problems. These issues include:

  • Corpus Noise: The data contains OCR noise (Optical Character Recognition errors) and malformed characters.

  • Textual Artifact Errors: There are instances of truncated sentences, inconsistent lemmatization, POS-tagging errors, and missed targets.

  • The authors quantify these issues in Table 1, showing that abnormal sentence boundaries (cut-off and oddly merged sentences) are pervasive across ENG 1 and 2 (CCOHA), GER 2 (BZ + ND) AND SWE 1 (Kubhist).

These issues collectively can distort model behaviour, complicate linguistic analysis, and reduce reproducibility.

Benchmark Design Limitations

The third limitation concerns the design of the test itself. The small curated target sets—37 targets for English, 48 for German, 40 for Latin, and 31 for Swedish—are noted to reduce realism and increase statistical uncertainty. This is compounded by the limited language coverage (only four European languages).

Conclusion and Call to Action

Taken together, these limitations suggest that the benchmark should be treated as a useful but partial test bed rather than a definitive measure of progress. The authors therefore call for future datasets and shared tasks to adopt broader theories of semantic change, document preprocessing transparently, expand cross-linguistic coverage, and use more realistic evaluation settings. Such steps are deemed necessary for more valid, interpretable, and generalisable progress in lexical semantic change detection.

Improvements for AI systems

Based on a meticulous analysis of the limitations and critiques presented in Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection, I have developed four specific, actionable improvements for AI systems designed to detect semantic change.

These improvements move beyond simply classifying if a word has changed to diagnosing how and why, while ensuring the robustness of the underlying data processing.


The original system is only rewarded for matching a binary (Changed/Unchanged) or gradient (Divergence Score) label. The improved system must be designed to provide linguistic justification for the detected change.

Specific Improvements:

  • Feature Integration: The model will not rely solely on lemma-level co-occurrence statistics. It will incorporate a multi-dimensional loss function that explicitly penalizes and rewards the shift in:
  1. Collocational Patterns: Detecting changes in adjacent words (e.g., grasp shifting from physical objects to abstract concepts).

  2. Syntactic/POS Environments: Identifying shifts in how a word functions grammatically (e.g., a noun becoming an adjective).

  3. Discourse Domains/Registers: Recognizing the contextual shift in which the word is typically used (e.g., board moving from a table to a governing body).

  • Change Type Classification: The system will be trained to distinguish between specific mechanisms of change (e.g., metonymic shift, gradual extension, pragmatic strengthening) rather than just reporting that a sense inventory has changed.

What the Improved AI System Can Do:

The system can generate an Explanatory Semantic Change Report, stating not only that the word X has changed, but also providing the specific linguistic evidence (e.g., X's shift is driven by its increasing association with legal contexts and its use of verbs indicating abstract reasoning).

The reliance on raw, noisy historical corpora is a critical failure point. The improved system must treat preprocessing as an integral part of the model's reliability, not merely a one-time data cleaning step.

The current pairwise comparison (T 1 vs T 2) simplifies a complex, continuous process into an endpoint comparison. The improved system must model semantic change as a trajectory.

Small target sets lead to high statistical uncertainty, making comparisons between models unreliable. The evaluation framework must account for this fragility.

Sources

Related papers