Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection

summary

Video file (mp4)

The gist

The paper re-examines SemEval-2020 Task 1, a highly influential shared benchmark for lexical semantic change detection, through a three-part evaluative framework: operationalisation, data quality,

In short

This episode reviews a paper titled "Evaluating the Evaluator," which critically examines SemEval-2020 Task 1 for Lexical Semantic Change Detection. Hosts discuss major flaws, including narrow definitions of change, noisy input data (OCR errors), and limited test sets. The discussion concludes with a call for future AI research to adopt continuous modeling and greater data transparency to accurately track language evolution.

Key concepts

Lexical Semantic Change Detection
This field studies how the meaning of words evolves over time. The paper examines current methods for identifying these shifts, noting that existing benchmarks often fail to capture subtle or gradual linguistic changes accurately.
SemEval-2020 Task 1 Issues
The authors critique this benchmark for being flawed. Problems include operationalizing change too narrowly (only discrete senses), using small target sets, and suffering from data quality issues like OCR noise and inconsistent tagging.
Continuous Modeling
This suggests moving away from looking at a word having one or two specific meanings (discrete senses). Instead, it involves modeling how the meaning changes gradually over time, allowing for a more nuanced understanding of linguistic evolution.
Data Transparency
Researchers must fully document their preprocessing steps and release normalization scripts. This ensures that other scientists can replicate results, which is crucial since hidden steps make performance comparisons meaningless.

Terminology used across episodes

This episode discusses

The paper

Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection · Read on arXiv

Bach Phan-Tat, Kris Heylen, Dirk Geeraerts, Stefano De Pascale, Dirk Speelmana

Institute for Dutch Language, Leiden, The Netherlands · Department of Linguistics and Literary Studies, Vrije Universiteit Brussel, Brussels, Belgium · Department of Linguistics, KU Leuven, Leuven, Belgium

This discussion paper re-examines SemEval-2020 Task 1, the most influential shared benchmark for lexical semantic change detection, through a three-part evaluative framework: operationalisation, data quality, and benchmark design. First, at the level of operationalisation, we argue that the benchmark models semantic change mainly as gain, loss, or redistribution of discrete senses. While practical for annotation and evaluation, this framing is too narrow to capture gradual, constructional, collocational, and discourse-level change. Also, the gold labels are outcomes of annotation decisions, clustering procedures, and threshold settings, which could potentially limit the validity of the task. Second, at the level of data quality, we show that the benchmark is affected by substantial corpus and preprocessing problems, including OCR noise, malformed characters, truncated sentences, inconsistent lemmatisation, POS-tagging errors, and missed targets. These issues can distort model behaviour, complicate linguistic analysis, and reduce reproducibility. Third, at the level of bench-mark design, we argue the small curated target sets and limited language coverage reduce realism and increase statistical uncertainty. Taken together, these limitations suggest that the benchmark should be treated as a useful but partial test bed rather than a definitive measure of progress. We therefore call for future datasets and shared tasks to adopt broader theories of semantic change, document pre-processing transparently, expand cross-linguistic coverage, and use more realistic evaluation settings. Such steps are necessary for more valid, interpretable, and generalisable progress in lexical semantic change detection

DOI: 10.5334/johd.547

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Evaluating the Evaluator: Problems with SemEval-2020 Task 1 for Lexical Semantic Change Detection".

Jane: The paper was written by Bach Phan-Tat, Kris Heylen, Dirk Geeraerts, Stefano De Pascale and Dirk Speelmana from Institute for Dutch Language, Leiden, The Netherlands and Department of Linguistics and Literary Studies, Vrije Universiteit Brussel, Brussels, Belgium and Department of Linguistics, KU Leuven, Leuven, Belgium.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary of Core Arguments: Tom: So, building on that critique of the title, let’s break down what the paper actually argues in its core findings. They focus on three main areas where they see issues with SemEval-two thousand twenty Task one.

Jane: First, they operationalize semantic change in a very narrow way—they only look at whether a word has gained or lost a discrete sense inventory. But the paper suggests this is too rigid for gradual change.

Meng: And that’s tied to their second major point about data quality, which I found really concerning. They found significant issues like OCR noise and truncated sentences in the corpora used for testing.

Lu: That noise issue completely undermines the notion of "discrete senses" because if the input data is corrupted, how can our models reliably detect subtle linguistic shifts?

Lalam: It’s a huge challenge to get a clean signal when your foundational data, as described in this study, is so messy. We are trying to measure subtle shifts but the raw materials are flawed.

Tom: The authors also pointed out limitations in the design itself, suggesting that these small target sets don't really represent real-world language use.

Jane: That’s right, they are arguing that using a small, curated list of words doesn' not give us a fair picture of how AI systems perform across the entire vocabulary.

Meng: They also mention issues with inconsistent lemmatization and POS tagging errors in the data, which makes it incredibly difficult to trust the resulting gold scores for any model.

Lu: So, we are seeing a convergence of these problems—a narrow theory, noisy data, and then weak experimental design—that is making the benchmark unreliable.

Lalam: We need to recognize that if the evaluation framework itself has these fundamental blind spots, we cannot use it to claim definitive progress.

Suggested Improvements: Tom: Given these serious critiques, what does the paper suggest as a path forward? They aren't just complaining; they' are proposing concrete ways to improve the field.

Jane: The authors really push for moving away from that rigid, discrete sense model towards continuous modeling of semantic change over time.

Meng: But how do we handle the practical complexity and cost of analyzing a continuous trajectory instead of just looking at two time slices?

Lu: We should be looking at broader linguistic theories too, like focusing on collocational or constructional changes rather than just the sense inventory.

Lalam: It's about giving AI models more explanatory targets—not just asking *if* a word changed, but *how* and why it might have evolved.

Tom: The paper also suggests that we must document our preprocessing steps with complete transparency, which is a huge operational shift for the data scientists.

Jane: That means researchers have to release their normalisation scripts so that when someone else runs the test, they know exactly how the original corpus was handled.

Meng: Transparency is key to reproducibility, and if you can’t replicate your results because of hidden preprocessing steps, the performance comparison is meaningless.

Lu: I think this requires a rethinking of what we consider a "valid" input, moving beyond just the raw text to include the full linguistic context.

Lalam: This shift allows us to build models that don' not just recognize a change but understand the mechanism of change itself, which will be transformative for cultural understanding.

Conclusion: Tom: As we wrap up our discussion on "Evaluating the Evaluator," it’s clear this paper is a necessary call for rethinking how we measure linguistic change.

Jane: The goal isn't to abandon benchmarks, but to make them far more comprehensive and statistically robust than what SemEval-two thousand twenty offered.

Meng: We need larger target inventories and a much broader scope of languages, not just the four European ones currently in the task.

Lu: I think we should also be looking at how these new, richer benchmarks will allow AI to detect more complex linguistic patterns that current models simply miss.

Lalam: By demanding greater rigor, this paper forces us toward a more holistic understanding of language evolution, ensuring our future work is truly representative of human history.

Tom: It’s a powerful reminder that the quality of our evaluation dictates the progress we make in AI semantic detection. We hope this has given you listeners a good overview of what the authors are suggesting for better computational linguistics.

Conclusion: Tom: So, we’ve spent a lot of time looking at the core weaknesses in SemEval-two thousand twenty Task one but it is absolutely critical to remember that this entire discussion about "Evaluating the Evaluator" isn't just an academic exercise.

Jane: It's a call for pushing boundaries, Tom, because we’ve seen how much these small target sets and noisy data can limit our ability to truly measure progress in the world of lexical semantic change detection.

Lu: I think what’s most exciting is that this failure to a paradigm shift—the realization that the old way wasn't working—is unlocking so many creative possibilities for how we model language evolution now.

Meng: For me, it means that as engineers, we are finally being given clear instructions on what to prioritize: building systems with robust pre-processing and a much greater attention to data quality rather than just relying on the flawed gold scores.

Lalam: We have learned that if we want our AI to understand language culturally, we must first understand its history accurately, and acknowledging these errors is how we begin building a more nuanced understanding of human thought.

Tom: It’s a humbling thing to see how much the limitations of the evaluation method can dictate what kind of models get built in the field.

Jane: We're hopeful that this critique will pave the way for future datasets that allow us to track those gradual, continuous shifts in meaning we've been missing.

Lu: I really love thinking about how much more complex language can be when we move beyond just semantic change and look at things like grammatical profiling.

Meng: It makes sense that if we’ are aiming for real-world applicability, the system needs to be able to handle the messy reality of historical data rather than idealized inputs.

Lalam: We want our future AI to reflect the full richness of human history, and this study is how we start making sure it doesn' that.

Tom: Well, we hope these insights into "Evaluating the Evaluator" give you a clear picture of what’s possible for the next time we talk about this topic.

More episodes

← Home