WiC is Not WSD: A Study on LLMs and Lexical Ambiguity Resolution
cs.CL
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: Accepted to AACL 2026 (main)
License: http://creativecommons.org/licenses/by/4.0/
The gist: Word-in-Context (WiC) remains challenging for language models, despite recent progress on lexical-semantic tasks.
Terminology
Abstract
Word-in-Context (WiC) remains challenging for language models, despite recent progress on lexical-semantic tasks. We hypothesise that this difficulty arises not only from comparing two contextual uses of a word, but also from the absence of an explicit sense inventory that specifies the relevant level of semantic granularity. We evaluate open LLMs on WiC and traditional Word Sense Disambiguation (WSD) under similar settings. We find that providing candidate senses, similar to what is done in traditional WSD, improves WiC performance in all settings. In general, explicit sense information helps models make more consistent and targeted judgements. Human evaluation further shows that many apparent WiC errors reflect label ambiguity or mismatches between model and annotator sense boundaries rather than simple failures of lexical understanding. In particular, results show that LLMs overthink the sense distinction often leading to errors based on overly fine-grained distinctions.
Sources
- Exploring the Word Sense Disambiguation Capabilities of Large Language Models
- DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
- The Llama 3 Herd of Models
- Mistral 7B
- Emergent Abilities of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering