Light or Full Verb? A Minimal-Pair Dataset for Probing Phraseological Competence in Language Models
cs.CL
Submitted: 2026-06-03
Updated: 2026-09-09
Code: https://github.com/XplainLing/LVC_
License: http://creativecommons.org/licenses/by/4.0/
The gist: Frequent verbs such as 'have' and 'make' can function either as collocates in light-verb constructions or as full lexical predicates, as in 'make a decision' vs.
Terminology
Abstract
Frequent verbs such as 'have' and 'make' can function either as collocates in light-verb constructions or as full lexical predicates, as in 'make a decision' vs. 'make a cake'. Whether language models represent this distinction, and whether such representations vary across languages, remains unclear. We introduce a large-scale controlled dataset in English, Spanish, and French, comprising minimally varying sentence series in which the same context contains the same verb in light-verb and full-verb uses. Two probing experiments show that language models differentiate between these uses even in minimal contexts and exhibit separable patterns across object types. We release the dataset, generation code, and materials as a reusable resource. The framework supports extensions to broader contexts, additional verbs, and other languages.
Sources
- Gemma 3 Technical Report
- Tracing the complexity profiles of different linguistic phenomena through the intrinsic dimension of LLM representations
- MultiBLiMP 1.0: A Massively Multilingual Benchmark of Linguistic Minimal Pairs
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering