All Entities are Not Created Equal: Examining the Long Tail for Ultra-Fine Entity Typing
cs.CL
Submitted: 2024-10-22
Updated: 2026-09-11
Journal ref: StarSEM 2025
DOI: 10.18653/v1/2025.starsem-1.15
Code: https://github.com/blast-cu/All-Entities-are-Not-Created-Equal
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: Due to their capacity to acquire world knowledge from large corpora, pre-trained language models (PLMs) are extensively used in ultra-fine entity typing tasks where the space of labels is extremely
Terminology
Abstract
Due to their capacity to acquire world knowledge from large corpora, pre-trained language models (PLMs) are extensively used in ultra-fine entity typing tasks where the space of labels is extremely large. In this work, we explore the limitations of the knowledge acquired by PLMs by proposing a novel heuristic to approximate the pre-training distribution of entities when the pre-training data is unknown. Then, we systematically demonstrate that entity-typing approaches that rely solely on the parametric knowledge of PLMs struggle significantly with entities at the long tail of the pre-training distribution, and that knowledge-infused approaches can account for some of these shortcomings. Our findings suggest that we need to go beyond PLMs to produce solutions that perform well for infrequent entities.
Sources
- Datasheet for the Pile
- The Llama 3 Herd of Models
- The Pile: An 800GB Dataset of Diverse Text for Language Modeling
- Context-Dependent Fine-Grained Entity Type Tagging
- Ultra-Fine Entity Typing with Prior Knowledge about Labels: A Simple Clustering Based Strategy
- Infini-gram: Scaling Unbounded n-gram Language Models to a Trillion Tokens
- Infini-gram mini: Exact n-gram Search at the Internet Scale with FM-Index
- Qwen3 Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering