From Outliers to Topics in Language Models: Anticipating Trends in News Corpora
cs.CL
Submitted: 2025-09-26
Updated: 2026-09-19
Comments: presented at ICNLSP 2025; to appear in the ACL Anthology; received the Best Full Paper Award
Code: https://github.com/evangeliazve/outliers-to-topics-icnlsp
Project page: https://maartengr.github.io/BERTopic
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: This paper examines how outliers, often dismissed as noise in topic modeling, can act as weak signals of emerging topics in dynamic news corpora.
Terminology
Abstract
This paper examines how outliers, often dismissed as noise in topic modeling, can act as weak signals of emerging topics in dynamic news corpora. Using vector embeddings from state-of-the-art language models and a cumulative clustering approach, we track their evolution over time in French and English news datasets focused on corporate social responsibility and climate change. The results reveal a consistent pattern: outliers tend to evolve into coherent topics over time across both models and languages.
Sources
- BERTrend: Neural Topic Modeling for Emerging Trends Detection
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- MTEB: Massive Text Embedding Benchmark
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering