Beyond Top Words: MonoTM for Topic Modeling with Interpretable Monosemantic Features

arXiv:2609.09575 · cs.CL, cs.HC · Submitted 2026-09-09 · Read on arXiv

cs.CL, cs.HC

Submitted: 2026-09-09

Updated: 2026-09-09

Comments: Accepted to appear in the Proceedings of AACL-IJCNLP 2026

License: http://creativecommons.org/licenses/by/4.0/

The gist: Topic models summarize large text corpora, but top-ranked words often provide only a limited representation of topic semantics.

Terminology

Abstract

Topic models summarize large text corpora, but top-ranked words often provide only a limited representation of topic semantics. Sparse autoencoders (SAEs) offer a way to move beyond word-level descriptors by extracting interpretable features from dense representations, yet how feature interpretability relates to topic-inference quality remains unclear. We introduce MonoTM, an interpretable topic modeling framework that decouples these roles. Across three benchmark corpora, we show that document--topic mixture estimation and semantic interpretation favor different SAE configurations and feature subsets. MonoTM estimates mixtures from the full SAE bag-of-features representation and, with them fixed, learns topic descriptors over a separate vocabulary of corpus-grounded semantic features. This design preserves global topic structure while representing topics with semantic units more meaningful than individual words, making them more useful for downstream corpus analysis.

Related papers