Iterative Topic Taxonomy Induction with LLMs: A Case Study of Electoral Advertising

arXiv:2510.15125 · cs.CL, cs.AI, cs.CY, cs.LG, cs.SI · Submitted 2025-10-16 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Iterative Topic Taxonomy Induction with LLMs".

Jane: By combining unsupervised clustering with iterative prompt-based inference from large language models, this research introduces an end-to-end framework for automatically inducing interpretable topic taxonomies from unlabeled political advertising text corpora.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, we've really dug deep into this paper about "Iterative Topic Taxonomy Induction with LLMs: A Case Study of Electoral Advertising," and now we’re wrapping up our discussion on what this means for us.

Jane: To put it simply, this research introduces a method that uses large language models to automatically discover organized topic labels from huge amounts of unlabeled text without needing any pre-made starting ideas. The authors focused heavily on showing that their approach consistently produced topic labels that human experts agreed with when compared against existing methods in political advertising.

Lu: The researchers really highlighted how they managed the iterative process, essentially letting the AI refine its understanding of the topics through constant questioning and self-correction until it achieved a coherent structure. That ability to build something complex iteratively is really something worth thinking about for modeling other kinds of unstructured data.

Meng: From my side, I’m still focused on how we can optimize that iterative inference speed so this framework isn't just useful for looking back at old data, but can actually handle the kind of continuous data streams we see in real-time. We need to make sure the engineering keeps up with the conceptual elegance.

Lalam: For me, this work really emphasizes how AI can systematically uncover hidden societal structures embedded within massive communication without relying on human bias to tell us what those structures even look like. That's a fundamental shift in how we think about understanding culture itself.

Tom: Exactly, Lalam, and Jane; it’s not just about organizing text anymore; it’s about gaining a new lens to look at political discourse by directly linking the content to human values and audience targeting patterns. It shows us where the money is going in terms of political messaging.

Jane: It really demonstrates the power of using iterative AI processes to create interpretable structures from messy, real-world data like political advertising, which is a major step forward for making complex information accessible.

Lu: And the implication is that this general framework could be applied across countless other large, unlabeled text sets where manual labeling would just be impossible or way too slow. That scalability is where the real research potential lies in AI theory.

Meng: I’m still thinking about how we can make that iterative inference process faster and more efficient so it’s ready for real-time analysis of data streams, which is a huge practical hurdle we need to clear.

Lalam: That speed, combined with the ability to uncover these deep structural connections, is what really excites me about where this technology can go next in shaping our understanding of society. We’re looking at a future where AI helps us map the hidden logic of communication itself.

Conclusion: Tom: So, we’ve just wrapped up our deep dive into "Iterative Topic Taxonomy Induction with LLMs: A Case Study of Electoral Advertising," and now we need to talk about what this paper is actually called and who came up with it.

Jane: The title itself tells us exactly what the authors were aiming for: they wanted to show how you can use iterative topic induction powered by Large Language Models to build organized topic labels from text, using political advertising as their test case.

Lu: I found the paper's methodology really clever; they used this combination of embedding-based clustering and prompt-based inference in a way that lets the AI refine its understanding of the topics through constant self-correction. It’s like giving the AI a conversation with itself to build structure.

Meng: From an engineering viewpoint, it’s impressive that they managed to bake that entire iterative loop—the clustering and the LLM prompting—into one coherent pipeline for handling huge datasets without needing constant human intervention in the middle of it.

Lalam: What strikes me most is how this work moves beyond simple text summarization; it’s about building a system that can systematically uncover societal structures embedded in communication without relying on human bias to tell us what those structures even look like.

Tom: Exactly, Lalam! Jane, you mentioned the core claim earlier—what’s the simplest way to explain what this paper actually *achieved* for our listeners? Keep it very accessible.

Jane: Well, essentially, this research provides a scalable method that lets AI automatically create clear topic labels from massive amounts of unlabeled text, and they proved these induced labels are actually quite consistent when compared to human experts looking at political ads.

Lu: That consistency is what makes the iterative process so powerful; it ensures the final structure isn't just a random grouping but one that has been vetted through a logical, self-correcting cycle. We could apply this structure induction idea to almost any domain where we have tons of text but no pre-defined categories.

Meng: And for practical application, this means researchers can build tools that don't need massive teams of people just to start labeling data; they can rapidly prototype analyses based on these discovered themes. It’s about accelerating the pace of discovery in many fields.

Lalam: The real cultural impact here is showing us how AI can reveal hidden patterns and value systems within advertising, helping us build a more nuanced understanding of society through its communication, rather than just seeing surface-level messaging.

Tom: So we’ve seen the technical setup and the impressive validation numbers; Jane, what's your final thought on why this paper matters right now in the context of our digital world?

Jane: This framework shows a practical way to take complex political discourse, break it down into consistent themes—like which moral foundations are being used—and then use that structure to understand targeting and spending patterns, making the analysis much richer.

Lu: The potential for scaling this iterative induction across completely different types of unstructured data sets is where the heavy theoretical lift is; that’s where the next big AI research questions live.

Meng: On a practical note, we still need to nail down how fast we can run those iterative loops in a way that handles continuous data streams efficiently, because retrospective analysis isn't enough for real-time decision making.

Lalam: I think the most important thing is demonstrating that AI can systematically uncover these underlying structures without needing prior knowledge of what those structures look like, which opens up new ways we can build better organizational intelligence into the digital landscape.

Department of Computer Science, ETH Zürich · Department of Computer Science, Purdue University

cs.CL, cs.AI, cs.CY, cs.LG, cs.SI

Submitted: 2025-10-16

Updated: 2026-10-01

Comments: Accepted to AACL-IJCNLP 2026 Findings. Camera-ready

Code: https://github.com/alexander-brady/llm-topic-synthesis

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 76/100

The gist: By combining unsupervised clustering with iterative prompt-based inference from large language models, this research introduces an end-to-end framework for automatically inducing interpretable topic

Key concepts

Embedding-based Clustering
This initial step uses a pre-trained model like Sentence-BERT to turn text documents into numerical vectors (embeddings). These vectors are then processed through UMAP for dimensionality reduction and HDBSCAN to group similar text together, forming preliminary topical structures from the raw data.
Iterative Topic Synthesis
The core innovation involves using an LLM in a loop. If a cluster lacks a label, the LLM creates one. For existing clusters, it is asked if labels are accurate ('yes'/'no'). If inaccurate, the LLM generates a new topic until all clusters have coherent and consistent labels.
SetFit
This technique is used for downstream tasks like classifying documents based on the newly generated topics. It involves fine-tuning a sentence transformer model on labeled data and then training a classification head to map document embeddings directly to the specific topic labels efficiently.

Terminology

Summary

By combining unsupervised clustering with iterative prompt-based inference from large language models, this research introduces an end-to-end framework for automatically inducing interpretable topic taxonomies from unlabeled political advertising text corpora. This method is significant because it enables the discovery of semantically rich and coherent topic labels without requiring predefined seed sets or domain expertise, providing a scalable tool for analyzing large-scale political discourse.

How it works

The proposed framework consists of three key components: (1) embedding-based clustering, (2) large language model (LLM) topic synthesis, and (3) LLM-based annotation. The process begins with unsupervised clustering where input documents are embedded using a pre-trained sentence embedding model like Sentence-BERT, followed by dimensionality reduction via UMAP to manage the curse of dimensionality. Subsequently, HDBSCAN is employed to group similar text data points into topical structures, extracting cluster representatives.

Iterative Topic Synthesis

The core innovation lies in the iterative process where LLMs are used to construct the taxonomy without initial labels. The process follows Algorithm 1: if a cluster has no existing label, the LLM generates one; for every other cluster, the LLM is prompted whether existing labels are accurate (yes or "no) using constrained decoding. If a label is deemed inaccurate (no"), the LLM is prompted to generate a new topic. This iterative loop continues until all clusters are processed, resulting in a set of coherent and interpretable document clusters.

Cluster Representative Labeling and Classification

After taxonomy generation, the next step involves annotating cluster representatives with the newly generated labels using constrained decoding to limit LLM output to a single label. Furthermore, the framework supports downstream supervised classification tasks. SetFit is identified as particularly effective for this purpose, utilizing a two-step process: fine-tuning a pre-trained sentence transformer model on labeled data points and then training a classification head to map embeddings to the label space for efficient annotation of remaining documents.

Case Study Validation

The framework was validated using a comprehensive case study of Meta political advertisements from one month before the 2024 U.S. presidential election, involving 8047 unique ads. The results demonstrated that the proposed method yielded more consistent and better-aligned topic labels than both BERTopic and the singleshot TopicGPT-style baseline under this evaluation setup, achieving a Cohen’s Kappa of 0.66 between annotators. Analysis revealed strong correlations between moral foundations (e.g., Fairness/Cheating with crime/justice) and induced topics, as well as clear demographic targeting patterns detected via positive PPMI analysis across states and age groups. Finally, ad-level analysis showed that topics like abortion, personal freedom, and voting rights received the highest average spend, while moral framing was highly polarized in abortion advertising.

Conclusion

The study successfully demonstrates a general, seed-free framework for iterative topic taxonomy induction using LLMs that enforces global consistency across clusters. This method enables interpretable downstream analyses of political advertising, including issue prevalence and moral framing, and its applicability extends beyond this domain to other large, unlabeled text corpora. While limitations exist regarding potential LLM biases and domain specificity, the framework offers a practical solution for scalable analysis in the era of digital microtargeting.


The gist

Structured, iterative labeling yields more consistent and interpretable topic labels than existing approaches under human evaluation for analyzing political advertising data.

How it works

  1. Embedding-based clustering: Documents are embedded using Sentence-BERT and reduced via UMAP, followed by HDBSCAN to identify topical structures.

  2. Iterative Topic Synthesis: An LLM iteratively generates a topic taxonomy from scratch by prompting it to generate new labels when existing ones are deemed inaccurate, constrained to output yes or no.

  3. Cluster Representative Labeling: The generated taxonomy is used to assign a single topic label to each cluster representative via constrained decoding.

  4. Supervised Classification: SetFit is employed for downstream classification tasks using the labeled cluster representatives as training data.

Case Study Validation

The framework was tested on 8047 political ads from the 2024 U.S. election, resulting in 72 clusters and a taxonomy of 14 topics. The method outperformed BERTopic and a singleshot TopicGPT baseline in topic alignment, achieving a Cohen’s Kappa of 0.66 between human annotators. Analysis showed strong correlations between moral foundations (e.g., Fairness/Cheating with crime/justice) and induced topics, and revealed demographic targeting strategies based on state-level impressions for issues like affordable housing among younger Florida audiences.

Ad-Level Analysis

Individual ads were annotated using the synthesized taxonomy to enable granular analysis of spending and funder behavior.

Improvements for AI systems

Here are specific improvements to AI systems based on the proposed framework, and what those improved systems can achieve:


  1. Improvement: Development of a general, seed-free, LLM-guided framework for interpretable topic taxonomy induction (Section 6).

  2. Improvement: Implementation of an iterative topic synthesis mechanism that yields more consistent and better-aligned labels than classical topic models and singleshot LLM labeling under human evaluation (Section 1 & 5.1).

  3. Improvement: Integration of embedding-based clustering (UMAP + HDBSCAN) with LLM inference for unsupervised, semantically rich topic generation, replacing traditional methods like LDA or NMF (Section 3.1 & 3.2).

  4. Improvement: Creation of a multi-pass framework that forces LLMs to iteratively validate and generate new topics based on existing ones, rather than single-shot label assignment (Algorithm 1).

  5. Improvement: Incorporation of Moral Foundation Theory (MFT) analysis by prompting LLMs to classify cluster representatives against predefined moral foundations, providing semantic grounding for political discourse (Section 4.3 & 5.2).

  6. Improvement: Implementation of a downstream supervised classification task using advanced models like SetFit to annotate large volumes of unlabeled data based on the induced taxonomy (Section 3.4).

  7. Improvement: Development of a microtargeting analysis module utilizing Positive Pointwise Mutual Information (PPMI) between ad content and demographic attributes (age, location, gender) to uncover specific targeting strategies (Section 5.2).

These improved AI systems can perform the following specific tasks:

  1. Predict the dominant political issues and moral framing structures within vast, unlabeled text corpora (e.g., social media advertising) without requiring pre-existing domain knowledge or manual labeling efforts.

  2. Generate a coherent, interpretable taxonomy of topics that is validated by human experts to be superior in consistency and alignment compared to traditional statistical topic models or single-shot LLM outputs.

  3. Systematically identify subtle, context-specific thematic shifts and arguments within dynamic political content through the iterative refinement process (LLM-in-the-loop).

  4. Perform granular ad-level analysis, linking specific topics and moral foundations to individual advertisements to understand spending patterns, funder concentration dynamics across issue domains, and precise demographic targeting strategies.

  5. Uncover the correlation between political issues (e.g., crime/justice) and underlying moral appeals (e.g., loyalty/betrayal) at the cluster level, revealing how different campaigns strategically leverage specific moral narratives to appeal to different voter segments in specific geographic regions.

  6. Automate the classification of new, unseen text into these discovered categories using a robust, fine-tuned model trained on the iteratively generated taxonomy.

Abstract

Social media platforms play a pivotal role in shaping political discourse, but the scale and rapid evolution of online content make systematic analysis difficult. We introduce an end-to-end framework for inducing an interpretable topic taxonomy from unlabeled text corpora. The framework combines embedding-based clustering with iterative large language model (LLM) inference to construct a topic taxonomy without requiring predefined labels or seed topics. It first synthesizes candidate topics from document clusters and then uses the resulting taxonomy to assign consistent topic labels across clusters. We evaluate the approach through a case study of political advertising ahead of the 2024 U.S. presidential election. We use the induced taxonomy to support downstream analyses of issue prevalence, moral framing, advertising spend, and demographic exposure patterns. These results suggest that iterative taxonomy construction can provide a scalable and interpretable approach to organizing large unlabeled text corpora while supporting substantive downstream analysis.

Sources

Related papers