TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification

arXiv:2608.11044 · cs.CL · Submitted 2026-08-11 · Read on arXiv

Jian Zhang, Zhuohao Yang, Songlin Lei, Bangli Liu, Ziwei Wang, Xufeng Weng, Gehan Amaratunga, Yu Lin, Hongwei Wang

Zhejiang University · ZJU-UIUC Institute · Shaoxing K3i Technology Co. Ltd · State Key Laboratory of CAD&CG

cs.CL

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: Accepted by IEEE CSCWD 2026

DOI: 10.1109/CSCWD68734.2026.11582679

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification proposes a weakly supervised hierarchical text

Terminology

Summary

TEAMMix: Taxonomy Enrichment Augmentation and Minority-augmented Mixing Strategy for LLM-enhanced Weak-Supervised Hierarchical Text Classification proposes a weakly supervised hierarchical text classification (HTC) framework enhanced by LLM-based data augmentation. The framework addresses challenges in HTC such as complex label hierarchies and class imbalance. The method first enriches the label hierarchy semantically through keyword generation and corpus mining, enhancing the model's understanding of labels. It then guides the LLM to generate pseudo-samples to mitigate the long-tail problem, and employs a Gaussian mixture model for confidence-based resampling to optimize the quality of generated data. Experimental results demonstrate that the proposed method effectively improves the reliability of LLM-generated pseudo-labels and significantly enhances classification performance on fine-grained and imbalanced datasets.

The core challenges addressed include large-scale label hierarchies, semantic distinctions between labels, and aligning labels with textual content. The method uses a taxonomy enrichment augmentation module that generates discriminative keywords for each category via LLM prompts, uses a top-down tree search algorithm to match candidate keywords, and mines corpus terms based on popularity, distinctiveness, and semantic similarity. The minority-augmented mixing module generates pseudo-texts for each path in the taxonomy tree, applies dynamic confidence-based sampling using a Gaussian mixture model for label screening, and relabels ambiguous labels based on prediction confidence.

The main contributions are: (1) an LLM-enhanced strategy requiring only category labels for HTC, (2) a minority-class augmented sampling strategy to screen and refine LLM-generated samples, and (3) experimental validation showing improved quality of pseudo-samples and labels, alleviating fine-grained sample scarcity.

Experiments were conducted on Amazon-531 and DBPedia-298 datasets. The method outperforms baseline zero-shot and weakly-supervised methods on Example-F1, P@3, and MRR metrics. Compared to TELEClass, TEAMMix achieves competitive P@1 but surpasses it substantially on other metrics, with improvements primarily from the minority-class sample mixing enhancement strategy. Sensitivity analysis on hyperparameters α and ϵ validates the operational range of the model. LLM prompt comparison shows TEAMMix slightly underperforms GPT-4 on Amazon-531 but significantly outperforms in time and cost efficiency (1 vs 800, 2 mins vs 400 mins), making it suitable for real-time applications.

Improvements for AI systems

Improvements to AI Systems:

  1. Hierarchical Label-Aware Data Augmentation
  • Integrate a taxonomy-enrichment module that auto-generates discriminative keywords per node via LLM prompts, then uses a top-down tree search to align corpus terms with label semantics. This improves the system’s ability to distinguish fine-grained, semantically similar categories in large label hierarchies.
  1. Confidence-Based Pseudo-Label Refinement
  • Implement a Gaussian mixture model (GMM) over LLM-generated pseudo-samples to dynamically resample based on prediction confidence. This filters out low-quality or mislabeled samples, reducing noise and improving the reliability of weak supervision without human annotation.
  1. Minority-Class Mixing for Imbalance Handling
  • Add a minority-augmented mixing strategy that synthesizes pseudo-texts for underrepresented paths in the taxonomy tree. This balances class distributions and mitigates the long-tail problem, leading to better performance on rare and fine-grained categories.
  1. Cost-Efficient LLM Prompting for Real-Time Weak Supervision
  • Replace expensive LLM calls (e.g., GPT-4) with a cheaper, faster model (e.g., GPT-3.5) while maintaining competitive accuracy via the GMM-based screening. This reduces cost from 800 to 1 and runtime from 400 mins to 2 mins, enabling real-time deployment in dynamic text classification pipelines.

What the Improved AI System Can Do:

  • Automatically classify hierarchical text (e.g., product categories, news topics, biomedical ontologies) using only label names—no labeled training data required.

  • Handle thousands of fine-grained categories with high semantic overlap, correctly distinguishing subtle differences (e.g., “wireless headphones” vs. “Bluetooth earbuds”).

  • Self-correct its own training data by scoring and resampling LLM-generated examples, ensuring robustness even when the LLM produces hallucinated or off-topic text.

  • Operate in low-resource or rapidly changing domains (e.g., e-commerce taxonomy updates, emerging news events) where manual labeling is infeasible.

  • Deploy in production with minimal latency and cost, making it suitable for real-time content moderation, recommendation tagging, or document routing in large-scale systems.

Related papers