ReLTEx: Reliable LLM-based Taxonomy Expansion

arXiv:2608.10970 · cs.CL, cs.AI · Submitted 2026-08-11 · Read on arXiv

Zeinab Ghamlouch, Mehwish Alam

Télécom Paris, Institut Polytechnique de Paris

cs.CL, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/zeinabGhamlouch/ReLTEx

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: ReLTEx: Reliable LLM-based Taxonomy Expansion Zeinab Ghamlouch, Mehwish Alam Télécom Paris, Institut Polytechnique de Paris, France Summary This paper presents ReLTEx, a framework for reliable

Terminology

Summary

ReLTEx: Reliable LLM-based Taxonomy Expansion

Zeinab Ghamlouch, Mehwish Alam

Télécom Paris, Institut Polytechnique de Paris, France

Summary

This paper presents ReLTEx, a framework for reliable LLM-based taxonomy expansion. The authors address the challenge that directly relying on LLM-generated expansions often leads to noisy, redundant, or hierarchically inconsistent structures, limiting their reliability for automated taxonomy expansion. The framework combines LLM-driven candidate generation with structure-aware validation and recursive expansion control to improve the consistency and quality of generated taxonomies by reducing hallucinations.

The paper identifies several key problems with existing taxonomy expansion approaches: "A generated concept may be semantically ambiguous, as in 'Bank', which can refer to either a financial institution or the side of a river. Candidate generation may also introduce noise or redundancy, for example by proposing both 'Question' and 'Inquiry'. Moreover, an otherwise valid concept may be attached to an inappropriate parent, such as placing 'Answer' under 'CreativeWork' rather than 'Comment'. These errors can compromise hierarchical coherence and propagate during recursive expansion."

ReLTEx operates in three stages: (A) LLM-based candidate generation, (B) structure-aware validation, and (C) recursive expansion using a stopping mechanism.

Stage A: LLM-based Candidate Generation. "ReLTEx performs recursive taxonomy expansion through a depth-first traversal of the taxonomy. For each target node, the framework constructs a prompt using its local hierarchical context, consisting of the path from the taxonomy root to the target node together with its existing children. The ancestor path provides the semantic context of the current taxonomy branch, while the existing children illustrate the desired level of abstraction, guiding the LLM to generate semantically relevant child concepts at the appropriate level of granularity. The LLM is instructed to generate a fixed number k of candidate child concepts per parent node, avoiding duplicates, synonyms, and simple rephrasing. After generation, Exact lexical duplicates under this canonical representation, as well as candidates matching existing children or sibling nodes, are removed before the remaining candidates are passed to the validation stage."

Stage B: Structure-Aware Candidate Validation. The authors investigate three validation strategies: (i) semantic similarity filtering, (ii) LLM-based validation, and (iii) classifier-guided validation. The proposed classifier learns structural compatibility from annotated taxonomy relations by jointly considering the hierarchical context, the parent concept, and the generated child concept. It is formulated as a binary classification task using a DistilRoBERTa encoder fine-tuned on labeled parent–child relation pairs. The classifier is trained on positive examples (valid taxonomy edges) and negative examples generated using hierarchy-aware perturbations that produce semantically plausible but structurally incorrect relations, including:

  • Reversed edges (e.g., Comment → Answer becomes Answer → Comment)

  • Sibling confusions (e.g., Comment → Answer becomes Book → Answer)

  • Grandparent–child confusions (e.g., Comment → Answer becomes CreativeWork → Answer)

  • Same-depth mismatches (e.g., Article → Answer)

  • Nearby hierarchy confusions (e.g., Review → Answer)

  • Random invalid relations (e.g., Organization → Answer)

The classifier outputs a confidence score s(p, v, ci) ∈ [0, 1], and Candidates with s(p, v, ci) ≥ τ, where τ denotes the validation threshold, are accepted and incorporated into the taxonomy; otherwise, they are discarded.

Stage C: Recursive Expansion. ReLTEx reuses the confidence scores assigned by the classifier during child validation to regulate recursive expansion at the branch level. Only accepted concepts are eligible for further recursive expansion. The framework requires a minimum number of accepted children Av ≥ m to continue expanding a branch. When this condition is met, it computes the mean confidence of accepted children S(v). The expansion below v continues only if S(v) ≥ S(parent(v)) − δ, where δ is the maximum allowable decrease in average classifier confidence between two consecutive expansion levels.

Evaluation Setup. The authors evaluate ReLTEx on two benchmarks: SemEval-2016 Task 13 Environment taxonomy and Schema.org, a large-scale real-world taxonomy. Experiments use four open-source LLMs: Llama3.2:3B, Mistral:7B, Qwen3:8B, and DeepSeek-R1:8B. The evaluation uses a masked taxonomy expansion protocol where 20% of the nodes from the seed taxonomy are removed and treated as hidden test concepts. Only leaf nodes are eligible for masking, with constraints to preserve structural context and balance the distribution of hidden concepts.

Evaluation Metrics. The paper adapts classical metrics to the generative setting:

  • Recall@K (R@K): measures the proportion of hidden concepts that are successfully recovered under their original parent nodes

  • SoftRecall@K (SR@K): accounts for semantic equivalence using cosine similarity with a threshold of ρ = 0.85

  • Mean Reciprocal Rank (MRR): reported as MRRoverall and MRRfound

  • Wu & Palmer (WuP) Similarity: evaluates the structural quality of node generation by comparing predicted parent placements against the original taxonomy structure

Results. In the masked taxonomy expansion benchmark, "Mistral achieves the strongest overall performance across both datasets, obtaining the highest R@K, MRRo, and WuPo scores on Schema.org while remaining highly competitive on SemEval. Llama3.2 performs best on the smaller SemEval taxonomy, achieving the highest R@K, S@K, and WuPo, indicating strong concept recovery and structural placement. Qwen3 consistently produces competitive semantic recovery performance, achieving the highest S@K on Schema.org, suggesting a greater ability to generate semantically related concepts beyond exact lexical matches. DeepSeek-R1 generally obtains lower recovery scores and exhibits larger standard deviations across several metrics, indicating lower stability across different masking configurations."

Human Evaluation. Three annotators evaluated generated expansions using criteria including Hierarchical Correctness (HC), Granularity Consistency (GC), Non-Redundancy (NR), Exact Recovery (ER), and Semantic Recovery (SR). The inter-annotator agreement was measured using Fleiss' κ, which is 0.695. Results show the generated expansions receive consistently high human judgments for hierarchical correctness, granularity consistency, and non-redundancy across both datasets.

LLM-Based Evaluation. Using the LITE framework with three open-source evaluator LLMs, the authors found that "Across both taxonomies, the enriched taxonomies achieve high SCA, indicating that recursive expansion preserves semantic quality. HRR is generally maintained or improved after enrichment, suggesting that generated parent–child relations remain logically coherent. Similar trends are observed for HRE/HRI, with most evaluators assigning equal or higher scores to the enriched taxonomies."

Ablation Study. The ablation study on the SemEval Environment taxonomy using Mistral:7B found that the Local prompting configuration consistently achieves the highest R@K and SR@K compared to Parent Subtree and Full Taxonomy contexts. For validation strategies, the proposed classifier consistently achieves the highest R@K and SR@K compared to semantic validators and an LLM judge.

Error Analysis. The authors identified remaining error categories: incorrect hierarchical placement (e.g., BankTransfer under PaymentCard), scope mismatch (e.g., Gallery under MediaGallery), semantic redundancy (e.g., TattooStudio alongside TattooParlor), and malformed concepts (e.g., Sho Emission Trading).

Expanded Taxonomy Statistics. Using Mistral:7B, ReLTEx expanded the SemEval Environment taxonomy from 211 nodes to 4,883 nodes, and Schema.org from 912 nodes to 38,603 nodes, each adding one additional hierarchy level.

Limitations. The authors note: "Our experiments rely on relatively compact open-source language models to enable efficient and reproducible experimentation. Larger and more capable models may further improve the quality of generated taxonomy expansions. In addition, the proposed classifier is a learned approximation of hierarchical validity and may occasionally accept incorrect parent–child relations. Finally, recursive LLM-based taxonomy generation remains an emerging task without a standardized benchmark or directly comparable generation baseline, making comprehensive system-level comparisons difficult."

Main Contributions. The paper proposes: (1) A zero-shot LLM-based taxonomy expansion framework that exploits contextual information from the taxonomy hierarchy to generate candidate concepts; (2) A structure-aware validation mechanism that mitigates hallucinated and structurally inconsistent parent–child relations, improving the reliability of recursive taxonomy expansion; and (3) A comprehensive evaluation protocol based on adapted evaluation metrics, human evaluation, and LLM based assessment.

Improvements for AI systems

Based on this paper, here are specific improvements for AI systems:

1. Structure-Aware Validation via Learned Classifier

  • Improvement: Replace generic LLM self-validation with a fine-tuned DistilRoBERTa binary classifier trained on hierarchy-aware negative perturbations (reversed edges, sibling confusions, grandparent–child confusions, same-depth mismatches, nearby hierarchy confusions, random invalid relations).

  • Capability: The AI system can reject semantically plausible but structurally incorrect parent–child relations (e.g., placing Answer under CreativeWork instead of Comment) before they propagate into recursive expansion, reducing hallucination cascades.

2. Confidence-Gated Recursive Expansion Control

  • Improvement: Implement a branch-level stopping mechanism that uses the classifier's mean confidence scores, requiring a minimum number of accepted children (m) and allowing expansion only if the average confidence does not drop below the parent's average by more than a threshold δ.

  • Capability: The system autonomously decides when to stop deepening a taxonomy branch, preventing uncontrolled growth of low-quality, noisy subtrees while preserving high-confidence hierarchical structures.

3. Context-Aware Candidate Generation with Local Hierarchical Prompting

  • Improvement: Construct prompts using the ancestor path from root to target node plus existing children, rather than full taxonomy or parent-subtree context. This provides semantic branch context and granularity exemplars.

  • Capability: The system generates child concepts at the appropriate abstraction level, reducing scope mismatches (e.g., generating Gallery under MediaGallery instead of a broader context) and improving lexical diversity while avoiding duplicates and synonyms.

4. Multi-Stage Redundancy and Ambiguity Filtering

  • Improvement: Apply canonical lexical deduplication, removal of candidates matching existing children or siblings, and semantic similarity filtering before validation.

  • Capability: The system eliminates redundant concepts (e.g., Question vs. Inquiry) and ambiguous terms (e.g., Bank as financial vs. river) early in the pipeline, improving precision and reducing downstream validation load.

5. Adaptive Evaluation with Soft Matching and Structural Quality Metrics

  • Improvement: Use SoftRecall@K with cosine similarity threshold (ρ=0.85) to account for semantic equivalence, and Wu & Palmer similarity to evaluate structural placement quality, alongside exact-match metrics.

  • Capability: The AI system can measure not just whether a concept was recovered, but whether it was placed at the correct hierarchical depth and under the correct parent, enabling more nuanced self-assessment of taxonomy expansion quality.

6. Hierarchical Negative Sampling for Robust Relation Learning

  • Improvement: Train the validation classifier using a curriculum of increasingly subtle negative examples—from random invalid relations to same-depth mismatches and nearby hierarchy confusions.

  • Capability: The system learns to distinguish fine-grained structural errors, improving robustness against plausible-but-wrong placements that generic semantic similarity methods would miss.

7. Depth-First Traversal with Recursive Expansion Eligibility

  • Improvement: Only accept concepts that pass validation become eligible for further expansion; enforce a minimum accepted-children threshold per node before recursing deeper.

  • Capability: The system avoids exponential blow-up of low-confidence branches and maintains hierarchical coherence across multiple levels of expansion, as demonstrated by expanding Schema.org from 912 to 38,603 nodes while preserving high human-rated hierarchical correctness.

8. Human-Aligned Quality Criteria Integration

  • Improvement: Incorporate the five human evaluation criteria (Hierarchical Correctness, Granularity Consistency, Non-Redundancy, Exact Recovery, Semantic Recovery) as explicit optimization targets during candidate generation and validation.

  • Capability: The system produces expansions that align with human judgment (Fleiss' κ=0.695), ensuring that generated taxonomies are not only technically valid but also practically useful for downstream applications like search, recommendation, and knowledge organization.

Abstract

Recent advances in Large Language Models (LLMs) have demonstrated strong capabilities in generating semantically relevant concepts and relations, making them promising tools for taxonomy enrichment. However, directly relying on LLM-generated expansions often leads to noisy, redundant, or hierarchically inconsistent structures, limiting their reliability for automated taxonomy expansion. In this paper, we present ReLTEx, a framework for reliable LLM-based taxonomy expansion. ReLTEx combines LLM-driven candidate generation with structure-aware validation and recursive expansion control to improve the consistency and quality of generated taxonomies by reducing hallucinations. We evaluate the proposed framework using benchmark taxonomies under a masked taxonomy expansion setting and compare multiple validation strategies. Experimental results, supported by both adapted evaluation metrics and human evaluation, demonstrate that ReLTEx produces more reliable and semantically coherent taxonomy expansions.

Sources

Related papers