Interpretable Unsupervised Community Detection with LLM-Symbolized Structured Processes

arXiv:2608.06402 · cs.AI, cs.LG · Submitted 2026-08-02 · Read on arXiv

Aoting Zeng, Kai Wang, Jianwei Wang, Yuxiang Sun, Yizhang He, Wenjie Zhang

Shanghai Jiao Tong University · University of New South Wales

cs.AI, cs.LG

Submitted: 2026-08-02

Updated: 2026-08-10

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 37/100

The gist: LUCID is proposed as "an LLM-guided interpretable, training-free and Unsupervised CommunIty Detection method" designed to address the limitations of classic objective-driven methods that "struggle

Terminology

Summary

LUCID is proposed as an LLM-guided interpretable, training-free and Unsupervised CommunIty Detection method designed to address the limitations of classic objective-driven methods that struggle with complex graph structures and deep-learning approaches that improve performance at the expense of interpretability and their reliance on labeled data and training. Inspired by phase-transition kinetics in natural systems, where complex structures emerge through initialization, merging, refinement, and selection, LUCID is structured as a four-stage pipeline where the LLM acts as an automatic rule inducer, which converts implicit model knowledge into explicit and interpretable logical structures.

The four stages of the LUCID framework are:

  1. Local-View Community Initialization: This stage encodes local graph structure with k-ego contexts and unsupervised node roles. It transforms the graph into localized k-ego representations and assigns unsupervised node role labels (via K-means clustering on Jaccard similarity percentiles) to help the LLM understand local structural semantics.

  2. Multi-factor Community Merge: In this stage, the LLM first derives a multi-factor decision tree for evaluating structural compatibility between communities. This decision tree is then applied to guide merge decisions between neighboring communities. The process utilizes similarity-based sampling to restrict merge candidates to structurally relevant nodes and density-based scheduling to prioritize merges in dense regions and regulate execution order.

  3. Multi-grain Community Refinement: This stage applies LLM-induced coarse-to-fine rules in parallel to reduce boundary noise. The LLM is guided to induce coarse-to-fine rules for refining node memberships in the community and removing boundary noise, establishing a hierarchy of coarse-to-fine node refinement rules that includes coarse, medium-grained, and fine-grained rules.

  4. Global-view Community Selection: This stage identifies high-quality communities with topological compactness and boundary clarity. It utilizes RDC (Ranking with Density and Conductance), a density-regularized conductance scoring function that incorporates a penalty derived from the minimum spanning tree constraint to balance internal cohesion with external separation.

Extensive experiments on diverse real-world datasets (Facebook, Amazon, Livejournal, DBLP, and Twitter) demonstrate that LUCID, as an unsupervised approach, achieves state-of-the-art performance and consistently outperforms leading unsupervised and semi-supervised baselines. Specifically, LUCID significantly outperforms existing unsupervised methods across all datasets, achieving average improvements of 20.7% in F1 score and 31.9% in Jaccard score compared to the best unsupervised baseline. Furthermore, LUCID consistently outperforms state-of-the-art semi-supervised baselines, achieving average relative gains of 10.1% in F1 and 16.6% in Jaccard over the strongest semi-supervised baseline on each dataset.

Improvements for AI systems

1. Dynamic Multi-Agent Orchestration System

  • Improvement: Integrate the Multi-factor Community Merge and Multi-grain Community Refinement stages into multi-agent frameworks.

  • Capability: The system can autonomously organize a swarm of LLM-based agents into specialized, hierarchical functional units. It can detect emerging task requirements through agent interaction patterns and dynamically restructure agent communities using explicit logical rules, allowing for real-time, self-organizing coordination in complex, evolving environments.

2. Self-Explaining Knowledge Graph (KG) Reasoning Engine

  • Improvement: Implement the Automatic Rule Inducer mechanism to bridge the gap between implicit graph embeddings and symbolic logic.

  • Capability: Instead of providing black-box relationship predictions, the system can convert latent structural patterns within a KG into human-readable, multi-factor decision trees. This allows the AI to provide explicit, logical justifications (e.g., Entity A is linked to Entity B because they share X local role and satisfy Y structural density constraints) for every discovered connection or cluster.

3. Automated Data-Centric Curation Pipeline

  • Improvement: Apply the Global-view Community Selection (RDC scoring) to massive, unstructured relational datasets.

  • Capability: The system can automatically identify and extract high-quality, topologically compact subsets of data to serve as training sets. By using density-regularized conductance to select representative clusters, the system can eliminate redundant or noisy data, ensuring that training sets are structurally diverse and minimizing the risk of model bias caused by over-represented data clusters.

4. Interpretable Graph Neural Network (GNN) Auditor

  • Improvement: Use the Local-View Community Initialization and Multi-grain Refinement stages as a post-hoc interpretability layer for deep GNNs.

  • Capability: The system can provide a structural audit trail for GNN predictions. It can decompose a node's classification into a hierarchy of coarse-to-fine logical rules, explaining whether a prediction was driven by local k-ego neighborhood semantics or broader community-level structural roles, thereby making deep graph learning models transparent to human operators.

5. Autonomous Network Anomaly Detection System

  • Improvement: Utilize the four-stage pipeline to model real-time communication or transaction graphs.

  • Capability: The system can establish a baseline of normal community structures using unsupervised role labeling. It can then detect sophisticated, low-signal attacks (such as coordinated botnets or money laundering rings) by identifying nodes that fail to conform to the LLM-induced structural rules or that disrupt the topological compactness of established functional communities.

Sources

Related papers