Towards Multi-Label Graph Foundation Models: from Single-Vector Representation Learning to Multi-Semantic Basis Learning
Dongxiao He, Jiayu Zhang, Jitao Zhao, Yi Wang, Di Jin
Tianjin University
cs.AI
Submitted: 2026-07-31
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 40/100
The gist: The paper "Towards Multi-Label Graph Foundation Models: from Single-Vector Representation Learning to Multi-Semantic Basis Learning" proposes a novel framework called the Multi-Semantic Basis Graph
Terminology
Summary
The paper Towards Multi-Label Graph Foundation Models: from Single-Vector Representation Learning to Multi-Semantic Basis Learning
proposes a novel framework called the Multi-Semantic Basis Graph Foundation Model (MSB-GFM) to address the challenges of cross-domain multi-label node classification.
Problem Statement
The authors identify a critical gap in existing research: while multi-label node classification is essential because nodes often naturally exhibit multiple semantic attributes rather than a single semantic identity,
existing methods are often limited to in-domain scenarios
and fail to generalize to new domains. Conversely, existing Graph Foundation Models (GFMs) are built upon single-label assumption, where all nodes are arbitrarily regarded as containing only one class of semantic and embedded into a single representation.
For multi-label nodes, this approach essentially approximates multiple semantics with a single point in the representation space, inevitably leading to semantic entanglement and making simultaneous discrimination of multiple labels difficult.
Proposed Methodology: MSB-GFM
To address these limitations, the MSB-GFM framework is introduced, which consists of three core modules:
- Multi-Semantic Basis Learning: This module
models each multi-label node as an adaptive composition of semantic bases, thereby enabling flexible representational capacity for modeling multiple semantics.
The model maintains aglobally learnable set of semantic bases B = b 1, b 2,, b m,
where each basiscorresponds to one semantic direction.
The process involves:
-
Dimension Alignment: Using PCA to project node features from different domains into a
unified feature space.
-
Threshold-based Activation: Computing cosine similarity between projected embeddings and semantic bases, using a
threshold-based activation mechanism
to allow each node toactivate a variable number of semantic bases according to its semantic complexity.
-
Reconstruction and Fusion: Reconstructing the representation through
aggregation weights
and combining it with the original embedding viaresidual fusion.
-
Loss Functions: Optimization is driven by a
semantic matching loss
(L match), which encourages bases to align with semantics, and asemantic alignment loss
(L align), whichregularizes the enhanced representation to remain close to the original node representation.
-
Structure-aware Prototype Learning (SAP): This module aims to
extract topological information
that islargely invariant across different graph domains.
It utilizes DeepWalk to obtaininitial structural embeddings Z stru
and introducesa set of learnable structural prototypes P = p 1, p 2,, p n,
where each prototyperepresents a prototypical topological role shared across different graph domains.
The module usesTop- k structural prototypes
andweighted aggregation
to create an enhanced structural representation. It is optimized using astructural alignment objective
(L align) and adiversity objective
(L div), whichmaximizes the entropy of the prototype assignment... preventing prototype degeneration.
-
Domain-invariant Learning (DI): To mitigate
domain shifts
andextract domain-invariant representations,
the authorsintroduce a domain-invariant learning based on adversarial domain adaptation.
This involves adomain discriminator
and aGradient Reversal Layer (GRL)
thatencourages the encoder to discard domain-specific patterns and retain only transferable domain-invariant knowledge shared across graph domains.
Pre-Training and Downstream Adaptation
During pre-training, the model jointly optimizes the semantic representation learning, structural prototype learning, and domain-invariant learning objectives
via a total loss: L total = L sem + gamma 1 L stru + gamma 2 L adv. For downstream tasks, the encoder parameters are frozen, and a MLP classifier
is trained on the target domain using a small number of available labeled nodes
(one-shot setting) optimized with binary cross-entropy loss.
Experimental Results
The model was evaluated on four datasets: Humloc, PCG, Blogcatalog, and PPI. The results demonstrate that:
-
MSB-GFM demonstrates great performance, achieving the best or competitive results across nearly all datasets.
-
Compared to multi-label graph learning methods,
MSB-GFM shows significant improvements,
because existing methodsrely on domain-specific data to capture label correlations, which hinders their generalization to target domains.
-
Compared to existing GFMs,
MSB-GFM also achieves consistent improvements,
revealing that the single-vector representation paradigm of existing GFMs isinsufficient for multi-label scenarios.
-
Ablation studies confirm that
all components contribute positively to the performance,
specifically validating the necessity of the MSB, SAP, and DI modules foreffectively alleviating semantic entanglement and enabling universal graph encoders to support both multi-semantic modeling and cross-domain generalization.
Improvements for AI systems
1. Hierarchical Semantic Basis Expansion
-
Improvement: Replace the flat, globally learnable set of semantic bases with a multi-level, hierarchical basis structure (e.g., a tree-based or capsule-based hierarchy).
-
Capability: The system can perform multi-scale multi-label classification, allowing it to simultaneously identify broad category labels (e.g.,
Biology
) and fine-grained sub-labels (e.g.,CRISPR Gene Editing
) by activating different levels of the semantic hierarchy.
2. Temporal-Semantic Evolution Modeling
-
Improvement: Integrate a temporal encoding mechanism (such as Time-Aware Graph Neural Networks) into the MSB module to allow the semantic bases to evolve over time.
-
Capability: The system can model and predict how a node's multi-label identity shifts dynamically, such as tracking the evolving professional interests of a user in a social network or the changing topicality of a research paper in a citation graph.
3. Probabilistic/Uncertainty-Aware Semantic Activation
-
Improvement: Replace the hard
threshold-based activation
with a probabilistic mechanism, such as a Dirichlet distribution or a Bayesian approach, to model the intensity of basis activation. -
Capability: The system can provide a
confidence score
for each semantic component, enabling it to signal when a node's multi-label identity is ambiguous or when the current set of semantic bases is insufficient to explain the node's complexity.
4. Heterogeneous Relation-Aware Basis Learning
-
Improvement: Condition the semantic bases on edge types/relations, moving from a general MSB to a relation-specific MSB.
-
Capability: The system can resolve
semantic context switching,
where a node exhibits different semantic identities depending on the type of connection it has (e.g., a person node exhibitingacademic
semantics when connected via a co-authored edge, butsocial
semantics when connected via a friend edge).
5. Explainable Semantic Attribution
-
Improvement: Implement a decoding layer that maps the activated semantic bases and their corresponding aggregation weights back to human-readable feature subsets or subgraph patterns.
-
Capability: The system can provide
semantic justifications
for its predictions, moving beyond simple label output to explain why a node was assigned multiple labels (e.g.,Node X is labeled 'A' and 'B' because its structural prototype matches 'Pattern 1' and its features align with 'Semantic Basis 3'
).
6. Scalable Structural Prototype Distillation
-
Improvement: Replace the DeepWalk-based structural embedding with a lightweight, sampling-based structural encoder (like GraphSAGE) integrated with a vector quantization (VQ) layer for the prototypes.
-
Capability: The system can scale to massive, web-scale graphs (billions of nodes) while maintaining the ability to extract domain-invariant topological roles without the computational bottleneck of global structural embedding.
Sources
- Graph Neural Networks for Vehicular Social Networks: Trends, Challenges, and Opportunities
- Graph Neural Network for Crawling Target Nodes in Social Networks
- Foundation Models for Natural Language Processing -- Pre-trained Language Models Integrating Media
- Graph Foundation Models: A Comprehensive Survey
- AnyGraph: Graph Foundation Model in the Wild
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection