Attacking Graph Foundation Models Through Their Shared Representation
Pankaj Kumar, Subhankar Mishra
cs.AI, cs.CR, cs.LG
Submitted: 2026-07-20
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: A graph foundation model generalizes across graph domains by mapping every input into one shared representation before any task reasoning.
Terminology
Abstract
A graph foundation model generalizes across graph domains by mapping every input into one shared representation before any task reasoning. We call this map the alignment layer, the component that separates a graph foundation model from a graph neural network, and we show it is a distinct attack surface that prior work has not studied. We attack it at inference time, with no access to training, on six public models spanning spectral tokenizers, text embedding spaces, and a discrete codebook. A directed representation-space perturbation collapses every model, but at a budget comparable to the representation norm a plain graph network also needs, with one exception: OpenGraph, whose spectral tokenizer collapses at a fifth of that budget, an alignment-specific fragility a plain network does not share and which a same-representation control traces to the tokenizer rather than the decoder. A realizable input-space attack that edits edges, features, or text removes at least half the correct predictions on three of the six models at peak. How much of this fragility an input-access attacker realizes tracks how directly the decoder reads the representation, and not the clean accuracy a task leaves; we measure this carrier gain structurally from the decoder's local Lipschitz sensitivity, and report clean-accuracy headroom as a within-model ordering heuristic that does not survive on realizable attacks.
Sources
- Understanding Adversarial Transfer: Why Representation-Space Attacks Fail Where Data-Space Attacks Succeed
- Adversarial Attacks and Defenses on Graphs: A Review, A Tool and Empirical Studies
- One for All: Towards Training One Graph Model for All Classification Tasks
- Graph Foundation Models: Concepts, Opportunities and Challenges
- Convergence Without Understanding: When Language Models Agree on Representations but Disagree on Reasoning
- GFT: Graph Foundation Model with Transferable Tree Vocabulary
- TrustGLM: Evaluating the Robustness of GraphLLMs Against Prompt, Text, and Structure Attacks
- Fully-inductive Node Classification on Arbitrary Graphs
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection