Graph Tokenization for Bridging Graphs and Transformers
cs.LG, cs.AI
Submitted: 2026-03-11
Updated: 2026-03-11
Comments: Accepted as a poster at ICLR 2026. Code is available at https://github.com/BUPT-GAMMA/Graph-Tokenization-for-Bridging-Graphs-and-Transformers
Journal ref: The Fourteenth International Conference on Learning Representations (ICLR), 2026
Code: https://github.com/BUPT-GAMMA/Graph-Tokenization-for-Bridging-Graphs-and-Transformers
License: http://creativecommons.org/licenses/by/4.0/
The gist: The success of large pretrained Transformers is closely tied to tokenizers, which convert raw input into discrete symbols.
Terminology
Abstract
The success of large pretrained Transformers is closely tied to tokenizers, which convert raw input into discrete symbols. Extending these models to graph-structured data remains a significant challenge. In this work, we introduce a graph tokenization framework that generates sequential representations of graphs by combining reversible graph serialization, which preserves graph information, with Byte Pair Encoding (BPE), a widely adopted tokenizer in large language models (LLMs). To better capture structural information, the graph serialization process is guided by global statistics of graph substructures, ensuring that frequently occurring substructures appear more often in the sequence and can be merged by BPE into meaningful tokens. Empirical results demonstrate that the proposed tokenizer enables Transformers such as BERT to be directly applied to graph benchmarks without architectural modifications. The proposed approach achieves state-of-the-art results on 14 benchmark datasets and frequently outperforms both graph neural networks and specialized graph transformers. This work bridges the gap between graph-structured data and the ecosystem of sequence models. Our code is available at here.
Sources
- LLaGA: Large Language and Graph Assistant
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- A Large-Scale Database for Graph Representation Learning
- AST: Audio Spectrogram Transformer
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- Semi-Supervised Classification with Graph Convolutional Networks
- Tabular Foundation Models are Strong Graph Anomaly Detectors
- Can Classic GNNs Be Strong Baselines for Graph-level Tasks? Simple Architectures Meet Excellence
- Large Language Models: A Survey
- GraphBPE: Molecular Graphs Meet Byte-Pair Encoding
- Graph-Mamba: Towards Long-Range Graph Sequence Modeling with Selective State Spaces
- Data-centric Federated Graph Learning with Large Language Models
- VQGraph: Rethinking Graph Representation Space for Bridging GNNs and MLPs
- Graph2text or Graph2token: A Perspective of Large Language Models for Graph Learning
- Graph-Bert: Only Attention is Needed for Learning Graph Representations
- GPatcher: A Simple and Adaptive MLP Model for Alleviating Graph Heterophily
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks