Call Neighbours Yourself: Graph Walks with Destination-Conditioned On-Policy Self-Distillation
cs.AI, cs.CL
Submitted: 2026-08-30
Updated: 2026-08-30
Comments: EMNLP 2026
Code: https://github.com/superallen13/CNY
License: http://creativecommons.org/licenses/by/4.0/
The gist: Reasoning over text-attributed graphs (TAGs) requires large language models (LLMs) to combine a node's text with evidence distributed across its neighbourhood.
Terminology
Abstract
Reasoning over text-attributed graphs (TAGs) requires large language models (LLMs) to combine a node's text with evidence distributed across its neighbourhood. Existing methods fix the set of accessible neighbours before generation, forcing reasoning to operate over a static context and preventing the model from acquiring missing evidence during inference. We argue that neighbour selection should itself be part of the reasoning process. To this end, we propose Call Neighbours Yourself (CNY), a framework that enables LLMs to proactively explore graph neighbourhoods through topology-constrained graph-walk actions. Instead of reasoning over a pre-selected neighbour set, CNY exposes lightweight neighbour previews and learns when to expand candidate neighbours for additional evidence. To address the delayed-credit challenge of neighbour exploration, we introduce destination-conditioned on-policy self-distillation, which retrospectively evaluates a selected neighbour after its content is revealed and converts the resulting change in action preference into an action-level training signal. Experiments on standard TAG reasoning benchmarks under a unified raw-text setting show that CNY consistently outperforms fixed-context post-training baselines. Furthermore, the learned exploration policy transfers to unseen graphs and to a graph-level task not encountered during training. Code is available at https://github.com/superallen13/CNY.
Sources
- Mistral 7B
- GFMate: Empowering Graph Foundation Models with Test-time Prompt Tuning
- Why Does Self-Distillation (Sometimes) Degrade the Reasoning Capability of LLMs?
- Understanding R1-Zero-Like Training: A Critical Perspective
- The Llama 3 Herd of Models
- Wiki-CS: A Wikipedia-Based Benchmark for Graph Neural Networks
- Proximal Policy Optimization Algorithms
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Self-Distillation Enables Continual Learning
- R1-Searcher: Incentivizing the Search Capability in LLMs via Reinforcement Learning
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- What Information Matters? Graph Out-of-Distribution Detection via Tri-Component Information Decomposition
- OpenClaw-RL: Train Any Agent Simply by Talking
- NPG-Muse: Scaling Long Chain-of-Thought Reasoning with NP-Hard Graph Problems
- Qwen2.5 Technical Report
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection