Query-Side Attacks on GNN-Based KGQA: Tracing Failures from Entity Linking to Answer Generation
cs.CL, cs.AI, cs.IR
Submitted: 2026-08-26
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by-sa/4.0/
The gist: GNN-based Knowledge Graph Question Answering (KGQA) pipelines process queries through four discrete stages: entity linking, subgraph retrieval, GNN reasoning, and answer generation.
Terminology
Abstract
GNN-based Knowledge Graph Question Answering (KGQA) pipelines process queries through four discrete stages: entity linking, subgraph retrieval, GNN reasoning, and answer generation. Standard robustness evaluations conflate stage-level failures into a single end-to-end metric, obscuring both the source of brittleness and the appropriate mitigation target. We ask which stage fails, and why, when the pipeline is subjected to adversarial perturbations on the input question. We introduce a stage-isolation protocol with two answer-preserving adversarial perturbations verified against the knowledge graph: Compositional Restructuring (CR) and Relation Synonym Swap (RS) target distinct stages while leaving entity seeds intact. Evaluated across ComplexWebQuestions and WebQSP, the results run counter to prevailing assumptions: the GNN reasoning stage retains near-baseline accuracy when the subgraph is intact, while subgraph construction accounts for over 99% of the end-to-end collapse under CR, occurring even when the gold answer is present in 74% of retrieved subgraphs. This exposes a fundamental distinction between answer presence and answer reachability that end-to-end metrics cannot detect, and places the mitigation target firmly at the subgraph construction stage rather than the reasoning model. Perturbed datasets and evaluation infrastructure are released at https://anonymous.4open.science/r/atkgrag-E85C.
Sources
- Phantom: General Backdoor Attacks on Retrieval Augmented Language Generation
- Enhancing Complex Question Answering over Knowledge Graphs through Evidence Pattern Retrieval
- The Llama 3 Herd of Models
- HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models
- Harnessing Deep LLM Participation for Robust Entity Linking
- BYOKG-RAG: Multi-Strategy Graph Retrieval for Knowledge Graph Question Answering
- GNN-RAG: Graph Neural Retrieval for Large Language Model Reasoning
- Reasoning on Graphs: Faithful and Interpretable Large Language Model Reasoning
- Are LLM-Enhanced Graph Neural Networks Robust against Poisoning Attacks?
- RAG Safety: Exploring Knowledge Poisoning Attacks to Retrieval-Augmented Generation
- BadRAG: Identifying Vulnerabilities in Retrieval Augmented Generation of Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering