Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

arXiv:2609.05074 · cs.CL · Submitted 2026-09-04 · Read on arXiv

cs.CL

Submitted: 2026-09-04

Updated: 2026-09-04

License: http://creativecommons.org/licenses/by-sa/4.0/

The gist: We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection.

Terminology

Abstract

We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The score combines directional influence on the logits with structural contribution within the residual stream, enabling a multi-scale analysis at the head, layer, and network levels. Applied to a DeBERTa model specialized for prompt injection detection, our framework reveals distinct decision behaviours between correct and erroneous predictions. Our method provides an effective compromise between fine-grained circuit analysis and global output-based methods, and offers a systematic way to study decision mechanisms in Transformer classifiers.

Related papers