Skip to results
MLSift

Titles, abstracts, or an arXiv ID

← Back to results
routineNLP & Language ModelsDeBERTa2609.05074

Influence Score and Transformers interpretability: Measure of the Effective Impact of Attention Heads at inference time

Lisa Bouger, Yannick Teglia, Philippe Loubet Moundi

cs.CL

Abstract

We propose an influence score to quantify the contribution of attention heads to classification decisions in Transformer-based models designed for prompt injection detection. The score combines directional influence on the logits with structural contribution within the residual stream, enabling a multi-scale analysis at the head, layer, and network levels. Applied to a DeBERTa model specialized for prompt injection detection, our framework reveals distinct decision behaviours between correct and erroneous predictions. Our method provides an effective compromise between fine-grained circuit analysis and global output-based methods, and offers a systematic way to study decision mechanisms in Transformer classifiers.

Topics

Classified with taxonomy v2 on Mon, 7 Sept 2026.

Report a classification error

Loading the PDF downloads the document. Open it in your browser's viewer, or load it here.

Open PDF