Attention-Path Fragility as an Uncertainty Signal in Large Language Models

arXiv:2608.11138 · cs.CL, cs.AI · Submitted 2026-08-11 · Read on arXiv

Minsoo Kim, Sungyoung Ji, Kisung Moon, Ilyong Yoon

POSCO Holdings Future Technology Research Institute

cs.CL, cs.AI

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 19 pages, Under review

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: The paper proposes that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is fragile under perturbation of

Terminology

Summary

The paper proposes that a model's uncertainty about a token is reflected not only in the breadth of its output distribution but also in whether a confident prediction is fragile under perturbation of its attention pathways. This is instantiated as ASMI (Attention-Subnetwork Mutual Information), a training-free estimator that masks attention heads and measures the BALD mutual information among the resulting subnetworks, with a semantic-agreement kernel to discount surface-form disagreement.

The signal is not a restatement of output confidence: on grounded QA an out-of-fold test shows it adds error-predictive information beyond single-pass confidence and entropy, concentrated in confident-but-fragile predictions, where acting on it roughly halves the retained error of a confidence filter. The distinctness is regime-graded, so ASMI predicts its own domain of applicability, strong where answers are routed through provided context and bounded by design where they are recalled from parametric knowledge. Sem-ASMI reads the signal from a single greedy response, without the stochastic generations the strongest baselines require, and ties or beats Semantic Entropy on ten of the twelve grounded benchmark-backbone settings. Across the same twelve settings, the best ASMI variant, typically the adaptive one reusing the ten samples already drawn for the baselines, ties or leads the strongest baseline in eight, significantly in three under a paired test. On parametric QA all variants revert to or below the zero-cost MSP baseline, exactly as predicted, and the estimates are near-deterministic across reruns. A head-level analysis shows that what tracks this boundary is not the presence of head-level fragility but whether that fragility couples to errors.

The method works by, for each Monte Carlo sample, masking attention heads at a target layer and quantifying disagreement among the resulting predictive distributions via the BALD mutual-information decomposition, with a semantic-agreement kernel. The estimator is training-free and needs only the choice of target layer. The intuition rests on a well-documented property of multi-head attention: many heads can be pruned with minimal loss, yet a high-level behavior can hinge on a single head, and a small subset is uncertainty-aware. Masking leaves a genuinely supported prediction intact, since redundant pathways carry the same information, and destabilizes a fragile one. Crucially, this fragility is invisible to a single forward pass: a token can be assigned high confidence yet collapse under masking, and these confident-but-fragile tokens carry error information that output-confidence measures miss.

The paper characterizes existing estimators along two axes: their variation source (stochastic output sampling versus structural model perturbation) and their measurement target (output spread versus computational path dependence). ASMI occupies the structural-perturbation, path-dependence corner, and its position yields a testable hypothesis: the signal should be informative where correctness depends on attending to specific context tokens, as in retrieval-grounded QA, and end where uncertainty originates elsewhere, as in parametric knowledge recall. This hypothesis is tested behaviorally, with grounded against parametric QA as a designed contrast, and internally, asking whether the signal adds error information beyond output confidence and how its head-level signature tracks the boundary.

The contributions are: proposing ASMI, a training-free token-level uncertainty estimator that probes attention-path fragility by random head masking with a semantic-agreement kernel, and reads it from a single greedy response, competitive with the strongest sample-diversity baselines which each draw ten stochastic generations, with near-deterministic estimates across reruns; showing the signal is distinct from output confidence, not a proxy for it, confirmed by an out-of-fold test that it adds error-predictive information beyond single-pass confidence and entropy, concentrated in confident-but-fragile predictions; and showing the distinctness is regime-graded, so ASMI predicts its own domain of applicability, leading on context-routed QA and reverting to or below the free MSP baseline on parametric recall, with a head-level analysis showing that what tracks this boundary is whether fragility couples to errors.

The methodology involves attention subnetwork sampling, where for each Monte Carlo sample a binary mask is drawn with each head masked with probability p, applied to head outputs before the output projection, with no rescaling. The prefix below the target layer is computed once and shared across all subnetworks, reducing the Monte Carlo cost from S(Cpre + Csuf) to Cpre + S Csuf FLOPs. The mutual information is computed as the difference between the entropy of the ensemble mean distribution and the mean of the entropies of the individual subnetwork distributions, approximated by top-K candidates plus a tail bucket. The semantic agreement weighting computes pairwise agreement between samples using a similarity matrix based on the output projection matrix, and the token-level score is MIt(1 - At). The adaptive variant gates the semantic weighting by the diversity of N sampled responses, with the gate approaching 1 when responses are near-paraphrases and 0 when they genuinely diverge.

The experiments use a two-condition contrast: context-routed condition with CoQA, SQuAD, and BabiQA, and a parametric condition with closed-book TriviaQA as a negative control. All datasets follow the LM-Polygraph pipeline with default prompting and standard splits, scored by AlignScore. The paper compares against 17 baseline UE methods spanning four families: information-based, sample-diversity, probing, and attention-based. The primary backbone is Qwen3-4B-base, with Qwen3-8B-base, Llama-2-7B, and Mistral-7B for generality. The operating point is (p, S) = (0.15, 40) with top-K truncation K = 64, and the masked layer is at relative depth d = 60% as representative by aggregate PRR.

The main results show that Sem-ASMI, needing only the greedy response, is competitive on grounded QA with the strongest sample-diversity baselines which each draw ten stochastic generations per input. Measured cost tracks input length, favoring long-context inputs typical of context-routed QA. Under a cluster-respecting paired bootstrap the best ASMI variant ties or leads the per-column best baseline in eight of the twelve grounded columns, significantly in three, and in every remaining lead the sampling-free Sem-ASMI ties or beats Semantic Entropy from the greedy response alone. On parametric recall all variants revert to or below the free MSP baseline, and the paired test places the top variant significantly below the per-column best baseline in three of the four columns.

The distinct-signal analysis shows that the semantically-weighted score adds error-predictive information beyond confidence and entropy, concentrated in confident-but-fragile predictions. The residual detects errors above chance on every grounded dataset, in five of the six backbone-dataset cells at AUROCs of 0.54 to 0.60. The effect is localized: splitting sequences at the medians of MSP and usem, the confident-but-fragile cell is wrong three times as often as the confident-robust cell, on BabiQA 31% against 10%, a gap MSP alone cannot see. The signal is regime-graded exactly as the design predicts: the raw gap between fragile and robust cells also appears on parametric TriviaQA, but there it is entirely absorbed by single-pass confidence and entropy with the residual at chance on both backbones, whereas on grounded QA a residual beyond both remains.

In deployment, among the answers MSP rates as confident, abstaining the ASMI-fragile ones roughly halves the retained error on grounded QA, from 16.0% to 5.6% on BabiQA and from 13.6% to 7.2% on CoQA, against 9.2% and 10.8% for the same abstention budget spent on entropy, and the gap holds across the whole abstention range. On parametric TriviaQA the order reverses, so the boundary holds in deployment too. The gain is localized rather than global: added to a logistic selector over MSP and entropy fit across all coverage levels, ASMI leaves the overall risk-coverage curve unchanged, because its error information concentrates in the confident stratum.

The boundary is sharp: on the same SQuAD examples with the passage removed, Sem-ASMI sits significantly below MSP (-0.054, 95% CI [-0.088, -0.021]) while the open-book condition is a statistical tie, so the pattern reproduces within a single dataset when only the knowledge source moves. Every remaining exception is mapped, not noise. All TriviaQA cells sit at or below MSP as the boundary predicts, and the one grounded exception has an identified cause: on BabiQA with Llama-2-7B head masking barely moves the output, so correct and incorrect answers carry the same near-zero MI and ASMI has no fragility to read while MSP still ranks the errors, an over-robustness that is a model property rather than an artifact. The cell announces itself before any correctness label exists: the same masking response that produces the score also reports when a model is too robust for it to carry information.

The structural analysis shows that at the operating layer the criticality of an uncertain token is distributed across about ten heads rather than concentrated on a few. The coupling is modest in magnitude but consistent in sign, following the grounded and parametric divide. The association ρ(MIt, zt) is positive on all three grounded benchmarks and negative on parametric TriviaQA, with intervals excluding zero, and mutual information tracks this head structure more than single-pass entropy does on the grounded cells. A causal probe confirms that the fragility is head borne: ablating a token's few most critical heads flips the prediction for high-MI tokens while leaving low-MI tokens almost unchanged. The dependence is present on all four benchmarks, so MIt measures genuine reliance on specific heads. What follows the grounded and parametric divide is whether that reliance tracks uncertainty. For k ≥ 2 the flip rates are in fact highest on TriviaQA, so fragility itself is strongest under parametric recall, and it is the alignment with correctness that is missing there.

The limitations note that the BALD form admits an epistemic reading of MIt under the variational view of inference-time dropout, but entropy-based decompositions of epistemic and aleatoric uncertainty face formal objections, so MIt is used as a disagreement functional with no claim of empirical separation of the two. Distinctness replicates across both backbones and all three grounded datasets, with two caveats: the CoQA residual on Qwen3-8B does not exclude chance, and the SQuAD effect on Qwen3-4B lies at the estimator's own Monte Carlo resolution. On the parametric control the top-K truncation retains 77.3% of the high-MI mass against over 95% on grounded data, but re-scoring TriviaQA at K = 256 raises mean token coverage to 0.97 and leaves Sem-ASMI significantly below MSP, so the parametric-side deficit is not a truncation artifact. Extending the two-axis characterization to reasoning, open-ended generation, instruction-tuned models, and further parametric tasks is future work.

The conclusion states that attention-path fragility was introduced as an uncertainty signal, instantiated as ASMI, a training-free estimator that probes attention-pathway redundancy by random head masking and reads it from a single greedy response. The semantically-weighted score adds error-predictive information beyond confidence and entropy, concentrated in confident-but-fragile predictions. Its value is local rather than global: a global selector cannot exploit it, yet in the confident stratum, where a deployed system acts, it roughly halves the error. The distinctness is regime-graded and confirmed by a designed contrast: from that single greedy pass ASMI is competitive with the strongest sampling-based baselines on context-routed QA, and reverts to or below the free MSP baseline on parametric recall, exactly as the hypothesis requires. A head-level analysis traces this boundary to whether fragility couples to errors, and a causal ablation confirms the fragility is carried by identifiable critical heads. The lesson outlasts the estimator: for a signal tied to a specific computation, where it applies and where its value lands follow from the locus of a model's errors, and ASMI is offered as a case for treating a UE signal's domain of applicability as a predictable property rather than a discovered one.

Improvements for AI systems

Improvements to AI Systems:

  1. Confidence-calibrated abstention for context-grounded tasks: AI systems can now identify confident-but-fragile predictions—tokens assigned high probability that collapse under attention-head masking. By abstaining on these fragile tokens, the system reduces retained error by roughly half (e.g., from 16.0% to 5.6% on BabiQA) compared to entropy-based abstention, enabling safer deployment in retrieval-augmented generation, question answering over documents, and fact-checking pipelines.

  2. Regime-aware uncertainty estimation: The system can predict when its uncertainty signal is valid. It automatically switches between ASMI (for context-routed tasks where answers depend on specific input tokens) and simpler confidence baselines (for parametric recall from memorized knowledge). This prevents over-reliance on fragile-pathway signals in settings where they add no value, improving robustness across mixed workloads.

  3. Single-pass uncertainty without sampling overhead: The Sem-ASMI variant reads uncertainty from one greedy response, eliminating the need for multiple stochastic generations. This reduces inference cost by up to 10× while matching or beating Semantic Entropy on 10 of 12 grounded benchmark settings, enabling real-time uncertainty estimation in latency-sensitive applications like chatbots and interactive assistants.

  4. Head-level error localization for interpretability: The system identifies which specific attention heads carry critical information for each token's prediction. A causal probe can flip predictions by ablating the top critical heads for high-uncertainty tokens while leaving low-uncertainty tokens unaffected. This enables debugging of model behavior, identifying over-reliance on spurious attention patterns, and guiding targeted fine-tuning or head-pruning for error correction.

  5. Self-diagnosing model robustness: The masking response itself reports when the model is too robust for the estimator to be informative (e.g., on BabiQA with Llama-2-7B, where head masking barely moves outputs). The system can detect this over-robustness before any correctness label exists, allowing it to fall back to alternative uncertainty methods or flag that the model's predictions are not path-sensitive, preventing false confidence in brittle systems.

  6. Cost-adaptive uncertainty computation: The system can reuse Monte Carlo samples already drawn for other baselines (adaptive variant) or operate with zero extra samples (Sem-ASMI), with near-deterministic estimates across reruns. This makes uncertainty quantification practical for long-context inputs where sampling cost scales with sequence length, enabling deployment in large-scale document processing and multi-hop reasoning tasks.

  7. Risk-coverage optimization in confident strata: The system improves risk-coverage curves specifically in the high-confidence region where deployed systems typically act. By adding ASMI as a secondary filter on top of confidence thresholds, it reduces error in the confident stratum without altering global risk-coverage trade-offs, allowing safer high-precision operation in production settings like medical QA or legal document analysis.

  8. Transferable domain-of-applicability prediction: The system learns to predict where its uncertainty signal will be useful based on the locus of model errors (context-routed vs. parametric). This principle extends beyond ASMI to any computation-specific signal, enabling future estimators to be designed with built-in applicability boundaries rather than discovered post-hoc, improving generalization to new tasks and domains.

Sources

Related papers