ProbGuard: Calibrated Safety Risk Estimation from LLM Output Distributions
Xinzhe Huang, Biwu Yao, Kedong Xiu, Mengnan Zhao, Di Wang, Puning Zhao, Tianhang Zheng
Zhejiang University · Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security · University of Electronic Science and Technology of China · Anhui University · King Abdullah University of Science and Technology · Sun Yat-sen University
cs.LG
Submitted: 2026-08-11
Updated: 2026-08-12
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 100/100
The gist: ProbGuard is the first completely probabilistic, architecture-agnostic guardrail that leverages LLM early output distributional signals to estimate and calibrate the safety risk of continued
Terminology
Summary
ProbGuard is the first completely probabilistic, architecture-agnostic guardrail that leverages LLM early output distributional signals to estimate and calibrate the safety risk of continued generation, enabling early stopping of unsafe ongoing outputs. Existing guardrails (e.g., Llama-Guard3, ShieldGemma) formulate safety assessment as deterministic classification over discrete token sequences, which cannot express the uncertainty inherent in early generation and discards probabilistic information in the LLM output distribution. Streaming monitors (e.g., SCM) assess only observed prefixes, and feature-probing methods (e.g., ShieldHead, TPCs) rely on architecture-specific hidden states, limiting transferability.
ProbGuard formulates safety risk as the probability that continued generation from the current state will produce an unsafe response: Ck = P(J(R)=1 x, Sk), where J(R) is a binary safety judge. This is rewritten as an expectation over the distribution of possible full responses induced by the LLM's generation dynamics, and estimated via Monte Carlo sampling of N=16 independent continuations (up to 512 tokens, temperature 1.0). The calibration target is constructed by averaging binary safety judgments from a trained judge model called CalibEval, which achieves an F1 of 0.943 on the PKU evaluation set in 4.9s, outperforming LLM-based judges (DeepSeek-R1: 0.786, GPT-5: 0.827, Gemini3-Pro: 0.833) while having a much lower false positive rate (0.058 vs. 0.392–0.490).
For the probabilistic input, ProbGuard encodes each LLM output distribution as a probability-weighted representation in its own embedding space. At each decoding step, the top-K=50 candidate tokens are decoded to text, retokenized with ProbGuard's tokenizer, averaged into embeddings, and weighted by normalized probabilities. This avoids dependence on the protected LLM's hidden states and tokenizer, making the approach architecture-agnostic. ProbGuard is trained with a negative log-likelihood loss to align its predicted risk with the Monte Carlo calibration target.
Experiments cover three target LLMs (Llama3-8B-it, Qwen3-8B, Gemma2-9B-it), three safety datasets (PKU, WildGuard, SEval), and six jailbreak attacks (GCG, COLD-Attack, AdvPrefix, AdvPrompter, PAIR, ECLIPSE). ProbGuard is compared against 13 baselines across four categories: confidence-based (Verbalized, P(True)), guardrails (Llama-Guard3, ShieldGemma, Qwen3Guard, GPT-Safeguard, XGuard, PolyGuard), streaming monitors (Qwen3Guard-stream, SCM), and feature-probing (linear probes, ShieldHead, TPCs). Metrics are Brier score and Expected Calibration Error (ECE).
Main results: ProbGuard achieves the best calibration performance across all nine model–dataset combinations at prefix length k=10, reducing average Brier score by 79.6% and ECE by 71.9% over the best baseline. For example, on Qwen3-8B with PKU, ProbGuard achieves Brier/ECE of 0.0141/0.0249, while the best baseline (Qwen3Guard-stream) achieves 0.1151/0.0979. ProbGuard reduces average Brier by approximately 91%, 81%, and 85% relative to Llama-Guard3, ShieldGemma, and Qwen3Guard, respectively.
For early intervention against jailbreak attacks, ProbGuard-8B limits the attack success rate (ASR) to at most 1% across all six attacks on both AdvBench and HarmBench after observing only the first ten decoding steps, with an overall average ASR of 0.75% (AdvBench) and 0.67% (HarmBench). The best baseline (Llama-Guard3) achieves 3.67% and 1.17% respectively. ProbGuard-4B achieves 0.83% and 1.17%, and ProbGuard-0.6B achieves 2.83% on both datasets.
Further analysis shows: (1) Efficiency: ProbGuard-8B processes 1,000 samples in 36.4s at k=10, reducing latency by 52.7% compared to GPT-Safeguard (76.9s), while ProbGuard-0.6B uses 3.32GB GPU memory and 12.8s. (2) Calibration across decoding steps: From k=5 to k=20 (including extrapolation beyond training range), Brier scores decrease from 0.0174 to 0.0157 (8B), 0.0195 to 0.0167 (4B), and 0.0241 to 0.0232 (0.6B); ECE decreases from 0.0227 to 0.0140, 0.0300 to 0.0146, and 0.0412 to 0.0383 respectively. (3) Monte Carlo sampling budget: N=16 achieves a Flip Rate of 3.60% and Spearman correlation of 0.9197 relative to N=128, providing a favorable accuracy-efficiency trade-off. (4) Input representations: Output distributional signals consistently outperform token-based probabilities (which deteriorate on Llama3-8B-it) and hidden-state activations (which do not transfer across architectures).
Improvements for AI systems
Improvements to AI Systems:
- Add Probabilistic Safety Guardrails to LLM Inference Pipelines
-
Integrate ProbGuard as a lightweight, architecture-agnostic safety layer that monitors the LLM’s output distribution at every decoding step (e.g., after 10 tokens).
-
The improved system can early-stop unsafe generations (e.g., jailbreak responses) with a success rate below 1% across diverse attacks (GCG, PAIR, etc.), while reducing false positives by up to 8× compared to deterministic guardrails like Llama-Guard3.
-
This enables real-time safety enforcement without retraining the target LLM or accessing its hidden states.
- Replace Deterministic Safety Classifiers with Calibrated Risk Estimators
-
Use ProbGuard’s Monte Carlo–calibrated probability (P(unsafe prefix)) instead of binary “safe/unsafe” labels from models like ShieldGemma.
-
The improved system can quantify uncertainty in safety decisions, allowing downstream applications to set risk thresholds (e.g., block only if risk > 0.9) and reduce over-blocking of benign content—critical for open-ended chatbots, code assistants, and creative writing tools.
- Enable Cross-Architecture Safety Transfer
-
Deploy ProbGuard’s probability-weighted token embeddings (top-K=50, retokenized) as a universal safety signal, independent of the target LLM’s tokenizer or hidden states.
-
The improved system can protect any LLM (Llama, Qwen, Gemma) without fine-tuning, making safety monitoring plug-and-play for new models—ideal for rapidly evolving open-source ecosystems.
- Optimize Latency and Memory for Streaming Safety
-
Use ProbGuard-0.6B (3.32GB GPU, 12.8s per 1,000 samples) for edge or real-time deployments, or ProbGuard-8B (36.4s per 1,000 samples) for high-accuracy cloud serving.
-
The improved system can perform safety checks in under 40ms per generation step, enabling seamless integration into interactive agents, API gateways, and multi-turn dialogue systems without noticeable user-perceived lag.
- Improve Jailbreak Defense via Early Detection
-
Leverage ProbGuard’s ability to detect unsafe trajectories from the first 10 tokens, even for attacks that evolve over longer contexts (e.g., COLD-Attack, ECLIPSE).
-
The improved system can preemptively terminate harmful outputs before they are fully generated, reducing exposure to toxic content and lowering moderation costs by 80% compared to post-hoc filtering.
- Calibrate Safety Confidence for Human-AI Collaboration
-
Use ProbGuard’s well-calibrated risk scores (ECE < 0.025) to inform human reviewers or automated escalation systems.
-
The improved system can flag high-risk generations with precise confidence, enabling selective human review only for ambiguous cases, thus improving workflow efficiency in content moderation pipelines.
- Extend to Multi-Turn and Long-Context Safety
-
Apply ProbGuard’s distributional monitoring at every turn or after each context window extension (k=5 to k=20, including extrapolation).
-
The improved system can maintain calibrated safety across long conversations (e.g., 20+ turns) without performance degradation, preventing delayed jailbreaks that exploit context accumulation.
- Reduce Computational Overhead of Safety Ensembles
-
Replace multiple baseline guardrails (e.g., Llama-Guard3 + SCM + probes) with a single ProbGuard model that achieves 79.6% lower Brier score and 71.9% lower ECE.
-
The improved system can cut safety-related compute by up to 50% while improving accuracy, freeing resources for core generation tasks or higher throughput.
Abstract
Recent research on Large Language Model (LLM) safety has widely adopted guardrails to identify unsafe LLM outputs. Existing guardrails typically formulate safety assessment as a deterministic classification task, mapping a discrete token sequence to a discrete safety label. However, this paradigm has two limitations: First, safety assessment is inherently an uncertain problem, particularly during the early generation state. Second, relying solely on discrete token sequences discards the rich probabilistic information embedded in the LLM output distribution. To address these limitations, we propose the first completely probabilistic architecture-agnostic guardrail ProbGuard to leverage the LLM early output distributional signals for estimating and calibrating the safety probability, thereby enabling early stopping of unsafe ongoing outputs. Specifically, given an LLM's generated prefix distribution, we formulate the safety risk as the unsafe probability of its continued generation dynamics and estimate this risk by Monte-Carlo sampling. Through post-training on the distributional signals and calibrated safety risk, ProbGuard achieves the best calibration performance across all nine model--dataset combination settings, reducing the average Brier score and ECE by 79.6% and 71.9%, respectively, over the best baseline. ProbGuard further limits the attack success rate to at most 1% across six representative jailbreak attacks after observing the LLM early output distributions from only the first ten decoding steps.
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Understanding intermediate layers using linear classifier probes
- Jailbreaking Black Box Large Language Models in Twenty Queries
- The Llama 3 Herd of Models
- NonTextual Target Attack
- DualBreach: Efficient Dual-Jailbreaking via Target-Driven Initialization and Multi-Target Optimization
- Language Models (Mostly) Know What They Know
- Understanding the Effects of RLHF on LLM Generalisation and Diversity
- PolyGuard: A Multilingual Safety Moderation Tool for 17 Languages
- From Judgment to Interference: Early Stopping LLM Harmful Outputs via Streaming Content Monitoring
- YuFeng-XGuard: A Reasoning-Centric, Interpretable, and Flexible Guardrail Model for Large Language Models
- Teaching Models to Express Their Uncertainty in Words
- Detecting High-Stakes Interactions with Activation Probes
- Beyond Linear Probes: Dynamic Safety Monitoring for Language Models
- DARWIN: Evolving Jailbreak Adversary and Guardrail for LLM Safety Evaluation and Protection
- Gemma 2: Improving Open Language Models at a Practical Size
- AgentSnare: Learning to Delay, Divert, and Defuse Autonomous Penetration Agents
- DataShield: Uncovering Risky Fine-Tuning Data Across LLMs Through Consensus Subspace Alignment
- Dynamic Jailbreaking Attack
- On Calibration of Large Language Models: From Response To Capability
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks