How Semantically Stable Are LLM Refusals? Measuring Confusion in Local Safety Boundaries
cs.CL, cs.AI
Submitted: 2025-11-30
Updated: 2026-09-12
Comments: Accepted at the 2026 IEEE International Conference on Data Science and Advanced Analytics (DSAA 2026)
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- ShieldGemma: Generative AI Content Moderation Based on Gemma
- Safety-Tuned LLaMAs: Lessons From Improving the Safety of Large Language Models that Follow Instructions
- Towards Safety and Helpfulness Balanced Responses via Controllable Large Language Models
- Constitutional AI: Harmlessness from AI Feedback
- Llama Guard: LLM-based Input-Output Safeguard for Human-AI Conversations
- Automatic Pseudo-Harmful Prompt Generation for Evaluating False Refusals in Large Language Models
- OR-Bench: An Over-Refusal Benchmark for Large Language Models
- FalseReject: A Resource for Improving Contextual Safety and Mitigating Over-Refusals in LLMs via Structured Reasoning
- The Llama 3 Herd of Models
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Red Teaming Language Models with Language Models
- Curiosity-driven Red-teaming for Large Language Models
- Universal and Transferable Adversarial Attacks on Aligned Language Models
- Jailbreaking ChatGPT via Prompt Engineering: An Empirical Study
- Jailbreaking Leading Safety-Aligned LLMs with Simple Adaptive Attacks
- SmoothLLM: Defending Large Language Models Against Jailbreaking Attacks
- Jailbreak and Guard Aligned Language Models with Only Few In-Context Demonstrations
- Defending LLMs against Jailbreaking Attacks via Backtranslation
- Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components
- Can LLMs Refuse Questions They Do Not Know? Measuring Knowledge-Aware Refusal in Factual Tasks
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering