Steering Multimodal Large Language Models Decoding for Context-Aware Safety
cs.CL, cs.AI
Submitted: 2025-09-23
Updated: 2026-08-28
Comments: EMNLP 2026 Main
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
- OR-Bench: An Over-Refusal Benchmark for Large Language Models
- DreamLLM: Synergistic Multimodal Comprehension and Creation
- Planting a SEED of Vision in Large Language Model
- Baseline Defenses for Adversarial Attacks Against Aligned Language Models
- VideoChat: Chat-Centric Video Understanding
- MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
- Safety of Multimodal Large Language Models on Images and Texts
- Protecting Privacy in Multimodal Large Language Models with MLLMU-Bench
- Towards Safer Large Language Models through Machine Unlearning
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
- Valley: Video Assistant with Large Language model Enhanced abilitY
- Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Towards Trustworthy Retrieval Augmented Generation for Large Language Models: A Survey
- Kosmos-2: Grounding Multimodal Large Language Models to the World
- MIA-Bench: Towards Better Instruction Following Evaluation of Multimodal LLMs
- Is ChatGPT a General-Purpose Natural Language Processing Task Solver?
- Latent Jailbreak: A Benchmark for Evaluating Text Safety and Output Robustness of Large Language Models
- XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering