Can Vision-Language Models Stay Helpful When Facing Implicit Risks? Intent-Privilege OPSD for Efficient Safety-Helpfulness Alignment
cs.CL
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/Tunanzzz/Meerkat-VL
Terminology
Sources
- On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes
- FigStep: Jailbreaking Large Vision-Language Models via Typographic Visual Prompts
- Distilling the Knowledge in a Neural Network
- CASA: Classification Augmented with Safety Attention for Robust Multimodal Alignment
- SGuard-v1: Safety Guardrail for Large Language Models
- HoliSafe: Holistic Safety Benchmarking and Modeling for Vision-Language Model
- Safe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs
- MOSSBench: Is Your Multimodal Language Model Oversensitive to Safe Queries?
- DREAM: Disentangling Risks to Enhance Safety Alignment in Multimodal Large Language Models
- MM-SafetyBench: A Benchmark for Safety Evaluation of Multimodal Large Language Models
- Safety Alignment for Vision Language Models
- SIA: Enhancing Safety via Intent Awareness for Vision-Language Models
- Qwen3 Technical Report
- SafeGRPO: Self-Rewarded Multimodal Safety Alignment via Rule-Governed Policy Optimization
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Learning by Distilling Context
- MMAligner: Safeguarding Multimodal Large Language Models through Representation Calibration
- CrossGuard: Safeguarding MLLMs against Joint-Modal Implicit Malicious Attacks
- SPA-VL: A Comprehensive Safety Preference Alignment Dataset for Vision Language Model
- Self-Distilled Reasoner: On-Policy Self-Distillation for Large Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering