Narrow Multimodal Fine-Tuning Can Induce Emergent Misalignment
cs.LG, cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Code: https://github.com/shunchang-liu/vlm-em
Terminology
Sources
- Qwen3-VL Technical Report
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- LoRA Learns Less and Forgets Less
- Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
- Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Emerging Properties in Unified Multimodal Pretraining
- Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models
- Conditional misalignment: common interventions can hide emergent misalignment behind contextual triggers
- KTO: Model Alignment as Prospect Theoretic Optimization
- Narrow Fine-Tuning Erodes Safety Alignment in Vision-Language Agents
- GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning
- GPT-4o System Card
- LoRA Fine-tuning Efficiently Undoes Safety Training in Llama 2-Chat 70B
- Preference Instability in Reward Models: Detection and Mitigation via Sparse Autoencoders
- Natural Emergent Misalignment from Reward Hacking in Production RL
- Training language models to follow instructions with human feedback
- Convergent Linear Representations of Emergent Misalignment
- Autoregressive Model Beats Diffusion: Llama for Scalable Image Generation
- Inoculation Prompting: Eliciting traits from LLMs during training can suppress them at test-time
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks