Evolving Safety Landscape of Multi-Modal Large Language Models: A Survey of Emerging Threats and Safeguards

summary

Video file (mp4)

The gist

This paper presents a systematic survey of the evolving safety landscape of multi-modal large language models (MLLMs).

In short

This episode surveys emerging safety threats for multi-modal large language models—systems that process text, images, and audio. Hosts discuss how these systems face new risks, such as compromising an image to trick the model, and propose advanced defenses that monitor internal model states rather than just inputs.

Key concepts

Multi-Modal Large Language Models
These are AI systems capable of understanding more than just text. They can process and interpret multiple types of data simultaneously, such as images, audio recordings, and video content.
Compromised Modality Integration
A new threat where an attacker exploits one input type (like an image) to introduce a harmful instruction or error that propagates through the entire multi-modal system.
Modality Misalignment
This occurs when an attacker manipulates the internal representation of data so that the model incorrectly links safe inputs with dangerous content, causing misinterpretation at a deeper level.
Internal Safety Intervention
A defense strategy that monitors what is happening inside the AI model—at its alignment and fusion stages—rather than just checking the initial input or final output.

Terminology used across episodes

This episode discusses

The paper

Evolving Safety Landscape of Multi-modal Large Language Models: A Survey of Emerging Threats and Safeguards · Read on arXiv

Xi Li, Shu Zhao, Xiaohan Zou, Fei Zhao, Fuxiao Liu, Yusen Zhang, Cheng Han, Yushun Dong, Jiaqi Wang

University of Alabama at Birmingham · NVIDIA · Penn State University · Columbia University · University of Missouri-Kansas City · Florida State University · Auburn University

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Evolving Safety Landscape of Multi-Modal Large Language Models: A Survey of Emerging Threats and Safeguards".

Jane: The paper was written by Xi Li, Shu Zhao, Xiaohan Zou, Fei Zhao, Fuxiao Liu et al. from University of Alabama at Birmingham and NVIDIA and Penn State University and Columbia University and University of Missouri-Kansas City and Florida State University and Auburn University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everybody. Today we’re digging into a big one: “Evolving Safety Landscape of Multi-Modal Large Language Models: A Survey of Emerging Threats and Safeguards.” Jane, I’ve got to say, the title alone tells you we’re not in Kansas anymore.

Jane: Absolutely, Tom. And for our listeners who might be new to this, multi-modal large language models are systems that don’t just read text—they also look at images, listen to audio, even watch video. Think of a model that can look at a photo of a dog and describe it, or listen to a voice note and summarize it. That’s the kind of system this paper is about.

Tom: Right, and the word “evolving” in the title is doing a lot of heavy lifting. The authors are saying that the safety playbook we used for text-only models just doesn’t cut it anymore. When you add images and audio, you’re adding whole new ways for things to go wrong.

Jane: Exactly. And the authors come from a bunch of universities—Alabama at Birmingham, Penn State, Columbia, Florida State, Auburn. It’s a solid group of researchers who’ve clearly been watching this space closely.

Tom: So what’s the big deal? Why can’t we just bolt on the old safety measures?

Jane: Well, think of it like this. With text-only models, an attacker might try to trick the model with a cleverly worded prompt. But with multi-modal models, an attacker could hide a harmful instruction inside an image that looks totally innocent. The model reads the image, gets the hidden message, and acts on it—even though the text prompt was perfectly safe.

Tom: That’s wild. So the image becomes a backdoor into the model’s brain.

Jane: Exactly. And that’s just one example. The paper lays out a whole taxonomy of these new threats. They call them compromised modality integration, modality misalignment, and fused safety risks. We’ll get into those in a bit, but the key point is that the attack surface has expanded dramatically.

Tom: And that means the defenses have to change too. You can’t just check the text input and call it a day.

Jane: Right. The paper argues that we need safety solutions that are aware of the multi-modal nature of these systems. That means monitoring what’s happening inside the model, not just at the input and output. It’s a whole new mindset.

Tom: So this isn’t just a survey that lists attacks and defenses. It’s really about how our thinking about safety has to shift.

Jane: That’s the core contribution. They’re not just cataloging problems; they’re proposing a new framework for understanding them. And that framework is what we’re going to unpack over the next few segments.

Tom: Can’t wait. So stick around, because next we’re going to talk about the summary of the paper and what the authors see as the biggest shifts in threat modeling. This is going to get interesting.

Summary: Tom: Alright, we’re back with “Evolving Safety Landscape of Multi-Modal Large Language Models: A Survey of Emerging Threats and Safeguards.” Jane, we talked about the title, but now let’s get into the meat of the summary. What are the authors really saying?

Jane: So the paper’s summary really hammers home three big shifts. First, the attack surface has expanded. In text-only models, you’re mostly worried about the words going in. But now, with images and audio, an attacker can compromise just one modality—say, the image encoder—and that’s enough to mess up the whole system.

Tom: So you don’t need to hack the entire model. You just need to find the weakest link.

Jane: Precisely. They call this “relaxed capability.” The attacker doesn’t need full access anymore. Second, there’s the idea of cross-modal threats. An image might be used to trigger a jailbreak, even when the text prompt is perfectly safe. The danger comes from the interaction between modalities.

Tom: And that’s something we never had to worry about before. In a text-only model, there’s no image to hide a secret message in.

Jane: Right. And third, they talk about how the fusion process itself can be exploited. That’s where the model combines information from different modalities to make a decision. An attacker could craft signals that are harmless on their own, but when combined, they trigger something dangerous.

Tom: Like two chemicals that are safe separately but explosive when mixed.

Jane: Exactly that. And the paper gives concrete examples. There’s a backdoor attack where both an image and a text prompt are needed to trigger the malicious behavior. Each one looks innocent, but together they’re dangerous.

Tom: So the summary is really about how the threat model has changed. It’s not just about adding more attack methods to a list. It’s about a fundamental shift in how we think about what can go wrong.

Jane: And that’s what makes this survey different from others. They’re not just saying “here are fifty attacks.” They’re saying “here’s why the old way of thinking about attacks doesn’t work anymore, and here’s a new way to categorize them.”

Tom: That’s a big deal. Because if you don’t have the right mental model, you’re going to miss the threats that actually matter.

Jane: Exactly. And that’s why they propose that new taxonomy—compromised modality integration, modality misalignment, and fused safety risks. We touched on those earlier, but in the next segment, we’re going to dig into the first page of the paper where they lay out this framework in detail.

Tom: Sounds good. So for now, the takeaway is that multi-modal models have fundamentally changed the game, and we need to rethink our defenses from the ground up. Stay with us.

Improvements: Tom: Welcome back. We’re still on “Evolving Safety Landscape of Multi-Modal Large Language Models: A Survey of Emerging Threats and Safeguards.” Jane, we’ve talked about the threats, but the paper also suggests some improvements. What are they proposing?

Jane: Great question. The paper doesn’t just stop at identifying problems. They also lay out a roadmap for better safety solutions. And they organize these around three new constraints that any good defense has to satisfy.

Tom: Okay, let’s hear them.

Jane: First, there’s “modality-agnostic readiness.” That means a defense can’t assume the attack will come through a specific modality. You can’t just protect the text input and ignore the image input, because attackers will find that gap.

Tom: So you need a defense that’s ready for anything, regardless of which modality gets hit.

Jane: Exactly. Second, there’s “cross-modal safety.” This is about making sure that when you defend one modality, you don’t break the interaction with another. If you filter out too much from the image, the model might lose important context that it needs to understand the text.

Tom: So it’s a balancing act. You want to be safe, but you also want the model to still work well.

Jane: Right. And third, there’s “internal safety intervention.” This is the big one. Instead of just checking inputs and outputs, you need to monitor what’s happening inside the model—at the alignment and fusion stages. That’s where a lot of these new threats actually take effect.

Tom: That sounds harder. You’re not just looking at the surface; you’re looking at the model’s internal state.

Jane: It is harder, but it’s necessary. And the paper reviews several approaches that do this. For example, some methods look at the embedding space—the internal representation of the input—to detect when something is off. If an image’s embedding is too far from what the text suggests, that’s a red flag.

Tom: So they’re using the model’s own internal consistency as a safety check.

Jane: Exactly. And then there are training-based approaches, like safety fine-tuning, where you train the model on safe and unsafe examples to make it more robust. There’s also preference-based optimization, where you teach the model to prefer safe responses over unsafe ones.

Tom: And what about the training-free stuff? I remember seeing something about that in the paper.

Jane: Yes, those are inference-time solutions. You don’t retrain the model; you just add a safety layer at the input, during processing, or at the output. For example, you might have a separate model that checks the response before it’s sent to the user.

Tom: So there’s a whole toolkit of improvements. The paper isn’t just saying “be more careful.” It’s giving concrete strategies.

Jane: Exactly. And the key is that these strategies are designed with the multi-modal nature in mind. They’re not just patching up old text-only defenses. They’re built for the new reality.

Tom: That’s a solid foundation. Next, we’re going to look at the actual first page of the paper to see how they set up this whole argument. Stay tuned.

First Page: Tom: Alright, we’re back on “Evolving Safety Landscape of Multi-Modal Large Language Models: A Survey of Emerging Threats and Safeguards.” Jane, let’s get into the nitty-gritty of the first page. What do the authors set up right from the start?

Jane: So the first page is really about laying the foundation. They start by explaining what multi-modal learning is and why it’s so powerful. You’ve got models like Flamingo, BLIP-two and LLaVA that can understand both images and text. That’s a huge step forward in capability.

Tom: But with great power comes great responsibility, right?

Jane: Exactly. And they immediately pivot to the safety challenges. They point out that adding more modalities doesn’t just add more features—it adds more ways for things to go wrong. They list three specific problems: new vulnerabilities in each modality, semantic inconsistencies between modalities, and the fusion process itself being exploitable.

Tom: And that’s where they introduce the idea of “compromised modality integration.” Can you break that down for us?

Jane: Sure. Imagine you have a model that looks at an image and reads a text prompt. The image goes through a visual encoder, and the text goes through a text encoder. Then those two streams are combined in a fusion block. If an attacker can mess with the image encoder—say, by adding a tiny perturbation that’s invisible to the human eye—that error propagates through the whole system.

Tom: So a tiny nudge in the image can cause the model to say something completely wrong.

Jane: Exactly. The paper gives examples where a small perturbation can make the model describe a bomb-making instruction instead of a harmless scene. And the scary part is that the perturbation is so subtle, a human wouldn’t even notice it.

Tom: That’s the “compromised modality integration” category. What about the other two?

Jane: The second is “modality misalignment.” This is when the attacker messes with the alignment between the image and the text. For example, they might make the image’s embedding look like the embedding of a harmful text prompt. So the model thinks it’s seeing something safe, but the internal representation is actually pointing toward dangerous content.

Tom: So it’s like a disguise at the representation level.

Jane: Exactly. And the third category is “fused safety risks.” This is where the danger only appears when the modalities are combined. Each modality on its own is fine, but together they trigger something harmful. The paper mentions backdoor attacks where you need both a specific image and a specific text prompt to activate the malicious behavior.

Tom: So the first page really sets up this whole framework. It’s not just a list of attacks; it’s a way of thinking about where attacks can happen.

Jane: Right. And they also start talking about how this changes the threat model. Attackers don’t need full access anymore. They just need to compromise one part of the pipeline. That’s a fundamental shift.

Tom: And that shift is what makes this paper so important. It’s not just about the attacks themselves; it’s about how we need to rethink safety from the ground up.

Jane: Exactly. And in the next segment, we’re going to wrap up by talking about the future directions they propose. That’s where things get really exciting.

Conclusion: Tom: Alright, we’ve reached the final segment of our discussion on “Evolving Safety Landscape of Multi-Modal Large Language Models: A Survey of Emerging Threats and Safeguards.” Jane, let’s bring it all together.

Jane: So the paper really gives us a new lens for looking at multi-modal safety. Instead of just listing attacks, they’ve organized them into three categories: compromised modality integration, modality misalignment, and fused safety risks. That gives researchers a clear framework for understanding where the dangers are.

Tom: And they’ve also shown that the threat model has shifted. Attackers don’t need full access anymore. They can target just one modality or exploit the interaction between modalities.

Jane: Right. And that means our defenses have to be smarter. They need to be modality-agnostic, they need to preserve cross-modal coherence, and they need to intervene at the internal stages of the model, not just at the input and output.

Tom: The paper also points toward some exciting future directions. They talk about building defenses that are resilient to partial corruption, and they even suggest moving toward agent-driven red-blue teaming, where AI systems automatically find vulnerabilities and patch them.

Jane: That last one is really interesting. Instead of relying on humans to manually discover every possible attack, you could have an AI system that constantly probes the model for weaknesses and then automatically develops defenses. That could be a game-changer for keeping up with the ever-evolving threat landscape.

Tom: So what’s the big takeaway for our listeners?

Jane: The big takeaway is that multi-modal AI is incredibly powerful, but that power comes with new and complex risks. This paper gives us a roadmap for understanding those risks and developing better safeguards. It’s a must-read for anyone working in AI safety.

Tom: And with that, we’re wrapping up our discussion on this paper. Jane, it’s been a great conversation.

Jane: It really has, Tom. Thanks to everyone for listening. We’ll be back soon with another paper to break down. Until then, stay curious.

Tom: And stay safe. See you next time.

More episodes

← Home