TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability

arXiv:2405.17678 · cs.CV, cs.AI · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability".

Jane: The paper was written by Fengji Ma, Hei Victor Cheng, Chenxing Li and Li Liu from The Hong Kong University of Science and Technology (Guangzhou) and Aarhus University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone! Today we're looking at a fresh arXiv paper that's got a real mouthful of a title: "TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability." Jane, I'm going to need you to unpack that for me.

Jane: Happy to, Tom! So the paper is about CLIP, which is that big AI model that can look at an image and match it to text descriptions, like "a photo of a dog." It's amazing at recognizing things it has never been trained on before, which is what "zero-shot" means. But there's a catch: you can trick it by adding tiny, invisible changes to an image, and suddenly it thinks a dog is a cat. Those tricks are called adversarial attacks.

Tom: Right, and that's the "adversarial robustness" part. The model needs to be tough against those attacks. But here's the tricky bit that this paper tackles: if you make the model tough against attacks, it often gets worse at its normal job of recognizing clean images. That's the "generalization ability" part. It's a seesaw.

Jane: Exactly. And the authors, Fengji Ma, Li Liu, and Hei Victor Cheng, they're saying that previous methods could balance that seesaw pretty well when the attacks were small. But when you crank up the attack strength, the whole thing falls apart. The model either becomes fragile again, or it loses its ability to recognize normal images.

Tom: So they're not just trying to make it a little better; they're trying to fix that balance for big, strong attacks. And their idea is called "Text-Image Mutual Awareness." Sounds like the model needs to pay attention to both its text side and its image side at the same time.

Jane: That's the gist. They're saying that to be robust against big attacks, you need to push the different classes further apart in the model's "mind." Think of it like parking spaces. If the spaces are too close together, a little bump can push a car into the wrong spot. But if you make the spaces huge, even a big bump won't cause a mix-up.

Tom: I like that analogy. So they're widening the parking spaces for both the text descriptions and the images themselves. And they're doing it in a way that doesn't mess up the original meaning of the classes. That's the "mutual awareness" part—the text and image sides are helping each other stay in their lanes.

Jane: Right. And the results they show in the paper are pretty impressive. They tested it on a bunch of different datasets, and their method, TIMA, beats the previous best methods on both clean accuracy and robustness against attacks, especially the big ones. We're talking about a big jump in performance when the attack is strong enough to be a real problem.

Tom: So this isn't just a minor tweak. This could be a real step forward for making these powerful models actually safe to use in the real world. I'm curious to see how they actually built this thing. Let's get into the details in the next segment.

Summary: Jane: So, Tom, we've established that TIMA is about making CLIP robust to attacks without losing its smarts. But how do they actually do it? The paper has two main parts, and they call them tuning mechanisms. The first one is for the text side, called Image-Aware Text tuning.

Tom: Okay, so they're adjusting the text embeddings, which are basically the model's internal representation of the words "dog," "cat," "fox," and so on. And they want to push those representations further apart. But they're using this thing called Minimum Hyperspherical Energy to do it. What does that mean in plain English?

Jane: It's a fancy way of saying they want to spread the text points out as evenly as possible on a sphere. Imagine putting stickers on a ball. You want them all as far away from each other as possible, so they're not clumped up. That maximizes the distance between every class.

Tom: But here's the thing—if you just spread them out randomly, you might lose the fact that "dog" and "fox" are more similar to each other than "dog" and "airplane." The paper says previous methods, like LAAT, made that mistake. They got the distance but lost the meaning.

Jane: Exactly. And that's where their "Image-Aware" part comes in. They use the original, unmodified CLIP model's image embeddings as a teacher. They make sure that the new, spread-out text embeddings still have the same relationship with the images as the old ones did. It's like they're saying, "You can move the text points, but you have to keep the same connections to the images."

Tom: So the images are the anchor that keeps the text from going crazy. That's clever. Now, what about the second part? They also have a Text-Aware Image tuning mechanism.

Jane: This is the part that I think is really novel. Previous work only adjusted the text side. But this paper says, "Hey, the image side needs work too." When you attack a model, the image embeddings get all bunched up together, making it hard to tell classes apart. So they add an adaptive margin to the training loss.

Tom: An adaptive margin. So it's like adding a penalty to the model when it gets confused between similar classes. And it's "adaptive" because the penalty is bigger for classes that are semantically closer, like "dog" and "fox," and smaller for totally different classes.

Jane: Precisely. And again, they use a knowledge distillation trick to keep the fine-tuned image embeddings from losing their connection to the original text embeddings. So both sides are being pushed apart, but they're both being held accountable to the original model's knowledge.

Tom: So it's a two-way street. The text helps the images, and the images help the text. That's the "mutual awareness" in the title. It's a really elegant way to think about it. Let's see how well it actually works in practice in the next segment.

Improvements: Tom: Alright, so we know the recipe. But does it actually cook a good meal? The paper has a ton of experiments, and I want to get into the numbers. Jane, what's the headline result?

Jane: The headline is that TIMA gets the best of both worlds. When they fine-tune on ImageNet, their method gets a clean accuracy of about sixty-four point three seven percent on average across ten datasets, which is higher than the previous best methods like TeCoA and PMG. But more importantly, under a strong PGD attack with a small perturbation, they get a robust accuracy of forty-six point four one percent, which is a huge jump from the six point five one percent that the original CLIP gets.

Tom: Whoa, that's a massive improvement. But what about those big perturbations we talked about? That's where the other methods were failing.

Jane: That's the real win. When they fine-tune on Tiny-ImageNet and test with a much larger perturbation, like eight/two hundred fifty-five TIMA gets an average robust accuracy of twenty-two point six four percent. The next best method, LAAT, only gets eleven point five eight percent. And TeCoA is way down at seven point eight seven percent. So they're more than doubling the robustness of the best existing method under those strong attacks.

Tom: So it's not just a small edge. It's a clear, decisive improvement in the exact scenario that was broken before. And they didn't sacrifice clean accuracy to get it. That's the balance everyone's been chasing.

Jane: Right. And they also show that their Text-Aware Image tuning mechanism is a plug-and-play component. They added it to TeCoA and LAAT, and it improved both their robust and clean accuracy. That's a strong sign that their idea about the image embeddings is fundamentally sound and can help other methods too.

Tom: That's a great point. It's not just a one-off trick; it's a building block. They also ran ablations to show that both the text and image tuning are necessary. If you only do one, you don't get the full benefit. The magic is in the combination.

Jane: Exactly. And they even tested with a different temperature setting for CLIP, and TIMA still came out on top. So it's not overfitting to one specific configuration. It's a robust solution.

Tom: This is really promising. It feels like they've cracked a code that a lot of people were stuck on. But I'm wondering about the bigger picture. What does this mean for actually using these models? Let's bring in the rest of the team for that.

Conclusion: Tom: Alright, let's wrap this up. We've been deep in the weeds on "TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability." Jane, give us the final takeaway.

Jane: The final takeaway is that TIMA shows you can have your cake and eat it too. You can make a CLIP model that is both highly accurate on clean images and highly resistant to even strong adversarial attacks. The secret is to increase the distance between classes in both the text and image spaces, while using knowledge distillation to keep the semantic meaning intact.

Tom: And that's a big deal. Lu, from your research perspective, what's the most exciting implication here?

Lu: For me, it's the idea that we can now think about robustness not as a trade-off, but as a design goal. This paper gives us a clear, principled way to achieve it. It opens up the possibility of deploying these powerful zero-shot models in safety-critical applications, like autonomous driving or medical imaging, where a single misclassification could be disastrous.

Meng: From an engineering standpoint, I'm also encouraged. The fact that the Text-Aware Image part is plug-and-play is huge. It means we can potentially bolt this onto existing fine-tuning pipelines without a complete rewrite. That lowers the barrier to adoption significantly.

Lalam: And if we look at the cultural impact, this is about trust. For AI to be integrated into our daily lives, people need to trust that it won't be fooled by a slightly altered photo. TIMA is a step toward building that trust, making AI not just more capable, but more reliable and dependable in the real world.

Tom: Well said, everyone. So we're saying goodbye to TIMA, but we're definitely keeping an eye on this line of research. It feels like a real milestone. Thanks for joining us, and we'll see you for the next paper.

Jane: Bye, everyone!

Fengji Ma, Hei Victor Cheng, Chenxing Li, Li Liu

The Hong Kong University of Science and Technology (Guangzhou) · Aarhus University

cs.CV, cs.AI

Submitted: 2026-08-17

Updated: 2026-08-18

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 60/100

The gist: TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability This paper addresses the challenge of achieving zero-shot adversarial robustness while

Key concepts

Zero-Shot Adversarial Robustness
This refers to a model's ability to remain accurate when faced with adversarial attacks, which are tiny, invisible changes made to an image designed to trick the AI into misclassifying it. The paper addresses how making a model robust against these attacks can negatively affect its normal performance on clean images.
Text-Image Mutual Awareness
This is the core idea where both the text and image sides of a model pay attention to each other. The goal is to push different classes further apart in both the text and image representations, ensuring they are well-separated without losing their original semantic meaning.
Minimum Hyperspherical Energy
This is a technique used for Image-Aware Text tuning. It spreads the text embeddings out as evenly as possible on a sphere, maximizing the distance between class representations. This helps separate classes while maintaining their relationships with the original images.

Terminology

Summary

TIMA: Text-Image Mutual Awareness for Balancing Zero-Shot Adversarial Robustness and Generalization Ability

This paper addresses the challenge of achieving zero-shot adversarial robustness while preserving zero-shot generalization in large-scale foundation models, with a focus on the Contrastive Language-Image Pre-training (CLIP) model. The authors note that while foundation models have exceptional zero-shot generalization, they are highly vulnerable to adversarial perturbations. Existing methods achieve a comparable good tradeoff between zero-shot adversarial robustness and generalization under small adversarial perturbations, but fail to do so under large adversarial perturbations.

To address this, the paper proposes a novel Text-Image Mutual Awareness (TIMA) method that strikes a balance between zero-shot adversarial robustness and generalization. The method consists of two main tuning mechanisms:

  1. Image-Aware Text (IAT) tuning mechanism: This mechanism increases the inter-class distance of text embeddings by incorporating the Minimum Hyperspherical Energy (MHE) principle. Simultaneously, fixed pre-trained image embeddings are used as cross-modal auxiliary supervision to maintain the similarity between the MHE-tuned and original text embeddings through knowledge distillation, preserving semantic information between different classes.

  2. Text-Aware Image (TAI) tuning mechanism: This mechanism increases inter-class distance between image embeddings during the training stage by using Text-distance based Adaptive Margin (TAM). Similarly, a knowledge distillation is utilized to retain the similarity between fine-tuned and pre-trained image embeddings.

The central hypothesis of TIMA is that Increasing the inter-class distances within the pretrained CLIP text and image embeddings is the key to improving zero-shot adversarial robustness, especially under large perturbation.

The main contributions are summarized as follows:

  • The proposed MHE method finds the largest inter-class distances for text embeddings, making CLIP more resilient to both small and large perturbations.

  • This is the first work to propose a working mechanism for enlarging the inter-class distance in image embeddings. The proposed TAI can be applied in a plug-and-play manner to combine with any text embedding adjustment schemes.

  • The authors design novel knowledge distillation methods for text and image embeddings to keep semantic and interactive information between the modalities after enlarging inter-class distances.

  • Extensive experiments show the proposed TIMA's superiority in zero-shot robust accuracy and clean accuracy across multiple datasets and under different perturbations.

The method is evaluated against state-of-the-art methods TeCoA, PMG, and LAAT, using adversarial fine-tuning on both ImageNet and Tiny-ImageNet. The evaluation uses multiple zero-shot test datasets including CIFAR10, CIFAR100, STL10, OxfordPets, Food101, SUN397, DTD, and EuroSAT.

Key experimental results show that when adversarially fine-tuned on ImageNet, TIMA achieves a 39.9% improvement in performance under PGD-10 attacks (from 6.51% to 46.41%) and a 43.7% enhancement against AutoAttack (from 0.84% to 44.54%) compared to the original CLIP, while only experiencing a marginal 3.01% decline in zero-shot clean accuracy. Compared to existing SOTA methods, TIMA improves zero-shot clean accuracy by 4.52% over TeCoA, 6.06% over PMG, and 12.77% over LAAT. Under PGD attack, TIMA's zero-shot robust accuracy exceeds TeCoA by 4.48%, PMG by 2.64%, and LAAT by 10.98%.

Under large perturbation radius (epsilon = 8/255), TIMA's zero-shot robust accuracy improved by 14.77% compared to TeCoA and by 11.06% relative to LAAT. The method also remains efficacious across various CLIP temperature settings, demonstrating versatility.

The ablation study confirms that both tuning mechanisms contribute to the performance, with IAT markedly improving adversarial robustness under various perturbations, and TAI being a plug-and-play tuning mechanism that can be introduced into other methods like TeCoA and LAAT, leading to increased both zero-shot adversarial robustness and generalization. The results demonstrate that increasing both text and image embedding inter-class distances improves zero-shot adversarial robustness, particularly under large perturbations, providing evidence for the paper's hypothesis.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, along with what the improved system can do:


1. Dual-Modality Inter-Class Distance Enlargement

  • Implementation: Add a training module that simultaneously increases the Euclidean distance between text embeddings (using Minimum Hyperspherical Energy) and image embeddings (using Text-distance Adaptive Margin) during adversarial fine-tuning.

  • Benefit: The model becomes significantly more robust to large adversarial perturbations (up to 8/255) without sacrificing clean accuracy, unlike existing methods that only adjust one modality.

2. Cross-Modal Knowledge Distillation for Semantic Preservation

  • Implementation: Integrate two distillation losses—Image-Aware Knowledge Distillation (IAKD) for text embeddings and Text-Aware Knowledge Distillation (TAKD) for image embeddings—using frozen pre-trained CLIP embeddings as teachers.

  • Benefit: The model retains the original semantic relationships between classes and the contrastive alignment between text and image, preventing the catastrophic drop in zero-shot generalization seen in methods like LAAT.

3. Adaptive Margin Based on Text Semantics

  • Implementation: Replace fixed margins with a dynamic margin that scales with the cosine similarity between the correct class text embedding and the incorrect class text embedding. Only apply the margin when the incorrect class is semantically close (above a threshold).

  • Benefit: The model avoids over-penalizing dissimilar classes, reducing misclassifications among semantically similar classes (e.g., dog vs. fox) under adversarial attack.

4. Plug-and-Play Image Tuning Module

  • Implementation: Design the Text-Aware Image (TAI) tuning mechanism as a standalone module that can be added to any existing text-embedding adjustment method (e.g., TeCoA, LAAT).

  • Benefit: Researchers and practitioners can easily integrate this module into their existing pipelines to boost adversarial robustness without redesigning their entire training framework.

5. Temperature-Invariant Robustness

  • Implementation: Ensure the training objective is stable across different CLIP temperature values (e.g., τ = 0.01 and τ = 1.0) by using the same distillation and margin formulation.

  • Benefit: The model maintains high zero-shot robust accuracy and clean accuracy regardless of the temperature setting, making it more reliable in deployment scenarios where temperature may be tuned for other purposes.

  • Achieve state-of-the-art zero-shot adversarial robustness under both small (1/255) and large (4/255, 8/255) perturbation budgets, with improvements of up to 14.77% over existing methods under 8/255 attacks.

  • Preserve zero-shot generalization on clean images, with only a 3% drop from the original CLIP model, while existing methods lose 10–15% or more.

  • Transfer robustness across datasets without retraining—the system fine-tuned on ImageNet or Tiny-ImageNet performs well on unseen datasets like CIFAR-10, OxfordPets, Food101, EuroSAT, and DTD.

  • Maintain semantic consistency between similar classes (e.g., bird vs. tiger) even after adversarial training, preventing the collapse of inter-class relationships.

  • Operate effectively under different CLIP temperatures, making it robust to hyperparameter changes in deployment.

  • Serve as a drop-in enhancement for existing zero-shot adversarial robustness methods, improving both their robustness and generalization without requiring architectural changes.

These improvements directly address the paper’s core hypothesis—that enlarging inter-class distances in both text and image embeddings, while preserving semantic information via cross-modal distillation, is the key to balancing zero-shot adversarial robustness and generalization.

Abstract

This work addresses the challenge of achieving zero-shot adversarial robustness while preserving zero-shot generalization in large-scale foundation models, with a focus on the popular Contrastive Language-Image Pre-training (CLIP). Although foundation models were reported to have exceptional zero-shot generalization, they are highly vulnerable to adversarial perturbations. Existing methods achieve a comparable good tradeoff between zero-shot adversarial robustness and generalization under small adversarial perturbations. However, they fail to achieve a good tradeoff under large adversarial perturbations. To this end, we propose a novel Text-Image Mutual Awareness (TIMA) method that strikes a balance between zero-shot adversarial robustness and generalization. More precisely, we propose an Image-Aware Text (IAT) tuning mechanism that increases the inter-class distance of text embeddings by incorporating the Minimum Hyperspherical Energy (MHE). Simultaneously, fixed pre-trained image embeddings are used as cross-modal auxiliary supervision to maintain the similarity between the MHE-tuned and original text embeddings by the knowledge distillation, preserving semantic information between different classes. Besides, we introduce a Text-Aware Image (TAI) tuning mechanism, which increases inter-class distance between image embeddings during the training stage by Text-distance based Adaptive Margin (TAM). Similarly, a knowledge distillation is utilized to retain the similarity between fine-tuned and pre-trained image embeddings. Extensive experimental results demonstrate the effectiveness of our approach, showing impressive zero-shot performance against a wide range of adversarial perturbations while preserving the zero-shot generalization capabilities of the original CLIP model.

Sources

Related papers