LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation

summary

Video file (mp4)

The gist

The gist: LINGUDISTILL proposes an adapter-free distillation method that restores native linguistic capability in Vision-Language Models by leveraging a frozen LM teacher through layer-wise KV-cache

In short

LINGUDISTILL is an adapter-free method to restore native language skills in vision-language models by using a frozen language model as a teacher. It achieves this through sharing key memory caches between the student and teacher decoders and selectively distilling knowledge based on data type. This results in about 10% performance recovery on language tasks while keeping vision performance relatively stable.

Key concepts

KV-cache Sharing
This mechanism allows the teacher decoder to reuse the memory (KV-cache) generated by the student decoder. The student processes multimodal input, creating this memory, and then the teacher uses its queries to attend over this shared memory. This effectively conditions the text-only teacher on the visual context processed by both models.
Selective Distillation Objective
The training objective mixes soft distillation from a frozen language model (teacher) with hard supervision from the student's own next-token predictions. The key is 'selective'—the distillation signal is routed based on data type, prioritizing language-relevant data for better linguistic recovery while relying more on hard visual supervision.
Adapter-Free Distillation
Unlike previous methods that add extra modules or parameters, LINGUDISTILL restores language capability without any additional architectural components. It achieves this by cleverly leveraging the existing structure of the VLM and the frozen LM teacher through memory sharing and targeted training signals, resulting in zero new parameters at inference time.

Terminology used across episodes

This episode discusses

The paper

LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation · Read on arXiv

Patrick Amadeus Irawan, Erland Hilman Fuadi, Shanu Kumar, Alham Fikri Aji, Yova Kementchedjhieva

Mohamed bin Zayed University of Artificial Intelligence

Turning a pretrained language model (LM) into a vision-language model (VLM) through multimodal fine-tuning often erodes its native language ability, a form of catastrophic forgetting that shows up even on text-only tasks. This loss is hard to undo with further fine-tuning, and existing remedies add adapters or alignment modules that increase architectural complexity and inference cost. We propose LinguDistill, an adapter-free knowledge distillation method that uses the original frozen LM as the teacher during multimodal post-training. To let a text-only teacher supervise vision-conditioned outputs, we introduce layer-wise KV-cache sharing, which exposes the teacher to the student's multimodal representations without changing either architecture. We then apply distillation selectively, on language-heavy data only, so the teacher restores linguistic ability while the student keeps its visual grounding on document and OCR tasks. LinguDistill recovers the language and knowledge performance lost during multimodal fine-tuning, matching the original VLM on average over text-only benchmarks (ARC, HellaSwag) and exceeding it on ScienceQA, while keeping vision-heavy performance close to standard fine-tuning. Since the teacher is dropped after training, the final model adds no parameters and no inference cost. More broadly, our results show that a model's own pre-adaptation backbone is a practical teacher for undoing forgetting, suggesting a simple recipe for keeping language ability intact as models are extended to new modalities.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation".

Jane: The gist: LINGUDISTILL proposes an adapter-free distillation method that restores native linguistic capability in Vision-Language Models by leveraging a frozen LM teacher through layer-wise KV-cache sharing and selective distillation,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We talked about how this paper, titled LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation, deals with the problem of lost language skills when building VLMs.

Jane: The authors are Patrick Amadeus Irawan, Erland Hilman Fuadi, Shanu Kumar, Alham Fikri Aji, and Yova Kementchedjhieva. They focused on how to recover those native language capabilities in vision-language models after they've been adapted for multimodal tasks.

Lu: The paper sets up a specific training framework where the student VLM decoder communicates with the frozen LM teacher by reusing the student’s representation via layer-wise KV sharing.

Meng: So, when you look at the core mechanism, it seems like they’re trying to bridge that gap between what vision sees and what language understands using this shared memory structure.

Tom: Right. It's about getting the teacher to attend over the student’s multimodal context through those shared KV caches so it can generate better supervision signals for the student model.

Jane: The implications here are significant because it suggests you don't need complex, extra modules to fix this representation shift; you can just use distillation to recover capability using only a frozen language model as the teacher.

Lu: It’s about showing that cross-modal knowledge distillation can work effectively even when the teacher and student operate on different input modalities, which is what this paper calls selective cross-modal KD.

Tom: And they found that this approach results in zero additional parameters being added to the final model after training, which is a big deal for deployment.

Jane: So, before we get into how they actually trained it, keep in mind that the goal is to fix language capability while keeping vision performance comparable to what it was before.

The paper's summary: Tom: So, summarizing LinguDistill, the main point is using a frozen LM teacher and layer-wise KV-cache sharing combined with selective distillation to restore lost language ability in VLMs.

Jane: Basically, they figure out that by letting the teacher attend to the student’s multimodal context via those shared caches, you can give it a better signal for language tasks without needing any new architecture.

Lu: They introduce a specific objective function where the student is optimized using both soft distillation from the frozen LM and hard next-token supervision based on labels.

Meng: That selective distillation part is important because it means they aren't just blindly copying everything; they are routing the distillation signal based on what kind of data you’re looking at.

Tom: They use this data-dependent weighting to prioritize applying the distillation signal where it helps language capability the most, while relying more on hard supervision for tasks that are heavy on vision.

Jane: The summary is that this method recovers about ten percent of the performance lost on language and knowledge benchmarks, which is a solid recovery without sacrificing too much in vision-heavy tasks.

Lu: It shows that distillation can target specific weaknesses—like language understanding—without causing a massive degradation in the areas where the model was already strong, like visual reasoning.

Tom: So, it’s about surgical recovery of language skills while preserving the existing visual competence of the VLM backbone.

The paper's improvements: Tom: The authors suggest a few key improvements to their method that really make it work better than standard fine-tuning approaches.

Jane: They focused on making the KV-cache sharing mechanism robust, ensuring the teacher can effectively utilize that shared memory across all layers of the student decoder.

Lu: One improvement is introducing a selective distillation objective based on source categorization, where language-heavy data gets the full distillation signal and OCR or doc-heavy data only gets cross-entropy loss.

Meng: That sounds like a smart way to handle the trade-off between language gains and visual preservation because it explicitly separates those objectives.

Tom: They also tuned the distillation hyperparameters by looking at training loss curves, finding that one specific setting reached the lowest overall cross-entropy curve during training.

Jane: This tuning shows that there is an optimal balance they can strike between how much language you recover and how much visual performance you sacrifice.

Lu: The main improvement they highlight is achieving the best language gains while maintaining controlled vision degradation, which they attribute specifically to their selective distillation strategy.

Conclusion: Tom: So to wrap up LinguDistill, the core idea is using an adapter-free distillation framework with layer-wise KV-cache sharing and selective distillation to restore native linguistic capability in vision-language models.

Jane: The results show that this approach recovers about ten percent of the performance lost on language and knowledge benchmarks while maintaining comparable performance on vision-heavy tasks without adding any new modules or parameters.

Lu: This work demonstrates that targeted cross-modal distillation can be a simple and effective way to preserve backbone capabilities in multimodal systems by leveraging a frozen LM teacher.

Meng: For practical applications, this means you can get an AI system that's better at understanding text prompts without needing a whole new, complex pipeline just for language enhancement.

Tom: It’s about preserving the original structure of the VLM while intelligently boosting its knowledge through distillation, which is a really clean approach.

Jane: So we’ve seen how they use selective KD to maximize language gains while controlling the degradation in vision-heavy areas. That's LinguDistill for now.

More episodes

← Home