LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation
summary
The gist
The gist: LINGUDISTILL proposes an adapter-free distillation method that restores native linguistic capability in Vision-Language Models by leveraging a frozen LM teacher through layer-wise KV-cache
In short
LINGUDISTILL is an adapter-free method to restore native language skills in vision-language models by using a frozen language model as a teacher. It achieves this through sharing key memory caches between the student and teacher decoders and selectively distilling knowledge based on data type. This results in about 10% performance recovery on language tasks while keeping vision performance relatively stable.
Key concepts
- KV-cache Sharing
- This mechanism allows the teacher decoder to reuse the memory (KV-cache) generated by the student decoder. The student processes multimodal input, creating this memory, and then the teacher uses its queries to attend over this shared memory. This effectively conditions the text-only teacher on the visual context processed by both models.
- Selective Distillation Objective
- The training objective mixes soft distillation from a frozen language model (teacher) with hard supervision from the student's own next-token predictions. The key is 'selective'—the distillation signal is routed based on data type, prioritizing language-relevant data for better linguistic recovery while relying more on hard visual supervision.
- Adapter-Free Distillation
- Unlike previous methods that add extra modules or parameters, LINGUDISTILL restores language capability without any additional architectural components. It achieves this by cleverly leveraging the existing structure of the VLM and the frozen LM teacher through memory sharing and targeted training signals, resulting in zero new parameters at inference time.
Terminology used across episodes
This episode discusses
- LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation · Paper Radio
- SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model
- Qwen3-VL Technical Report
- Microsoft COCO Captions: Data Collection and Evaluation Server
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Distilling the Knowledge in a Neural Network
- UniFusion: Vision-Language Model as Unified Encoder in Image Generation
- Multimodal Alignment and Fusion: A Survey
- OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- IAA: Inner-Adaptor Architecture Empowers Frozen Large Language Model with Multimodal Capabilities
- Cross-Modal Knowledge Distillation for Speech Large Language Models
- Investigating the Catastrophic Forgetting in Multimodal Large Language Models
- T5Gemma 2: Seeing, Reading, and Understanding Longer
- LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention
- Unified Multimodal Understanding and Generation Models: Advances, Challenges, and Opportunities
The paper
LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation · Read on arXiv
Patrick Amadeus Irawan, Erland Hilman Fuadi, Shanu Kumar, Alham Fikri Aji, Yova Kementchedjhieva
Mohamed bin Zayed University of Artificial Intelligence
Turning a pretrained language model (LM) into a vision-language model (VLM) through multimodal fine-tuning often erodes its native language ability, a form of catastrophic forgetting that shows up even on text-only tasks. This loss is hard to undo with further fine-tuning, and existing remedies add adapters or alignment modules that increase architectural complexity and inference cost. We propose LinguDistill, an adapter-free knowledge distillation method that uses the original frozen LM as the teacher during multimodal post-training. To let a text-only teacher supervise vision-conditioned outputs, we introduce layer-wise KV-cache sharing, which exposes the teacher to the student's multimodal representations without changing either architecture. We then apply distillation selectively, on language-heavy data only, so the teacher restores linguistic ability while the student keeps its visual grounding on document and OCR tasks. LinguDistill recovers the language and knowledge performance lost during multimodal fine-tuning, matching the original VLM on average over text-only benchmarks (ARC, HellaSwag) and exceeding it on ScienceQA, while keeping vision-heavy performance close to standard fine-tuning. Since the teacher is dropped after training, the final model adds no parameters and no inference cost. More broadly, our results show that a model's own pre-adaptation backbone is a practical teacher for undoing forgetting, suggesting a simple recipe for keeping language ability intact as models are extended to new modalities.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation".
Jane: The gist: LINGUDISTILL proposes an adapter-free distillation method that restores native linguistic capability in Vision-Language Models by leveraging a frozen LM teacher through layer-wise KV-cache sharing and selective distillation,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: We talked about how this paper, titled LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation, deals with the problem of lost language skills when building VLMs.
Jane: The authors are Patrick Amadeus Irawan, Erland Hilman Fuadi, Shanu Kumar, Alham Fikri Aji, and Yova Kementchedjhieva. They focused on how to recover those native language capabilities in vision-language models after they've been adapted for multimodal tasks.
Lu: The paper sets up a specific training framework where the student VLM decoder communicates with the frozen LM teacher by reusing the student’s representation via layer-wise KV sharing.
Meng: So, when you look at the core mechanism, it seems like they’re trying to bridge that gap between what vision sees and what language understands using this shared memory structure.
Tom: Right. It's about getting the teacher to attend over the student’s multimodal context through those shared KV caches so it can generate better supervision signals for the student model.
Jane: The implications here are significant because it suggests you don't need complex, extra modules to fix this representation shift; you can just use distillation to recover capability using only a frozen language model as the teacher.
Lu: It’s about showing that cross-modal knowledge distillation can work effectively even when the teacher and student operate on different input modalities, which is what this paper calls selective cross-modal KD.
Tom: And they found that this approach results in zero additional parameters being added to the final model after training, which is a big deal for deployment.
Jane: So, before we get into how they actually trained it, keep in mind that the goal is to fix language capability while keeping vision performance comparable to what it was before.
The paper's summary: Tom: So, summarizing LinguDistill, the main point is using a frozen LM teacher and layer-wise KV-cache sharing combined with selective distillation to restore lost language ability in VLMs.
Jane: Basically, they figure out that by letting the teacher attend to the student’s multimodal context via those shared caches, you can give it a better signal for language tasks without needing any new architecture.
Lu: They introduce a specific objective function where the student is optimized using both soft distillation from the frozen LM and hard next-token supervision based on labels.
Meng: That selective distillation part is important because it means they aren't just blindly copying everything; they are routing the distillation signal based on what kind of data you’re looking at.
Tom: They use this data-dependent weighting to prioritize applying the distillation signal where it helps language capability the most, while relying more on hard supervision for tasks that are heavy on vision.
Jane: The summary is that this method recovers about ten percent of the performance lost on language and knowledge benchmarks, which is a solid recovery without sacrificing too much in vision-heavy tasks.
Lu: It shows that distillation can target specific weaknesses—like language understanding—without causing a massive degradation in the areas where the model was already strong, like visual reasoning.
Tom: So, it’s about surgical recovery of language skills while preserving the existing visual competence of the VLM backbone.
The paper's improvements: Tom: The authors suggest a few key improvements to their method that really make it work better than standard fine-tuning approaches.
Jane: They focused on making the KV-cache sharing mechanism robust, ensuring the teacher can effectively utilize that shared memory across all layers of the student decoder.
Lu: One improvement is introducing a selective distillation objective based on source categorization, where language-heavy data gets the full distillation signal and OCR or doc-heavy data only gets cross-entropy loss.
Meng: That sounds like a smart way to handle the trade-off between language gains and visual preservation because it explicitly separates those objectives.
Tom: They also tuned the distillation hyperparameters by looking at training loss curves, finding that one specific setting reached the lowest overall cross-entropy curve during training.
Jane: This tuning shows that there is an optimal balance they can strike between how much language you recover and how much visual performance you sacrifice.
Lu: The main improvement they highlight is achieving the best language gains while maintaining controlled vision degradation, which they attribute specifically to their selective distillation strategy.
Conclusion: Tom: So to wrap up LinguDistill, the core idea is using an adapter-free distillation framework with layer-wise KV-cache sharing and selective distillation to restore native linguistic capability in vision-language models.
Jane: The results show that this approach recovers about ten percent of the performance lost on language and knowledge benchmarks while maintaining comparable performance on vision-heavy tasks without adding any new modules or parameters.
Lu: This work demonstrates that targeted cross-modal distillation can be a simple and effective way to preserve backbone capabilities in multimodal systems by leveraging a frozen LM teacher.
Meng: For practical applications, this means you can get an AI system that's better at understanding text prompts without needing a whole new, complex pipeline just for language enhancement.
Tom: It’s about preserving the original structure of the VLM while intelligently boosting its knowledge through distillation, which is a really clean approach.
Jane: So we’ve seen how they use selective KD to maximize language gains while controlling the degradation in vision-heavy areas. That's LinguDistill for now.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck