LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation

arXiv:2604.00829 · cs.CV, cs.CL · Submitted 2026-04-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation".

Jane: The gist: LINGUDISTILL proposes an adapter-free distillation method that restores native linguistic capability in Vision-Language Models by leveraging a frozen LM teacher through layer-wise KV-cache sharing and selective distillation,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We talked about how this paper, titled LinguDistill: Recovering Linguistic Ability in Vision-Language Models via Selective Cross-Modal Distillation, deals with the problem of lost language skills when building VLMs.

Jane: The authors are Patrick Amadeus Irawan, Erland Hilman Fuadi, Shanu Kumar, Alham Fikri Aji, and Yova Kementchedjhieva. They focused on how to recover those native language capabilities in vision-language models after they've been adapted for multimodal tasks.

Lu: The paper sets up a specific training framework where the student VLM decoder communicates with the frozen LM teacher by reusing the student’s representation via layer-wise KV sharing.

Meng: So, when you look at the core mechanism, it seems like they’re trying to bridge that gap between what vision sees and what language understands using this shared memory structure.

Tom: Right. It's about getting the teacher to attend over the student’s multimodal context through those shared KV caches so it can generate better supervision signals for the student model.

Jane: The implications here are significant because it suggests you don't need complex, extra modules to fix this representation shift; you can just use distillation to recover capability using only a frozen language model as the teacher.

Lu: It’s about showing that cross-modal knowledge distillation can work effectively even when the teacher and student operate on different input modalities, which is what this paper calls selective cross-modal KD.

Tom: And they found that this approach results in zero additional parameters being added to the final model after training, which is a big deal for deployment.

Jane: So, before we get into how they actually trained it, keep in mind that the goal is to fix language capability while keeping vision performance comparable to what it was before.

The paper's summary: Tom: So, summarizing LinguDistill, the main point is using a frozen LM teacher and layer-wise KV-cache sharing combined with selective distillation to restore lost language ability in VLMs.

Jane: Basically, they figure out that by letting the teacher attend to the student’s multimodal context via those shared caches, you can give it a better signal for language tasks without needing any new architecture.

Lu: They introduce a specific objective function where the student is optimized using both soft distillation from the frozen LM and hard next-token supervision based on labels.

Meng: That selective distillation part is important because it means they aren't just blindly copying everything; they are routing the distillation signal based on what kind of data you’re looking at.

Tom: They use this data-dependent weighting to prioritize applying the distillation signal where it helps language capability the most, while relying more on hard supervision for tasks that are heavy on vision.

Jane: The summary is that this method recovers about ten percent of the performance lost on language and knowledge benchmarks, which is a solid recovery without sacrificing too much in vision-heavy tasks.

Lu: It shows that distillation can target specific weaknesses—like language understanding—without causing a massive degradation in the areas where the model was already strong, like visual reasoning.

Tom: So, it’s about surgical recovery of language skills while preserving the existing visual competence of the VLM backbone.

The paper's improvements: Tom: The authors suggest a few key improvements to their method that really make it work better than standard fine-tuning approaches.

Jane: They focused on making the KV-cache sharing mechanism robust, ensuring the teacher can effectively utilize that shared memory across all layers of the student decoder.

Lu: One improvement is introducing a selective distillation objective based on source categorization, where language-heavy data gets the full distillation signal and OCR or doc-heavy data only gets cross-entropy loss.

Meng: That sounds like a smart way to handle the trade-off between language gains and visual preservation because it explicitly separates those objectives.

Tom: They also tuned the distillation hyperparameters by looking at training loss curves, finding that one specific setting reached the lowest overall cross-entropy curve during training.

Jane: This tuning shows that there is an optimal balance they can strike between how much language you recover and how much visual performance you sacrifice.

Lu: The main improvement they highlight is achieving the best language gains while maintaining controlled vision degradation, which they attribute specifically to their selective distillation strategy.

Conclusion: Tom: So to wrap up LinguDistill, the core idea is using an adapter-free distillation framework with layer-wise KV-cache sharing and selective distillation to restore native linguistic capability in vision-language models.

Jane: The results show that this approach recovers about ten percent of the performance lost on language and knowledge benchmarks while maintaining comparable performance on vision-heavy tasks without adding any new modules or parameters.

Lu: This work demonstrates that targeted cross-modal distillation can be a simple and effective way to preserve backbone capabilities in multimodal systems by leveraging a frozen LM teacher.

Meng: For practical applications, this means you can get an AI system that's better at understanding text prompts without needing a whole new, complex pipeline just for language enhancement.

Tom: It’s about preserving the original structure of the VLM while intelligently boosting its knowledge through distillation, which is a really clean approach.

Jane: So we’ve seen how they use selective KD to maximize language gains while controlling the degradation in vision-heavy areas. That's LinguDistill for now.

Patrick Amadeus Irawan, Erland Hilman Fuadi, Shanu Kumar, Alham Fikri Aji, Yova Kementchedjhieva

Mohamed bin Zayed University of Artificial Intelligence

cs.CV, cs.CL

Submitted: 2026-04-01

Updated: 2026-10-03

Code: https://github.com/huggingface/nanoVLM

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: The gist: LINGUDISTILL proposes an adapter-free distillation method that restores native linguistic capability in Vision-Language Models by leveraging a frozen LM teacher through layer-wise KV-cache

Key concepts

KV-cache Sharing
This mechanism allows the teacher decoder to reuse the memory (KV-cache) generated by the student decoder. The student processes multimodal input, creating this memory, and then the teacher uses its queries to attend over this shared memory. This effectively conditions the text-only teacher on the visual context processed by both models.
Selective Distillation Objective
The training objective mixes soft distillation from a frozen language model (teacher) with hard supervision from the student's own next-token predictions. The key is 'selective'—the distillation signal is routed based on data type, prioritizing language-relevant data for better linguistic recovery while relying more on hard visual supervision.
Adapter-Free Distillation
Unlike previous methods that add extra modules or parameters, LINGUDISTILL restores language capability without any additional architectural components. It achieves this by cleverly leveraging the existing structure of the VLM and the frozen LM teacher through memory sharing and targeted training signals, resulting in zero new parameters at inference time.

Terminology

Summary

The gist: LINGUDISTILL proposes an adapter-free distillation method that restores native linguistic capability in Vision-Language Models by leveraging a frozen LM teacher through layer-wise KV-cache sharing and selective distillation, recovering about 10% of the performance lost on language and knowledge benchmarks while maintaining comparable performance on vision-heavy tasks.

Introduction

Adapting pretrained language models (LMs) into vision-language models (VLMs) can degrade their native linguistic capability due to representation shift and cross-modal interference introduced during multimodal adaptation, such as representation shift and cross-modal interference introduced during multimodal adaptation. Prior recovery approaches typically introduce additional modules that act as intermediate alignment layers to maintain or isolate modality-specific subspaces, which increases architectural complexity, adds parameters at inference time, and limits flexibility across models and settings. LINGUDISTILL is a cross-modal knowledge distillation method to restore language capability in VLMs. A key challenge is that the teacher only operates on text, so its outputs are not aligned with vision-conditioned generation. To address this, we introduce a KV-cache sharing mechanism between the teacher and student decoders. After training, the teacher is removed, resulting in a standard VLM with improved language capability and no additional parameters or inference cost.

KV Sharing Architecture

The KV-cache sharing mechanism proceeds as follows:

  1. The student decoder Φ processes the full multimodal sequence X and produces KV caches: The student decoder Φ processes the full multimodal sequence X and produces KV caches.

  2. The teacher decoder omega reprocesses the same text prompt to form its query states: The teacher decoder omega reprocesses the same text prompt to form its query states.

  3. The teacher omega directly reuses the student KV cache at every layer: The teacher omega directly reuses the student KV cache at every layer.

  4. The teacher attends using its computed queries over the transported student memory: The teacher attends using its computed queries over the transported student memory.

This design allows the teacher to attend to the student’s multimodal context via shared KV caches, effectively conditioning it on multimodal input.

Selective Distillation Objective

The student (Φ) is optimized with a typical KD mixture of frozen LM’s (omega) soft distillation and its own hard next-token supervision. The final objective is L = 1 / omegapos ∑ (b,t)∈omegapos α(db) L(b,t) soft + (1 − α(db)) L(b,t) hard. This objective allows the teacher’s supervision signal to be sourced more from languagerelevant data, while relying more on hard supervision for visual-heavy sources.

Experimental Setup and Results

The experiments use the instruction-tuned nanoVLM-460M-8k model as the VLM (student) and its instruction-tuned LM backbone, SmolLM-360M-Instruct (teacher). The training uses a selective distillation objective described in Sec. 3.2. Selective KD (Ours) resolves the trade-off by routing the distillation signal based on data type. Table 7 reports benchmark scores across all variants. Selective KD achieves the strongest gains coupled with a controlled perception ability degradation. For instance, on language and knowledge-heavy benchmarks, ScienceQA reaches 0.676 (+14.2%).

Discussion

LINGUDISTILL recovers performance on language and knowledge benchmarks by ∼10%, while exhibiting comparable degradation in vision-heavy tasks relative to standard fine-tuning methods. The key insight is that distillation helps whenever a task needs knowledge, regardless of whether it also involves reading text. For OCRBench, the model corrects nonsense strings into real words when using LINGUDISTILL because The teacher’s language signal makes the model better at understanding text but worse at copying exactly what it sees.

Conclusion

LINGUDISTILL proposes an adapter-free distillation framework that restores native linguistic capability by leveraging a frozen LM teacher through layer-wise KV-cache sharing and selective distillation. Our results demonstrate that LINGUDISTILL recovers performance on language and knowledge benchmarks while maintaining comparable performance on vision-heavy tasks, surpassing standard finetuning objective with zero additional parameters. These findings suggest that targeted cross-modal distillation provides a simple and effective approach to preserving backbone capabilities in multimodal systems without introducing additional modules.

A Training Configuration

The shared training configuration across all variants includes:

(a) Hard loss across KD-strength variants.

(b) Soft loss across KD-strength variants.

Table 4 lists hyperparameters shared across all variants. All experiments use nanoVLM-full (Wiedmann et al., 2025) with a total size of 460M parameters, which pairs a SigLIP2-B/16 vision encoder (Tschannen et al., 2025) with a SmolLM2-360M-Instruct language decoder (Allal et al., 2025). The teacher is a frozen copy of the original SmolLM2-360M-Instruct backbone used in nanoVLM-full.

D Sequence Length Analysis

We train with a maximum sequence length of 1024 tokens and a single image per example. The overall mean across all sources is under 120 tokens, meaning the 1024-token limit provides ample headroom for most training examples.

E Detailed OCR Analysis

The regression in Key Information Extraction and Doc-oriented VQA together make up 89 of the 126-point total drop (71%), even though they are only 40% of the benchmark. We see the same pattern in DocVQA, where the biggest losses are in table/list questions (−109), layout questions (−84), and free-text extraction (−89).

F Expanded & Ablation Results

Selective Distillation achieves the best language gains with contained vision degradation. We attribute LINGUDISTILL’s success to selective distillation. The main LINGUDISTILL setting reaches the lowest CE curve overall, while High KD stays slightly above it and Low KD remains the highest throughout training <ref:2604.008291,9Figure 2Training-loss analysis for the three selective distillation variants. Left: the CE term is lowest for the main LINGUDISTILL setting, with High KD slightly above it and Low KD clearly worse <ref:2604.

Improvements for AI systems

  1. To recover native language capability in VLMs without additional modules, implement LINGUDISTILL to utilize the original frozen LM as a teacher. This results in a linguistically improved VLM with zero additional modules and recovers ∼10% of the performance lost on language and knowledge benchmarks.

  2. Implement layer-wise KV-cache sharing between the student decoder (Φ) and the frozen teacher decoder (omega) to enable vision-conditioned supervision. This mechanism allows the teacher to access the same multimodal context as the VLM and produce vision-aware supervision signals without modifying architecture or adding parameters at inference time.

  3. Employ data-dependent weighting in the selective distillation objective, using data-dependent weighting, where distillation is primarily applied to language-intensive data. This ensures that LINGUDISTILL restores linguistic capability by applying distillation only on languagecapability recovering data while preserving the original structure of the VLM on visual-heavy or general multimodal data.

  4. Utilize the selective KD objective based on source categorization, where Language-heavy (8) receive the full KD signal and OCR/doc-heavy (9) receive CE only, as summarized in Table 1. This strategy allows for a superior trade-off: it maximizes language gains while preserving visual capability significantly better than uniform distillation.

  5. Apply distillation hyperparameter tuning based on Figure 2, where the optimal setting is found by balancing the trade-off: The main LINGUDISTILL setting reaches the lowest CE curve overall, while High KD stays slightly above it and Low KD remains the highest throughout training.

Abstract

Turning a pretrained language model (LM) into a vision-language model (VLM) through multimodal fine-tuning often erodes its native language ability, a form of catastrophic forgetting that shows up even on text-only tasks. This loss is hard to undo with further fine-tuning, and existing remedies add adapters or alignment modules that increase architectural complexity and inference cost. We propose LinguDistill, an adapter-free knowledge distillation method that uses the original frozen LM as the teacher during multimodal post-training. To let a text-only teacher supervise vision-conditioned outputs, we introduce layer-wise KV-cache sharing, which exposes the teacher to the student's multimodal representations without changing either architecture. We then apply distillation selectively, on language-heavy data only, so the teacher restores linguistic ability while the student keeps its visual grounding on document and OCR tasks. LinguDistill recovers the language and knowledge performance lost during multimodal fine-tuning, matching the original VLM on average over text-only benchmarks (ARC, HellaSwag) and exceeding it on ScienceQA, while keeping vision-heavy performance close to standard fine-tuning. Since the teacher is dropped after training, the final model adds no parameters and no inference cost. More broadly, our results show that a model's own pre-adaptation backbone is a practical teacher for undoing forgetting, suggesting a simple recipe for keeping language ability intact as models are extended to new modalities.

Sources

Related papers