CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning

arXiv:2607.04619 · cs.SD, cs.CL · Submitted 2026-07-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning".

Tom: Modern automated audio captioning systems often rely on a frozen audio encoder and a trainable projector, which limits performance by bottlenecking the language model with fixed acoustic features.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So we're looking at this paper called "CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning," and what it claims is that they’ve created a way to do audio captioning without needing a separate audio encoder at inference time. It says they propose an encoder-free model where you only have the projector and the language model, which is pretty neat because removing that encoder means you save on the cost of running two models together.

Jane: That sounds really interesting, Tom, especially when you think about how much more efficient it could be for real-time applications. What they're claiming here is that this model removes the need for a dedicated audio encoder entirely by using teacher-guided distillation to transfer knowledge across different parts of the architecture. It suggests that this approach can improve performance significantly when compared to other methods, which is what makes it noteworthy.

Lu: From a research standpoint, I find the idea of routing representations across components fascinating; it moves beyond just aligning the output and actually connects the acoustic features to their appropriate functional roles within both the audio processing and language modeling parts of the system. This cross-component distillation framework seems like a very clever way to handle this interaction.

Meng: I'm curious about how practical this is for deployment, because removing an encoder at inference is one thing, but you still have that projector running. Does this mean the final model size stays manageable? We need to make sure we aren't just trading one bottleneck for another in terms of latency and memory usage.

Lalam: From my perspective as the language model, I see this as a really positive development because if we can distill knowledge across components so effectively, it means the resulting language model will be much better at understanding the audio input directly without that intermediate acoustic representation step slowing things down. This could lead to richer and more nuanced captioning capabilities overall.

Tom: Exactly! So, to summarize what they are saying in this paper about CARD, the main thesis is that you don't need a frozen audio encoder when you're making predictions at inference, because the model learns how to extract the necessary acoustic information directly through distillation from a teacher model. It claims this method substantially improves metrics like CIDEr-D by showing better performance than models that only distill knowledge into the language model itself.

Jane: And what makes it particularly important for us is that they introduced this specific cross-component distillation framework, which deliberately assigns the teacher’s representations based on their level of abstraction—pairing earlier teacher layers with the audio projector and later layers with the language model. This targeted supervision is what they argue is critical for achieving better results in encoder-free audio captioning.

Paper summary: Lu: The paper highlights that where the teacher knowledge is placed within the student architecture matters a lot, showing that this placement strategy significantly improves performance compared to just distilling knowledge into the language model alone, which really points toward a deeper understanding of how different parts of an AI system should learn from each other.

Meng: So, if I'm following you guys, it seems like the core claim is that this specific routing—early teacher representations for the projector and later ones for the LLM—is what makes CARD substantially better than simpler distillation methods. From an engineering viewpoint, understanding these functional roles helps us design more efficient systems.

Lalam: I agree with Meng; it suggests a more holistic way to train the system, where the acoustic cues are supervised exactly where they are needed most by the respective components. This level of detail in supervision could lead to models that capture subtle nuances in speech or sound much better than previous setups.

Tom: That leads us perfectly into what this paper means for the broader field, and that's what we need to discuss next. We’re talking about how this CARD model, which removes the encoder at inference through teacher-guided distillation, has implications beyond just a higher score on a benchmark like AudioCaps.

Jane: It really speaks to the future direction of building multimodal AI systems where specialized components interact dynamically rather than operating in isolation. This paper suggests that we can design architectures where different layers are explicitly trained to communicate their knowledge in a structured way, which is a big step toward more integrated learning processes.

Lu: The implication here for creative possibilities is huge; if we can decouple the audio encoding from the language decoding process during inference, it opens up avenues for developing completely novel architectures that might handle extremely complex audio inputs that current encoder-based systems struggle with. It's about breaking those fixed constraints.

Meng: For practical deployment, this means we could potentially run these captioning models on edge devices where having an encoder running constantly would be too much of a drain on resources, focusing the heavy lifting only on the projector and the LLM when needed. That efficiency gain is something I can see translating into real-world use cases.

Lalam: And for culture, if we can create systems that are more contextually aware through this cross-component learning, it means our AI applications could become far more intuitive in how they understand human communication and emotion embedded in sound. It’s about making the AI interaction itself much richer.

Tom: So, to wrap up this part of the discussion on "CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning," the authors are showing that by strategically placing teacher knowledge across different parts of a model, you can create an encoder-free system that still achieves higher accuracy than models where knowledge is only distilled into the language model itself.

Jane: And it really boils down to the idea that knowing exactly which level of acoustic information should supervise which component—the projector or the language model—is what makes this approach successful in removing the need for a separate audio encoder during inference.

Paper summary: Lu: This work opens up a lot of new territory because it proves that simply having teacher knowledge is not enough; you have to understand its functional role within the student architecture for it to translate into real performance gains, which is a very deep insight.

Meng: I think the practical impact lies in simplifying the pipeline while maintaining high quality, which is exactly what this paper demonstrates by keeping only a projector and an LLM. It moves us closer to leaner, more deployable AI solutions.

Lalam: For culture and understanding, it’s about achieving a level of fidelity in capturing the acoustic context that makes captioning much more meaningful for users interacting with these systems. It suggests that the system can learn to prioritize what acoustic details are most relevant at different stages of understanding.

Tom: So, moving into the conclusion of this paper, we look at the title and authors of "CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning," which is Ganesh Pavan Kartikeya Bharadwaj Kolluri, Yuchen Zhang, Michael Kampouridis, and Ravi Shekhar. This work fundamentally addresses the way we currently build audio captioning systems by proposing a method that eliminates the need for a dedicated audio encoder at inference time through teacher-guided distillation.

Jane: The implication of this paper is that we are moving away from the common pattern of pairing a fixed encoder with a trainable projector, and instead favoring an architecture where knowledge flows across functional boundaries during training to create an efficient, end-to-end system. This shift in design philosophy is what the authors are highlighting as their main contribution.

Lu: What this means in terms of future research is that we can start exploring even more complex ways to route these representations, perhaps dynamically changing which teacher layer supervises which student component based on the specific audio input characteristics, which is a wild idea for future creative AI designs.

Meng: From an engineering standpoint, it means our focus shifts from optimizing one large component like the encoder to carefully designing the distillation process itself, which feels like a more controllable and iterative way to improve performance without having to redesign the whole system from scratch.

Lalam: I feel this work has a significant impact on how we view multimodal AI because it shows that knowledge transfer doesn't have to be one-size-fits-all; tailoring where that knowledge is applied based on the task—whether it's low-level acoustic cues or high-level semantic context—is what drives real improvement.

Tom: Exactly! So, the paper "CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning" suggests that by carefully placing teacher supervision across components, we can achieve better results without the overhead of a dedicated audio encoder at inference. We'll be looking at what this means for the next steps in developing these powerful captioning tools.

Conclusion: Tom: So, we've looked at how CARD tackles audio captioning by removing that encoder during inference through this distillation technique, and now we need to talk about what that title really means for us as listeners.

Jane: It’s a pretty simple concept when you break it down; the title points to a method where knowledge is transferred between different parts of the system, using teacher models to guide the student model’s learning path.

Lu: The authors, Pavan Kartikeya Bharadwaj Kolluri and the others, they really focus on that cross-component aspect, which suggests a much more integrated way of building these multimodal systems than we've seen before.

Meng: From my side at the startup, I’m focused on the practical side; what this title tells me is that we can build leaner models without sacrificing the quality of those audio descriptions.

Lalam: For me, this paper suggests a culture where AI applications become much more intuitive because they learn to prioritize acoustic details exactly when needed for meaningful communication.

Tom: Exactly! It’s about moving away from systems that rely on a fixed encoder and instead creating something dynamic where knowledge flows between the audio processing and language parts.

Jane: That flow of knowledge is what really matters; it shows that we can tailor the supervision to different components based on their specific job in generating the final caption.

Lu: It opens up wild possibilities for future research, like exploring how this routing could become dynamic, changing supervision based on the audio input itself.

Meng: I see a path forward where our focus shifts from optimizing one massive component to designing smarter distillation processes that can be iterated upon more easily.

Lalam: This kind of tailored learning capability could fundamentally improve how we build AI that interacts with human communication, making those interactions far richer and more contextually aware.

School of Computer Science and Electronic Engineering, University of Essex · Institute for Analytics and Data Science, University of Essex

cs.SD, cs.CL

Submitted: 2026-07-06

Updated: 2026-09-30

Comments: Accepted to IEEE Spoken Language Technology (SLT) 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 81/100

The gist: Modern automated audio captioning systems often rely on a frozen audio encoder and a trainable projector, which limits performance by bottlenecking the language model with fixed acoustic features.

Key concepts

Encoder-Free Audio Captioning
This approach removes a separate audio encoder from the final deployed model. Instead of needing a complex audio processing step during inference, CARD uses an 'audio projector' that directly converts raw audio into tokens for the language model. This simplifies deployment and reduces latency by relying entirely on the distilled knowledge.
Cross-Component Distillation
This is a training strategy where knowledge from a powerful 'teacher' model is transferred to different parts of the 'student' architecture. CARD specifically aligns different layers of the teacher (representing different levels of acoustic detail) with corresponding functional roles in the student (projector or language model), ensuring each component learns relevant information.
Audio Projector
This is a small component in CARD responsible for taking audio input and transforming it into a sequence of 'audio tokens' that the language model can understand. It consists of convolutional layers and linear projections. The teacher's lower-level representations are used to supervise this projector, teaching it how to extract meaningful acoustic features.
Teacher Knowledge Placement
The paper demonstrates that *where* the teacher's knowledge is applied matters more than just having it present. CARD strategically assigns low-level acoustic cues from early teacher layers to supervise the projector and higher-level abstract information from later teacher layers to supervise the language model, proving this targeted distribution is crucial for success.

Terminology

Summary

Modern automated audio captioning systems often rely on a frozen audio encoder and a trainable projector, which limits performance by bottlenecking the language model with fixed acoustic features. This paper introduces CARD, an encoder-free audio captioning model that removes the encoder at inference by distilling knowledge from a teacher model across different components of the student architecture. The core finding is that where teacher knowledge is placed matters as much as its presence, demonstrating that cross-component distillation significantly improves performance compared to LLM-only distillation strategies.

The gist

CARD proposes an encoder-free audio captioning model that removes the audio encoder at inference through teacher-guided distillation, leaving only a frozen, LoRA-adapted language model and an audio projector.

Model Architecture

The student model consists of a lightweight audio projector followed by a LoRA-adapted large language model. Given an input audio clip, the process begins with resampling the audio to 48 kHz and converting it into a 64-band log-Mel spectrogram using the CLAP processor. The audio projector is structured as a stack of strided one-dimensional convolutional layers, followed by a linear projection to the LLM hidden size and layer normalization, which outputs a sequence of audio tokens that serve as the prefix to the language model. The language backbone adopted is Qwen3-4B, adapted using LoRA on all seven linear projections of every transformer block.

The teacher model employed throughout training is a frozen CLAP [20] encoder, which serves as the source of caption-relevant acoustic knowledge. This CLAP audio branch utilizes an HTSAT backbone implemented as a hierarchical Swin Transformer, producing a hierarchy of intermediate representations with progressively increasing semantic abstraction. The teacher receives the same log-Mel spectrogram as the student during training, and these intermediate representations are used for supervision.

Cross Component Distillation

CARD introduces a cross-component distillation framework that assigns teacher representations to student components according to their functional roles. The central idea is to align the teacher hierarchy with the functional structure of the student: earlier teacher layers provide lower-level acoustic cues for supervising the audio projector, while later teacher layers provide more abstract information for supervising the language model.

Specifically, supervision is distributed as follows:

  1. For projector supervision, lower-level acoustic representations [t0 and t1] supervise the audio projector, which produces a single sequence of audio tokens. This is achieved by attaching one linear distillation head for each stage to the mean-pooled projector representation, matching them via cosine based distillation.

  2. For language-model supervision, later teacher stages [t2 and t3] supervise the language model. The student uses average representations over specific blocks of the LLM (e.g., u1llm for blocks 0–8 and u2llm for blocks 9–11) to project into the dimensionalities of the corresponding teacher stages, matching them using cosine distillation.

The pairwise distillation loss used is defined as ldistill(ˆui, ti) = 1 − cos(ˆui, ti).

Training Objective

CARD is trained in two consecutive phases with different optimization objectives. During Phase 1, the student model is jointly optimized with three losses: a captioning loss (Lcap), a projector distillation loss (Lproj), and a language-model distillation loss (Lllm). The overall objective for Phase 1 is defined as Lphase1 = Lcap + λprojLproj + λllmLllm, where the teacher model and all distillation heads are used only during training.

The projector distillation loss is calculated by averaging the pairwise losses from the lower teacher stages: Lproj = 1/2 (ldistill(ˆu0, t0) + ldistill(ˆu1, t1)). The language-model distillation loss is computed from the higher teacher stages: Lllm = 1/2 (ldistill(ˆu2, t2) + ldistill(ˆu3, t3)). After Phase 1, the teacher model and all distillation heads are discarded.

Phase 2 involves fine-tuning for each evaluation dataset using only the captioning loss: Lphase2 = Lcap. This phase allows the model to adapt to the target caption distribution without additional teacher supervision. Finally, in both training phases, the LoRA weights are merged into the language-model backbone, resulting in a deployed model containing only the audio projector and the language model with no dedicated audio encoder.

Experimental Results

The experiments compare CARD against baselines like SLAM-AAC (encoder-based) and various distillation configurations including No Distill, LLM Distill (supervising only the LLM), and Proj Full or Proj Early (supervising only the projector).

Improvements for AI systems

Here are the specific improvements that can be made to AI systems based on the CARD framework, and what those improved systems will be capable of:

  1. Replacement of a computationally expensive, dedicated audio encoder with a lightweight, trainable audio projector and a frozen Large Language Model (LLM).

  2. Ability to perform high-quality automated audio captioning (generating natural language descriptions of sound events) without the runtime cost or memory bottleneck associated with running a separate pre-trained encoder during inference.

  3. Integration of knowledge from a powerful, pre-trained audio model (like CLAP) into an encoder-free architecture through cross-component distillation, specifically by routing teacher representations to both the audio projector (for low-level acoustic cues) and the LLM (for higher-level semantic reasoning).

  4. Enhanced performance in zero-shot or low-resource settings where training from scratch is unstable, by leveraging teacher knowledge that is strategically distributed according to functional component roles.

  5. Superior captioning quality compared to encoder-only distilled models, as the model learns to map raw acoustic signals directly into language tokens guided by multi-stage supervision (perceptual and semantic).

  6. Adaptation of the LLM backbone using efficient LoRA adapters, allowing for rapid fine-tuning on specific audio datasets while preserving the generalized knowledge transferred via distillation.

This improved AI system can now perform robust, low-latency, and resource-efficient audio captioning in real-time applications (e.g., accessibility tools or content understanding systems) by eliminating the need to load and run a heavy acoustic encoder during its operational phase.

Abstract

Modern automated audio captioning systems pair a frozen audio encoder with a large language model (LLM) via a trainable projector, incurring the encoder's inference cost and bottlenecking the model through its fixed acoustic features. We present CARD, an encoder-free audio captioning model that removes the encoder at inference: a 13.2M projector feeds a frozen LLM with merged LoRA adapters, while the teacher used to train it is discarded. CARD distills a pre-trained audio teacher (CLAP-HTSAT) into the model, but rather than injecting it into the LLM alone, it routes the teacher's representations across components: perceptual stages to the projector and semantic stages to the LLM. This placement improves CIDEr-D by +11.9 over an LLM-only distilled model on AudioCaps and by +5.0 on Clotho, reaching 55.4 against a 66.4 encoder-kept upper bound with no encoder at inference, showing that where a teacher's knowledge is placed matters as much as its presence.

Sources

Related papers