Discriminative and Consistent Representation Distillation

arXiv:2407.11802 · cs.CV, cs.AI · Submitted 2024-07-16 · Read on arXiv

Imperial College London

cs.CV, cs.AI

Submitted: 2024-07-16

Updated: 2026-09-02

Comments: Accepted at the 19th European Conference on Computer Vision (ECCV 2026) Workshops

Code: https://github.com/giakoumoglou/rrd

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 69/100

The gist: The paper "Discriminative and Consistent Representation Distillation" proposes a new method called Discriminative and Consistent Representation Distillation (DCD) to improve Knowledge Distillation

Terminology

Summary

The paper Discriminative and Consistent Representation Distillation proposes a new method called Discriminative and Consistent Representation Distillation (DCD) to improve Knowledge Distillation (KD). The authors identify that while contrastive objectives are effective for learning structured representations, their use in distillation is hindered by two practical shortcomings: the reliance on external memory banks for negative sampling, and fixed temperature hyperparameters that limit adaptability across training stages and teacher-student pairs. Furthermore, they note that "instance-discrimination objectives supervise only the diagonal of the cross-model similarity matrix, leaving its off-diagonal structure unconstrained: nothing prevents the student from matching its own teacher counterpart while relating inconsistently to the remaining samples in the batch."

To address these issues, the authors propose DCD, which combines contrastive instance discrimination with a consistency regularization term over the cross-model similarity matrix. The method's dual objective is formulated as follows:

  • Contrastive Component: This term aligns each student representation with its teacher counterpart using an InfoNCE objective. The authors employ contrastive learning to align teacher and student representations at the instance level, treating the task as an N-way classification problem where the student must identify its corresponding teacher representation among all teacher representations in the batch.

  • Consistency Component: This term penalizes asymmetry between the row-normalized and column-normalized views of that matrix, constraining the off-diagonal structure that instance discrimination alone leaves free. The consistency objective enforces agreement between the two [views] by minimizing the KL divergence between the student view (row-normalized) and the teacher view (column-normalized). The authors prove that this term vanishes precisely when this matrix is symmetric, meaning the affinity of student i towards teacher j agrees with the affinity of teacher i towards student j.

The similarity matrix is computed using learnable scale and bias parameters that adapt during training to control the sharpness and offset of the distillation signal, defined as ij = s times z i S, z j T + b, where s = (tau). This approach automatically adjusting the sharpness and offset of the distillation signal rather than relying on fixed hyperparameters.

Regarding implementation and efficiency, the method replaces external memory banks with an efficient in-batch sampling strategy, using only the negative samples that naturally co-exist within each mini-batch. This significantly reduces the negative storage from roughly 655MB on ImageNet (CRD) to 0.13MB per iteration. The method matches the training speed of standard KD while adding only 66K additional parameters, all of which are discarded at inference.

Experimental results demonstrate the effectiveness of DCD across several domains:

  • Image Classification: On CIFAR-100, "DCD combined with KD achieves strong performance, surpassing the teacher network by +0.45% in the same-architecture setting (WRN-40-2 to WRN-16-2) and by +0.90% in the cross-architecture setting (WRN-40-2 to ShuffleNet-v1). On ImageNet, the method consistently improves upon baselines and achieves competitive performance across different architectures."

  • Object Detection: On MS-COCO, Our method consistently improves upon baselines [13, 37, 67, 84] across all three settings, demonstrating that the gains of our approach extend beyond classification to the more challenging object detection task.

  • Transferability: In cross-dataset transfer tests (STL-10 and Tiny ImageNet), Our method, both standalone and combined with KD, consistently improves transferability over baselines, suggesting that the learned representations capture generalizable features rather than overfitting to the training distribution.

  • Efficiency: The authors show that our method matches KD and FitNet as the fastest approach at 8ms per batch, and is over 2× faster than CRD.

Improvements for AI systems

1. High-Fidelity Edge Vision Systems

  • The Improvement: Implementing DCD to distill massive, high-parameter teacher models (like Vision Transformers) into lightweight, mobile-optimized student architectures (like ShuffleNet or MobileNet) using learnable scale/bias parameters and consistency regularization.

  • What the improved system can do: It can perform real-time, high-accuracy object detection and image classification on resource-constrained hardware, such as smartphones, drones, or IoT devices, maintaining performance levels that previously required high-end GPUs.

2. Domain-Robust Foundation Models

  • The Improvement: Utilizing the consistency component (minimizing KL divergence between row-normalized and column-normalized similarity views) to ensure the student learns a structured, symmetric representation space rather than just mimicking label outputs.

  • What the improved system can do: It can achieve superior out-of-distribution generalization. For example, a medical imaging AI trained on one specific hospital's dataset can be deployed to a different hospital with different equipment and patient demographics without a significant drop in diagnostic accuracy, because its learned features are more generalizable and less overfitted.

3. Memory-Efficient Large-Scale Training Pipelines

  • The Improvement: Replacing heavy, memory-intensive external memory banks used in traditional contrastive distillation (like CRD) with DCD’s efficient in-batch sampling strategy.

  • What the improved system can do: It enables the training of large-scale models on limited hardware clusters by drastically reducing VRAM requirements (from hundreds of megabytes to nearly zero for negative sampling). This allows researchers to increase batch sizes or model depth within the same hardware budget, accelerating the development cycle of large vision models.

4. Cross-Architecture Knowledge Transfer Engines

  • The Improvement: Applying the dual-objective (contrastive + consistency) to bridge the structural gap between fundamentally different neural architectures (e.g., transferring knowledge from a Transformer-based teacher to a CNN-based student).

  • What the improved system can do: It can create highly specialized expert models. For instance, a system can take the complex relational knowledge of a heavy Transformer and compress it into a fast, specialized CNN designed for specific low-latency tasks like autonomous vehicle obstacle avoidance, ensuring the student understands not just what an object is, but how it relates to other objects in the scene.

Abstract

Knowledge Distillation (KD) transfers knowledge from a large teacher to a smaller student model. While contrastive objectives have proven effective for learning structured representations in self-supervised settings, their use in distillation is hindered by two practical shortcomings: the reliance on external memory banks for negative sampling, and fixed temperature hyperparameters that limit adaptability across training stages and teacher-student pairs. We therefore propose Discriminative and Consistent Representation Distillation (DCD), which combines contrastive instance discrimination with a consistency regularization term over the cross-model similarity matrix. The contrastive term aligns each student representation with its teacher counterpart, while the consistency term penalizes asymmetry between the row-normalized and column-normalized views of that matrix, constraining the off-diagonal structure that instance discrimination alone leaves free; we show that it vanishes precisely when this matrix is symmetric. We further introduce an efficient in-batch sampling that eliminates external memory banks, and learnable scale and bias parameters that adapt during training to control the sharpness and offset of the distillation signal. The method matches the training speed of standard KD while adding only 66K additional parameters. Through extensive experiments on CIFAR-100, ImageNet, and MS-COCO, together with cross-dataset transfer to STL-10 and Tiny ImageNet, we show that our approach achieves competitive performance in classification, object detection, and transfer, while substantially reducing memory consumption and training time compared to existing contrastive distillation methods.

Sources

Related papers