Diffusable Latents from Structure-Agnostic Distillation
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Diffusable Latents from Structure-Agnostic Distillation".
Jane: Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Okay, so we're looking at the paper "Diffusable Latents from Structure-Agnostic Distillation," and the title itself tells you right away what they're focusing on—making latents more diffusible by being structure-agnostic during distillation. The authors are Ramanana Rahary, Nicolas Dufour, Patrick Pérez, and David Picard from LIGM at ENPC in Paris.
Jane: That title really captures the essence of what they are doing: they are taking a concept about latent diffusion and making it more robust by focusing on the structure-agnostic aspect of how we transfer knowledge from a teacher model to a student model. It sounds complicated, but let's break down what that means for us.
Lu: What’s interesting is the authors' motivation in moving away from the standard way of doing things, which is position-wise matching, where every single latent position has to align perfectly with its corresponding feature in the teacher model. They are arguing that this constraint is unnecessary for many situations and that pooling works just as well.
Meng: From a research perspective, it’s interesting how they frame this as a way to remove unnecessary constraints on the latent space itself. It suggests that the internal layout of our student latent doesn't need to conform rigidly to the teacher's organizational structure for us to get good results.
Lalam: I see this as a step towards building more adaptable AI systems; instead of designing models specifically around rigid structural requirements, we can design them around powerful, transferable image-level concepts that are distilled efficiently. It makes the resulting AI feel less constrained and more capable.
The paper's summary: Tom: Now let's get into the specifics of what they actually propose in "Diffusable Latents from Structure-Agnostic Distillation." They summarize it by showing that distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, which directly helps diffusion models converge faster and reach higher sample quality.
Jane: So the main point is that this distillation technique is a way to make the latent space more 'diffusable,' meaning it allows for better sampling trajectories during generation. This leads to faster convergence and better final outputs from our generative models.
Lu: They specifically contrast their approach with dense position-wise distillation by showing that aligning a single pooled image-level descriptor to the teacher’s descriptor achieves the same performance as, or even slightly surpasses, dense position-wise matching in some settings.
Meng: The paper details a few specific objectives they test: first-order matching where each descriptor matches directly, and relational objectives which compare images only through similarity matrices without needing any projector at all. They show these relational methods are also very effective for improving diffusability.
Lalam: It's fascinating how they cover different scenarios, including 1D token latents and distillation from text encoders into image autoencoders, proving that this method isn't limited to one type of setup but is quite versatile across various modalities.
The paper's improvements: Tom: Moving on to the actual technical improvements they propose, the authors introduce three specific distillation methods: Pool-Align, which aligns a pooled descriptor using a learned linear map W, and two relational objectives called Centered Kernel Alignment and Softmaxed Similarity KL.
Jane: Pool-Align seems like their star objective because it uses cosine distance to align those image descriptors directly with the teacher's space via a linear map W, which is quite clean mathematically. It simplifies the process by focusing on that single pooled descriptor per image rather than dealing with every token separately.
Lu: And then you have the relational objectives, like CKA which matches the entire between-image similarity kernel as a whole and Soft-KL which matches each image's soft ranking of its nearest neighbors. These methods are particularly interesting because they are invariant to the descriptor space, meaning they don't need a projector to work.
Meng: From an implementation viewpoint, CKA and Soft-KL are appealing because they don't rely on that shared grid or a specific channel projector for their calculation, which cuts down on the complexity of the infrastructure needed to support these distillation methods.
Lalam: The authors' conclusion is that structure-agnostic distillation is a simple and flexible alternative to position-wise alignment, allowing us to keep our latent layout completely free while still benefiting from the teacher's knowledge. That freedom in design is what makes it so compelling for future AI development.
Conclusion: Tom: So, wrapping up this discussion on "Diffusable Latents from Structure-Agnostic Distillation," the main takeaway is that pooling image descriptors provides a flexible and effective way to boost latent diffusibility without needing to constrain the internal structure to match a teacher's spatial layout.
Jane: It really boils down to using these structure-agnostic techniques, like Pool-Align, which helps generative models converge faster and reach higher sample quality because the distilled latent is more diffusable. This means our sampling process will be smoother overall.
Lu: I think this work has huge implications because it shows that we can leverage powerful visual guides without having to impose a teacher's spatial organization onto our student model, opening up possibilities for much broader applications of foundation models in generation tasks.
Meng: Practically speaking, this means we can build more versatile AI systems that don't get locked into specific latent structures and can handle a wider variety of input modalities with the same improved performance metrics.
Lalam: I think the impact is that we get more reliable and adaptable generative AI components across the board because they are built on distilled knowledge in a way that respects their inherent structure while leaving their own layout open.
Tom: Fantastic summary, everyone; we’ve covered how this paper introduces "Diffusable Latents from Structure-Agnostic Distillation," and I think we have a clear picture of why this work is so important for the generative community.
Adrien Ramanana Rahary, Nicolas Dufour, Patrick Pérez, David Picard
LIGM · ENPC · IP Paris · CNRS
cs.CV
Submitted: 2026-09-30
Updated: 2026-09-30
Code: https://github.com/AdrienRR/structure-agnostic-distillation
Importance score: 88/100
The gist: Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality.
Key concepts
- Latent Diffusability
- This refers to how easily a latent representation can be manipulated by diffusion models during training. Better diffusability means the model can learn the underlying data distribution more effectively, leading to faster convergence and higher quality generated samples.
- Position-wise Distillation
- A previous method that matched every single token in the student's latent space to a corresponding token in the teacher's latent space. This approach was constrained by the specific layout of both latents, which is unnecessary for many scenarios like compact latents or different teacher types.
- Structure-Agnostic Distillation
- This proposed method pools all image latent tokens into one single descriptor before matching it to the teacher's pooled descriptor. This makes the distillation process independent of the specific token count or layout of either latent, allowing for flexible student and teacher shapes.
- Relational Objectives
- These are distillation methods that focus on matching the relationships between images rather than their exact token positions. Examples include matching similarity matrices or using soft ranking losses to ensure the distilled latent captures the same semantic structure as the teacher's.
Terminology
Summary
Distilling pretrained foundation models into an autoencoder bottleneck improves latent diffusability, enabling diffusion models to converge faster and reach higher sample quality. The gist: Aligning a single pooled image-level descriptor to the teacher’s performs as well as or slightly better than dense position-wise distillation.
The Core Problem and Proposed Solution
Latent diffusion trains a prior in an autoencoder's latent space, but sample quality depends on how diffusable
that latent is. Previous work improved diffusability by matching each latent position to its co-located patch feature (position-wise distillation), but this constraint is unnecessary for various scenarios, such as compact 1D token latents or teacher modalities from different domains like text encoders. The authors propose structure-agnostic distillation, which pools each image's latent tokens into one image-level descriptor and defines the objective on these descriptors alone, leaving the latent’s internal layout free. This approach is independent of the latent’s token count and layout, allowing student and teacher to have any shape.
Distillation Objectives Compared
The paper compares a family of pooled objectives across different latent shapes and teacher modalities:
-
First-order matching: Aligns each descriptor directly to its teacher descriptor.
-
Relational objectives: Match the matrix of between-image similarities, comparing images only through their similarity matrices, which are invariant to the descriptor space.
The paper instantiates three specific distillation methods:
(Pool-Align):
A first-order match that aligns each image’s pooled student descriptor with the teacher’s descriptor for the same image using a learned linear map W, minimizing cosine distance: Lpool = P i 1 − vˆ i⊤ u˜ i
.
(Relational Objectives):
A relational objective reproduces the teacher’s between-image similarity structure. This includes:
-
Centered kernel alignment (CKA): Matches the two matrices as a whole, using the unbiased estimator:
LCKA = 1 − CKA(Ss, St)
. -
Softmaxed-similarity KL (Soft-KL): Matches each image’s soft ranking of its nearest neighbours by matching a softmax distribution to a KL divergence loss:
LsoftKL = P i KL(p t i p s i)
.
Experimental Setup and Results
The experiments compare structure-agnostic methods against two references: a control with no distillation (No Distillation) and the position-wise baseline, VA-VAE’s vision-foundation alignment (VF), which requires a shared latent–teacher grid. The study varied the student’s shape or the teacher’s modality across three settings:
(A) Shared grid:
A 2D grid latent with a DINOv2 image teacher where the latent lines up with the teacher's. Here, Pool-Align performs as well as or slightly better than position-wise VF (Pool-Align 5.77 vs. VF 6.04)
. Every structure-agnostic objective beats the undistilled No-Distillation (9.52).
(B) No grid:
A 1D sequence latent produced by a Perceiver-resampler with no spatial grid, where position-wise VF drops out, but the four layout-free methods still apply. Pool-Align 15.62 vs. 26.49
beats No-Distillation in this setting as well.
(C) Cross-modal:
A cross-modal 2D grid latent aligned to per-image captions from the captioned ImageNet of Degeorge et al., embedded with a BGE-large text teacher (no spatial grid). For two of the three objectives, Pool-Align (6.97) and Soft-KL (8.51) improve over No-Distillation (9.52), while CKA (10.41) lands above it.
Key Findings on Latent Structure
The results demonstrate that pooling representations can match position-wise alignment by averaging out distractors carried in the teacher’s per-token features, such as positional information and high-norm artifact tokens. A latent PCA analysis confirms this: "the undistilled latent is nearly featureless; the position-wise VA-VAE (VF) latent carries a smooth horizontal spatial gradient, whereas the pooled latents show no such gradient and instead pick out the foreground object from the background. Furthermore, when aligning to a text teacher,
Pool-Align brings gFID from 9.52 to 6.97, showing that representation distillation can transfer useful semantic geometry without inheriting the teacher’s structural organization. The paper concludes that structure-agnostic distillation is a
simple, flexible alternative to position-wise alignment that leaves the latent’s layout free.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements that can be made to AI systems by implementing Structure-Agnostic Distillation:
-
The core improvement is increasing the quality and
diffusability
of generative models (like latent diffusion models) without constraining their internal latent structure to match a teacher's spatial layout. -
By employing structure-agnostic distillation (specifically the Pool-Align objective), AI systems can leverage powerful, frozen foundation models (like DINOv2) as high-quality semantic guides, even when the student model operates on significantly different latent structures (e.g., 1D sequences or cross-modal embeddings).
Specific capabilities of these improved AI systems:
-
The system will achieve higher sample quality in generative tasks (lower gFID) because the distillation process transfers useful semantic geometry from a powerful teacher without forcing the student to adopt the teacher's specific spatial grid.
-
The diffusion model prior will converge faster and reach higher sample quality because the distilled latent space is more
diffusable,
meaning it allows for smoother, more effective sampling trajectories during generation. -
The system can effectively generate high-fidelity images or data from complex, non-grid latents (like Perceiver-resampler outputs or text embeddings) by aligning them to a teacher representation based on image descriptors rather than requiring a token-to-token spatial correspondence.
-
The system will be more robust across different latent modalities: it can successfully improve the diffusability of 1D token sequences and even align an image latent to a text encoder (cross-modal distillation).
-
For tasks involving visual understanding, the improved latent representations will better capture content and semantics by averaging out positional artifacts present in the teacher's features, leading to more content-focused reconstructions.
Sources
- Deep ViT Features as Dense Visual Descriptors
- How far can we go with ImageNet for Text-to-Image generation?
- The FID Lottery: Quantifying Hidden Randomness in Generative-Model Evaluation
- V-RAE: Rethinking Video Latent Spaces for Generation
- GAIA-1: A Generative World Model for Autonomous Driving
- Multiplayer Interactive World Models with Representation Autoencoders
- BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
- GAIA-2: A Controllable Multi-View Generative World Model for Autonomous Driving
- GLU Variants Improve Transformer
- Improved Baselines with Representation Autoencoders
- RoFormer: Enhanced Transformer with Rotary Position Embedding
- Scaling Text-to-Image Diffusion Transformers with Representation Autoencoders
- Diffusion Transformers with Representation Autoencoders
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models