REPA-G: Test-Time Conditioning with Representation-Aligned Visual Features

arXiv:2602.03753 · cs.CV · Submitted 2026-02-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "REPA-G: Test-Time Conditioning with Representation-Aligned Visual Features".

Tom: Test-Time Conditioning with Representation-Aligned Visual Features introduces RepresentationAligned Guidance (REPA-G), a novel framework that leverages representation alignment learned during diffusion model training to enable test-time conditioning from features in generation.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So to recap for everyone, REPA-G introduces RepresentationAligned Guidance, which aims to enable test-time conditioning from features extracted during generation by leveraging representation alignment learned from diffusion model training (<ref:2602.03753#pg0>). The paper claims this provides a path toward achieving fine-grained control over image synthesis without depending on fixed class labels or potentially ambiguous text prompts (<ref:2602.03753#pg1>).

Jane: That's the central thesis, Tom; essentially, they are taking those aligned representations and using them to steer the denoising process toward a target feature representation at inference time through an optimization of a similarity objective (the potential) (<ref:2602.03753#pg1>).

Lu: What makes this significant is that it unlocks versatile control mechanisms, allowing users to guide generation using different levels of granularity, from single patches for texture matching to global image tokens for broad concepts like "rabbit" or "car" (<ref:2602.03753#pg1>).

Meng: It matters because it offers an alternative to the methods that rely on training-free guidance, which some researchers have looked into recently, by making test-time conditioning the main focus of their investigation (<ref:2602.03753#pg2>).

Lalam: This means we can achieve control over concrete visual concepts and abstract semantics in a way that's more precise than what standard text-to-image models offer when relying solely on textual input (<ref:2602.03753#pg1>).

Tom: It’s clear that the method is designed to provide flexible and precise alternatives to the often ambiguous nature of text prompts or coarse class labels during inference, which improves controllability while maintaining flexibility and generality (<ref:2602.03753#pg1>).

Jane: And they achieve this by extracting visual tokens from a real image using the same self-supervised learning network that was used for the representation alignment training, such as DINOv2 (<ref:2602.03753#pg1>). That link to the training data is what makes these features semantically meaningful.

Lu: The paper also validates this approach by analyzing the properties of the representation space, proving that self-supervised features are well-suited for this task because they exhibit the necessary semantic embedding and smooth mapping properties (<ref:2602.03753#pg2>).

Meng: I’m still thinking about how practical it is to extract these tokens on demand during inference without introducing significant latency, which would be a key hurdle for my team at the startup (<ref:2602.03753#pg0>).

Lalam: Ultimately, this framework suggests that visual features provide a denser and more informative signal than text captions when compared against other text-to-image models in terms of conveying specific visual information (<ref:2602.03753#pg1>).

Conclusion: Tom: So wrapping up, we’ve talked about how REPA-G moves conditioning into the test phase by using representation alignment to steer generation, offering granular control from texture patches to global concepts (<ref:2602.03753#pg1>). The authors and their team have shown that this technique works by validating the properties of the feature space, proving that these self-supervised features are indeed well-suited for this task (<ref:2602.03753#pg2>).

Jane: It seems the title itself points to the key innovation: Test-Time Conditioning with Representation-Aligned Visual Features, which is significant because it solves the problem of achieving versatile control without retraining or relying on ambiguous prompts (<ref:2602.03753#pg0>). The implications are that we can create much more controllable visual synthesis systems by using features directly from the generation process itself.

Lu: From a creative perspective, I see this as an opening for entirely new ways to compose images, especially with the multi-concept composition extension they discuss, which could allow for incredibly faithful blending of distinct visual ideas (<ref:2602.03753#pg1>).

Meng: For me, the real impact lies in how this affects deployment; if we can condition on features instead of just text, it opens up possibilities for systems that need to maintain specific structural elements like pose or shape consistently across many different outputs (<ref:2602.03753#pg1>).

Lalam: I think the big cultural impact is in democratizing high-fidelity control; it means less reliance on perfectly written prompts and more reliance on directly providing visual context, which could make AI art tools much more intuitive for a wider audience (<ref:2602.03753#pg1>).

Tom: Exactly, so the authors have shown that these aligned representations offer a "denser, more informative signal than text captions" (<ref:2602.03753#pg1>), and this capability to condition via features at inference time is what makes REPA-G noteworthy (<ref:2602.03753#pg0>).

Jane: It’s a solid piece of research because it provides a principled way to steer the diffusion process using learned semantic structure, moving beyond just standard denoising objectives (<ref:2602.03753#pg1>).

Lu: This work pushes the boundaries on how we use internal features to direct generative models, showing that representation alignment isn't just for training stability but is a powerful tool for inference-time manipulation (<ref:2602.03753#pg2>).

Meng: It’s a solid technical contribution that gives us concrete mechanisms to work with, even if the engineering challenges of implementing it efficiently are still present (<ref:2602.03753#pg0>).

Lalam: I think this research signals a shift toward more controllable and context-aware generative AI systems where visual semantics guide the creation process directly (<ref:2602.03753#pg1>).

cs.CV

Submitted: 2026-02-03

Updated: 2026-10-06

Comments: NeurIPS 2026

Code: https://github.com/valeoai/REPA-G

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 77/100

The gist: Test-Time Conditioning with Representation-Aligned Visual Features introduces RepresentationAligned Guidance (REPA-G), a novel framework that leverages representation alignment learned during

Key concepts

Representation Alignment Training
This training method restores semantic meaning within the internal feature spaces of diffusion models. It forces the model's features to be invariant to augmentations while ensuring they retain their core semantic content. This creates a robust feature space where similar concepts are positioned close together, which is crucial for effective guidance.
Potential Function
The potential function quantifies the similarity or alignment between the generated image's features and the conditioning data's features. By optimizing this potential during inference, the method guides the generation process toward a specific distribution that matches the desired semantic input.
Test-Time Conditioning
This technique allows users to condition image synthesis on visual features extracted from a real image at inference time. Instead of relying on vague text prompts, it uses specific visual tokens—like patches or averages—to precisely guide the model's output toward desired visual characteristics.
Tilted Distribution Sampling
The guidance term modifies the standard sampling process to sample from a 'tilted distribution.' This ensures that the generated image is steered away from a generic output and towards a specific semantic concept defined by the conditioning features, effectively controlling the synthesis direction.

Terminology

Summary

Test-Time Conditioning with Representation-Aligned Visual Features introduces RepresentationAligned Guidance (REPA-G), a novel framework that leverages representation alignment learned during diffusion model training to enable test-time conditioning from features in generation. This method is significant because it provides versatile, fine-grained control over image synthesis at inference time without requiring retraining or the ambiguity of text prompts or coarse class labels.

How it works

The core mechanism of REPA-G involves optimizing a similarity objective (the potential) at inference to steer the denoising process toward a conditioned representation extracted from a pre-trained feature extractor. This is achieved by computing the gradient of this potential, which quantifies the alignment between the representations of the generated sample and conditioning data points corresponding to semantic concepts tokens. The sampling procedure is then guided by this gradient, aiming to sample from a tilted distribution induced by the defined potential.

The framework operates entirely at inference time without requiring fine-tuning or retraining. It extracts visual tokens from a real image using the same self-supervised learning network used during representation alignment training, such as DINOv2. These tokens can be used for conditioning at multiple scales:

  1. A single token for fine-grained texture matching via patches.

  2. A set of tokens to preserve shape or pose when conditioned with a mask.

  3. The average of all image tokens to provide a broad concept signal, such as rabbit or car.

Representation Space Properties

The paper justifies that this guidance enables sampling from the steered density and that self-supervised features are well-suited for this task by analyzing the properties of the representation space. Effective conditioning requires two main properties:

  1. A semantically meaningful embedding space, where nearby embeddings correspond to similar concepts.

  2. A smooth mapping from conditioning features to conditional distributions, ensuring that similar conditioning embeddings yield similar conditional densities.

Representation alignment training restores semantic information in internal feature spaces of diffusion transformers by enforcing invariance to augmentations while preserving semantic content, which is a key advantage over the standard denoising objective alone. The Euclidean embedding property is validated by showing that the distance between conditional densities scales with the squared distance between their conditioning features: AlVert ϕ (x1) - ϕ (x2) rVert squared ≤ d 2(p1, p2) ≤ BlVert ϕ (x1) - ϕ (x2) rVert squared.

Design Choices for the Potential V

The potential function is defined to quantify the alignment between representations. The base potential is given by:

(h, h⋆)

(h, h⋆)

(h, h⋆)

The paper details three specific configurations for generalizing this potential during inference:

  1. Alignment with full feature map (IPA): Setting the weight matrix P = 1/N recovers the original objective, enforcing a dense, spatial constraint.

  2. Alignment with feature mask: A binary mask m is used to preserve information within a masked region while allowing for structural variation in surrounding areas.

  3. Alignment with an average concept: This potential aligns the spatial average of predicted features with the spatial average of conditioning features, preserving general semantic category while removing all spatial constraints.

Guidance at Inference

Inference is performed by introducing a guidance term that modifies the score function in the standard SDE sampling process to steer it toward a tilted distribution. The modified SDE is:

(d x t = (v⋆(x t, t) + t∇x log p t(x t)) dt + √2t dW¯ t)

This modification is equivalent to sampling from the tilted distribution:

(tilde p0(x; xc) ∝ p0(x) e λmathscr V (ϕ(x), ϕ(xc)))

The guidance term is formalized as:

(∇ x log p t(x t) + λ∇mathscr V x((hθ ◦ fθ)(xt,t),ϕ (xc)))

where λ > 0 controls the influence of the conditioning signal.

Experimental Results

Quantitative results on ImageNet and COCO demonstrate that REPA-G achieves high-quality, diverse generations. The method shows control over concrete and abstract visual concepts, handling objects, textures, and background semantics through various guidance granularities (Full Feature Map Masked Feature Map Average Feature Map). For multi-source composition, the Selective Patch Alignment (SPA) variant is shown to be most effective at blending sources using a temperature parameter T. Instance-level evaluations confirm that REPA and REPA-E nearly perfectly reconstruct the condition across all levels of granularity, whereas standard SiT models struggle due to an ambiguous feature space. Furthermore, visual features provide a denser, more informative signal than text captions when compared against text-to-image models.

Improvements for AI systems

As a fastidious researcher, I have analyzed the provided paper, Test-Time Conditioning with Representation-Aligned Visual Features (REPA-G), which introduces a novel inference-time guidance framework for diffusion models.

Here are the specific improvements and capabilities that can be realized by implementing REPA-G:


)

  1. The AI system will gain the capability for high-precision, controllable image synthesis based on external visual features rather than ambiguous text prompts or coarse class labels.

  2. The system can perform fine-grained texture matching by conditioning generation on specific patches extracted from an anchor image (e.g., matching a specific pattern or texture from a reference object).

  3. The system can achieve broad semantic guidance by using global image feature tokens to steer the entire generation toward a general concept (e.g., generating a rabbit regardless of exact pose or setting).

  4. The AI will support multi-concept composition, allowing for the faithful combination of distinct visual concepts derived from multiple source images (e.g., placing an object from Image A onto the background of Image B).

  5. The system can generate images conditioned on specific spatial structures (e.g., preserving the shape or pose of a subject by using feature maps extracted from a masked region of an anchor image).

  6. The AI will demonstrate superior performance in zero-shot or low-resource settings when conditioning on visual features, as shown by its robust results across ImageNet and COCO datasets compared to text-to-image baselines (CAD-I).

  7. The system's internal representation space becomes semantically coherent; the alignment training ensures that similar concepts cluster together in the feature space, leading to more stable and reliable conditioning mappings.

  8. The system can be deployed efficiently at inference time without requiring any additional retraining, fine-tuning, or complex architectural modifications beyond leveraging a pre-trained self-supervised encoder (like DINOv2).

  9. The system can adapt its guidance strategy dynamically using Selective Patch Alignment (SPA), allowing it to switch between rigid spatial constraints and flexible semantic blending based on the chosen temperature parameter.

Abstract

While representation alignment with self-supervised models has been shown to improve diffusion model training, its potential for enhancing inference-time conditioning remains largely unexplored. We introduce Representation-Aligned Guidance (REPA-G), a framework that leverages these aligned representations, with rich semantic properties, to enable test-time conditioning from features, in generation. By optimizing a similarity objective (the potential) at inference, we steer the denoising process toward a conditioned representation extracted from a pre-trained feature extractor. Our method provides versatile control at multiple levels of granularity, ranging from patch level matching via single patches to broad semantic guidance using global image feature tokens. We further extend this to multi-concept composition, allowing for the faithful combination of distinct concepts. REPA-G operates entirely at inference time with no additional training required, offering a flexible and precise alternative to often ambiguous text prompts or coarse class labels. Our approach achieves high-quality, diverse generations on ImageNet and COCO. Code is available at https://github.com/valeoai/REPA-G

Sources

Related papers