Mask-supervised Object-centric Representation Learning with LeJEPA

arXiv:2607.02404 · cs.CV, cs.LG · Submitted 2026-07-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Mask-supervised Object-centric Representation Learning with LeJEPA".

Tom: Object-centric learning promises better data efficiency than image-level learning by aligning representations at the level of objects rather than whole scenes,

Jane: First, who's behind it and why it matters.

Paper summary: Tom: So, to wrap up our discussion on "Mask-supervised Object-centric Representation Learning with LeJEPA," we’ve covered how this method tackles the stability issues in object learning by leveraging fixed masks during training and how it achieves strong performance across various benchmarks. Jane And we also explored the big idea behind shifting the focus from image-level alignment to object-level alignment, which is what sets this paper apart in terms of data efficiency. Lu The authors introduced a framework that uses an object-centric LeJEPA loss and an instance-level contrastive loss to create distinct semantic and instance representations for every entity in a scene.

Meng: From my standpoint, the most important thing is that the performance gains are achieved while training on just ten percent of COCO data, which makes it highly scalable for practical applications where data acquisition is a major bottleneck <ref:2607.02404#pg0>. Lalam I really see this as empowering future AI development by showing that we can build powerful perception systems without needing access to those enormous training datasets anymore.

Tom: It seems the authors have successfully developed a way to create object-centric representations that are both robust and highly data efficient for vision tasks, moving beyond just learning from whole images. Jane The title itself really summarizes the essence: masking supervision combined with an object-centric approach to representation learning using LeJEPA.

Lu: This work suggests that decomposing a scene into its constituent objects and learning unique representations for each one is a viable path forward for more intelligent scene understanding in AI systems.

Meng: I'm just thinking about the long term, how this capability will influence how we deploy vision AI in things like autonomous vehicles or remote monitoring systems where data constraints are severe.

Lalam: I believe this research paves the way for creating more nuanced and context-aware AI that understands the world as a collection of interacting objects rather than just a stream of pixels.

Conclusion: Tom: So, we've spent some time diving into how this paper tackles object learning by focusing on masks instead of whole images. Jane, what are your initial thoughts on the core idea behind "Mask-supervised Object-centric Representation Learning with LeJEPA"?

Jane: I think it makes sense because it shifts the focus to individual objects, which is a much more intuitive way for a model to understand a scene, and that's what this paper achieves by aligning representations at the object level. Lu, from your perspective as someone working on complex AI structures, does this object-centric view unlock any fundamentally new ways we can structure visual understanding?

Lu: Absolutely! Moving away from global image alignment to local object alignment opens up possibilities for truly compositional reasoning. Imagine an AI that can track a specific person across multiple scenes because it learns what that person *is*, not just what the whole background looks like. That level of detail is incredibly fertile ground for novel applications.

Meng: I'm curious about the practical side, though; if this method is so efficient, how does that translate into something we can actually deploy reliably in a real-world system? We need to know what its real-world performance looks like beyond just the numbers on paper.

Lalam: From a model perspective, I see this as significantly enhancing the AI's ability to build robust internal concepts. When you train it on ten percent of the data and it still performs well, it means the learned representations are much more general and less prone to overfitting on specific visual patterns, which is fantastic for building resilient cultural models.

Tom: That's a huge point about resilience; the idea that we can get high performance with less raw data is compelling. Jane, how should we explain this whole concept of object-centric LeJEPA to someone who isn't deeply familiar with deep learning architectures?

Jane: I think we can put it simply by saying the model stops looking at every pixel in a scene and starts focusing on defining what each distinct item in that scene is, creating a unique identity for every object. It’s like moving from describing a whole painting to meticulously cataloging and understanding each brushstroke as its own distinct element.

Lu: And that distinction between semantic object representations and instance-level representations, which the paper uses, suggests we're getting two different layers of understanding—what it is generally, and exactly which specific thing it is. That dual approach is really intriguing for complex tasks.

Meng: I just want to make sure we aren't over-promising on deployment; while the efficiency numbers are impressive, the real challenge will be fine-tuning these object representations for very specific industrial uses where precision is everything.

Lalam: And if we look at the implications for our future models, this suggests a path toward more compact and contextually aware AI systems that don't require massive datasets to build deep perceptual knowledge.

Tom: It sounds like this paper isn't just another incremental tweak; it presents a fundamentally different way to learn visual representations. Jane, what’s your take on the title itself—"Mask-supervised Object-centric Representation Learning with LeJEPA"?

Jane: I think the title perfectly captures the essence because it highlights both the specific technique used, LeJEPA, and the main innovation: moving from image alignment to object alignment guided by mask supervision.

Lu: Exactly! The combination of those elements tells a whole story about how they solved a major hurdle in scene understanding by changing where they applied their learning mechanism.

Meng: It sounds like the authors really nailed the technical implementation, which is always impressive when you're dealing with these kinds of complex representation spaces.

Lalam: For me, the implication is that we can build more nuanced and interpretable AI because we can actually see and understand what specific objects an AI is focusing on when it makes a decision.

Tom: It’s clear this work lays a solid foundation for future research in vision, moving us toward systems that think about scenes as collections of meaningful entities. But where do we go from here?

Jakob Geusen, Ender Konukoglu

ETH Zurich

cs.CV, cs.LG

Submitted: 2026-07-02

Updated: 2026-10-05

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 82/100

The gist: Object-centric learning promises better data efficiency than image-level learning by aligning representations at the level of objects rather than whole scenes, which allows models to exploit

Key concepts

Object-centric Learning
This approach focuses representation learning at the object level rather than the entire image. Instead of aligning features across whole scenes, it aligns features specifically for individual objects. This allows the model to better capture compositional structures and relate instances across different scenes effectively.
LeJEPA
LeJEPA is a framework that learns representations using a combination of semantic and instance spaces. It uses object-centric LeJEPA loss to shape a semantic space, while an instance-level contrastive loss shapes an instance space, creating richer object understanding.
LObjectLeJEPA
This is the primary loss function that shapes the semantic object representations. It encourages representations of the same object to be similar across different augmented views, using cross-view alignment. A regularization term (LSIGReg) prevents these representations from collapsing into a single point.
Instance-level Contrastive Loss (Linstance)
This loss ensures that all patches belonging to the same object agree on a common instance representation while remaining distinct from other objects. It uses a supervised contrastive formulation to enforce this separation and agreement at the patch level.

Terminology

Summary

Object-centric learning promises better data efficiency than image-level learning by aligning representations at the level of objects rather than whole scenes, which allows models to exploit compositional structures and connect instances across many scenes. The gist: Object-level LeJEPA outperforms image-level LeJEPA on tracking (DAVIS), classification (ImageNet-1k), segmentation (ADE20k), and re-identification (NAVI) across two model scales and 10–100% of COCO.

How it works

The method extends LeJEPA by moving the alignment and regularization term from the image to the object level, enabled by using cheap, off-the-shelf SAM proposals for object masks during training. This approach sidesteps instability inherent in joint partitioning and representation learning by accepting fixed partitioning during training in exchange for far fewer restrictions on the representation space.

The framework learns representations in two complementary spaces: a semantic space shaped by an object-centric LeJEPA loss, and an instance space shaped by a contrastive loss.

  1. Semantic Object Representations: A semantic object representation, denoted as a function of patch features, is computed as a weighted mean of independently projected patch features based on the patch-wise average-pooled object mask, followed by cross-attention to mask out external keys and values before passing through a residual MLP. This representation is shaped by the object-centric LeJEPA loss LObjectLeJEPA, which uses cross-view alignment to encourage representations of the same object across different augmented views to be similar, while preventing collapse via the regularization term LSIGReg.

  2. Instance-level Object Representations: The instance-level object predictions are obtained by applying an MLP, denoted as ϕ, to each patch feature before L2 normalization: yn,v,i = ϕ(g(xn,v)i). This representation is shaped by the instance-level contrastive loss Linstance, which uses a supervised contrastive formulation to ensure that all patches of the same object to agree on a common instance representation while remaining separable from those of other objects.

Key Training Components

The total training objective is the sum of two object-centric losses:

  1. LObjectLeJEPA: This loss combines an alignment term (Lpred) and a regularization term (LSIGReg), defined as LObjectLeJEPA = Lpred + λLeJEPA LSIGReg, where Lpred encourages representations across different views of the same object to be similar, and LSIGReg prevents collapse by regularizing the projected object representations.

  2. Linstance: This loss targets instance-specific structure within a single view, defined as l instance = −1/P(i) Σ p∈P(i) log exp(y n,v,p/τ) P a∈A(i) exp(y n,v,a/τ). This term ensures that patches sharing a dominant mask agree on an instance representation while being separable from those of other objects.

Evaluation and Findings

The performance of Object LeJEPA was evaluated across various downstream tasks using frozen patch features.

  1. Instance-Awareness Probing: When clustering normalized patch features with K-Means, Object LeJEPA not only lead[s] the COCO-trained models but overtook DINOv3, achieving a FG-ARI of 0.431 and mIoU of 0.355 on ADE20k, and reaching 0.915 accuracy and 0.954 AUC in the same quadratic probe task used to predict same-instance pairs, surpassing DINOv3 (0.914 / 0.957).

  2. Dense Semantic Tasks: On linear-probe semantic segmentation of ADE20k, Object LeJEPA attained the best mIoU among the COCOtrained models (0.418), edging out SlotMIM and improving over Image LeJEPA (0.339).

  3. Object-level Tasks: For object re-identification on NAVI, Object LeJEPA was the strongest model, reaching a balanced accuracy of 0.250 for top-1 classification and 0.420 for top-5 classification when using its native semantic object representation (z).

Data Efficiency and Generalization

A central finding is the data efficiency of the method: trained on only 10% of COCO, Object LeJEPA with a ViT-B architecture already matched image-level LeJEPA trained on the full COCO dataset on every task. This demonstrates that object-level alignment recovers performance from ten times less data compared to image-level alignment.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements to AI systems that can be derived from the Object-centric LeJEPA framework:


)Object-centric LeJEPA System Improvements

The core improvement lies in shifting representation learning from global scene alignment to object-level alignment, enabling superior data efficiency and enhanced instance discrimination. The improved system can achieve the following specific capabilities:

)Specific Capabilities of the Improved AI System:

  1. For a given image, the system can perform high-accuracy object detection and segmentation by leveraging object-level representations extracted via SAM proposals during training, leading to sharper object boundaries compared to prior methods like Image-level LeJEPA.

  2. The system exhibits significantly improved instance discriminability across different views and scenes (e.g., out-of-distribution objects like fish). It can accurately distinguish between instances of the same object category (intra-image) and separate them from other objects, even when those objects are unseen during training, by leveraging the instance-level contrastive loss.

  3. The system achieves superior performance in tracking and re-identification tasks (e.g., DAVIS and NAVI). It can maintain temporal stability in video sequences by propagating object identities across frames (tracking) and accurately re-identify the same object across diverse images with varying backgrounds, poses, and lighting (re-identification).

  4. The system is highly data-efficient. It achieves performance comparable to or exceeding image-level LeJEPA trained on 10% of the training data, meaning it requires drastically less labeled or collected data to achieve strong generalization across various downstream tasks.

  5. The system can perform robust classification and semantic segmentation by effectively capturing both object identity (semantic space) and instance-specific details (instance space), leading to better performance in dense prediction tasks on datasets like ADE20k compared to models relying solely on global image tokens.

Sources

Related papers