COMiT: Learning Structured Visual Tokens through Sequential Communication

arXiv:2602.20731 · cs.CV, cs.AI, cs.LG · Submitted 2026-02-24 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "COMiT: Learning Structured Visual Tokens through Sequential Communication".

Jane: Discrete image tokenizers are crucial for modern vision systems, but existing methods often optimize for reconstruction and compression, yielding tokens that capture local texture rather than object-level semantic structure.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: We've touched on what the paper is about, so now let's get into the specifics of COMiT: Learning Structured Visual Tokens through Sequential Communication. The authors are Davtyan, Sahin, Haghighi, Stapf, Acuaviva, and Favaro.

Jane: That sounds like a solid team of researchers tackling a deep problem in vision systems. What should we be focusing on when we look at the paper's summary?

Lu: The summary really hammers home that the goal is to construct a latent message within a fixed token budget by observing localized crops and recurrently updating it. This sequence of observations is key because it’s how they move from just seeing pixels to building a structured description.

Meng: It sounds like the paper is saying that by following this iterative observation and update process, the final message conditions a flow-matching decoder that reconstructs the full image, which connects the encoding and decoding parts in a unified way.

Lalam: I think what’s exciting here is how they’re moving away from just local details to capture high-level entities and their relations, because this structured approach should make our downstream applications much more reliable.

The paper's summary: Tom: So, looking at the summary again, the core mechanism involves an encoding process where the model acts as both speaker and listener, which they achieve through two main principles: attentive and sequential tokenization.

Jane: Can you break down what that means for us in simpler terms when we talk about those principles? I want to make sure everyone gets the mechanics.

Lu: Attentive and sequential tokenization means the encoder processes the image as a sequence of localized observations, where it pays attention to different regions at each step while incrementally adding information into that discrete latent message. It's like scanning a scene piece by piece while keeping track of what you’ve seen before.

Meng: That sounds like the model is building up context layer by layer, which should naturally help it organize the visual data into meaningful parts instead of just dumping raw pixel data into a vector.

Lalam: I agree, that sequential refinement implies that the final token sequence isn't random noise; it's built step-by-step based on a structured interaction with the image input.

The paper's improvements: Tom: Moving on to what they suggest are the specific improvements in COMiT, it seems they focus heavily on how this communication structure helps them achieve better results compared to previous tokenization methods.

Jane: What are these key design principles that actually lead to those suggested improvements? I want to understand the mechanism behind why it’s better than just a standard autoencoder setup.

Lu: The paper highlights homogeneous communication as a major point, contrasting it with traditional autoencoders that separate the encoder and decoder into different networks. COMiT uses one unified network for both roles, which is crucial because it forces the encoding and decoding to be intrinsically linked in their learning process.

Meng: That unified design is important from an engineering standpoint because it prevents the encoder from learning a representation that is only good for compression, forcing it to maintain semantic integrity throughout the entire flow.

Lalam: I think this unification means they are achieving a better balance between reconstruction quality and semantic organization simultaneously, which is something we’ve struggled with before.

Conclusion: Tom: Alright, so let's wrap up the discussion on COMiT: Learning Structured Visual Tokens through Sequential Communication. The main implication here is that this framework yields structured discrete token sequences that are more interpretable and object-centric than what prior methods produced, which is really significant for multimodal architectures.

Jane: It sounds like the ultimate goal is moving toward visual representations where tokens directly correspond to meaningful objects or object parts, rather than just patches of texture.

Lu: Indeed, the paper shows that this communication-inspired sequential encoding induces more interpretable and object-centric tokens because of how information is distributed across those tokens.

Meng: From a practical standpoint, having these structured sequences means that when we use COMiT for reasoning tasks, we should expect much better performance in compositional generalization and relational reasoning compared to what we see now.

Lalam: I think the existence of this structured latent space is incredibly promising for our culture because it suggests that future AI can operate with a deeper understanding of the visual world's composition.

Tom: That’s a lot to digest, but overall, COMiT really points toward a more grounded way to create visual tokens for complex AI tasks. We have covered the title, the summary, and the specific design improvements in this paper today.

Jane: It’s been fascinating seeing how they framed communication as a training mechanism rather than just an architectural choice. We’re ready now to see what other papers are out there.

Lu: I'm looking forward to seeing how others try to adapt this sequential reasoning idea into different domains, perhaps even temporal ones, which is where the real creative potential lies.

Meng: I’m curious about the practical challenges of scaling this up; how does that iterative process actually run efficiently on massive datasets compared to what we're used to?

Lalam: It’s exciting because it gives us a new way to think about building AI that doesn't just memorize data but learns a structured way to understand visual concepts.

Aram Davtyan, Yusuf Sahin, Yasaman Haghighi, Sebastian Stapf, Pablo Acuaviva, Alexandre Alahi, Paolo Favaro

University of Bern, Switzerland

cs.CV, cs.AI, cs.LG

Submitted: 2026-02-24

Updated: 2026-09-29

Comments: Project website: https://araachie.github.io/comit/

Code: https://github.com/Araachie/comit

Project page: https://araachie.github.io/comit

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Discrete image tokenizers are crucial for modern vision systems, but existing methods often optimize for reconstruction and compression, yielding tokens that capture local texture rather than

Key concepts

Discrete image tokenizers
These are methods used in vision systems that create discrete tokens from images. Existing methods often focus on reconstruction and compression, resulting in tokens that capture local texture instead of high-level semantic structure.
Attentive and sequential tokenization
This encoding process involves the model processing an image as a sequence of localized observations. It pays attention to different regions at each step while incrementally adding information into a discrete latent message, building context layer by layer.
Homogeneous communication
COMiT uses one unified network for both the encoder and decoder. This contrasts with traditional autoencoders that separate these networks, forcing the encoding and decoding processes to be intrinsically linked during learning.
Structured visual tokens
The goal is to move beyond local details to capture high-level entities and their relations. These structured tokens are more interpretable and object-centric than previous methods, suggesting future AI can operate with a deeper understanding of visual composition.

Terminology

Summary

Discrete image tokenizers are crucial for modern vision systems, but existing methods often optimize for reconstruction and compression, yielding tokens that capture local texture rather than object-level semantic structure. This paper introduces Communication-Inspired Tokenization (COMiT), a novel framework designed to learn structured discrete visual token sequences by framing the encoding process as an iterative communication-and-reconstruction game inspired by human communication. COMiT aims to induce interpretable, object-centric token structure and substantially improve compositional generalization and relational reasoning over prior methods.

How it works

COMiT constructs a latent message within a fixed token budget by iteratively observing localized image crops and recurrently updating its discrete representation. The core mechanism involves an encoding process where the model acts as both speaker and listener, mirroring human communication symmetry. This is achieved through two key design principles:

  1. Attentive and sequential tokenization: The encoder processes the image as a sequence of localized observations, attending to different regions at each step and incrementally integrating information into a discrete latent message, defined by the update rule:

mk = f mθ(ck, tk, ak, mk-1).

  1. Homogeneous communication: Unlike traditional autoencoders with separate encoder and decoder networks, COMiT uses a unified design where the same network performs both encoding and decoding.

Training Objectives

The entire pipeline is trained end-to-end using a combination of flow-matching reconstruction and semantic representation alignment losses. Specifically, the training loss is formulated as:

L = LFM + λREPALREPA + λSREPALSREPA (Equation 7).

**: LFM is the flow-matching loss, which utilizes a unified, differentiable flow objective to train both encoding and decoding stages. The semantic representation alignment objective (SREPA) encourages grounding by distilling high-level features from a frozen self-supervised vision model. Crucially, the paper argues that while alignment provides semantic signal, the attentive sequential tokenization determines how this information is distributed and localized across tokens. 3.2 Distilling Semantic Representation details this by projecting intermediate representations into the distilled semantic space using a loss based on cosine similarity: LSREPA = exp −Sim ψ(x),¯f mθ[j](xt, t, ag, mK). 5. Conclusion and Discussion notes that attentive tokenization is critical for inducing interpretable, object-centric tokens. 4.1 Ablations further reveal that communication-inspired sequential encoding induces more interpretable, object-centric tokens. The final training loss combines these elements: L = LFM + λREPALREPA + λSREPALSREPA (Equation 7). 5. Conclusion and Discussion also highlights the trade-off: COMiT consistently outperforms prior work on semantic tasks. This highlights a representation-reconstruction trade-off that contrasts with the conventional compression-reconstruction trade-off targeted in prior work. 4.2 Quantitative Results show COMiT consistently outperforms prior work on semantic probing across ImageNet1k, MSCOCO, and Visual Genome. 7. Communication-Inspired Tokenization for Structured Image Representations concludes that "structured discrete token sequences such as those learned by COMiT provide a promising interface for multimodal architectures, especially in settings where object-centric reasoning and compositional understanding are critical. 8. Future work includes extending COMiT to video, where temporal redundancy and long-range structure pose additional challenges and opportunities for discrete representation learning. The paper also explores cropping policies, noting that COMiT also naturally supports test-time scaling: adding local crops provides modest gains on compositional generalization and inter-object relations. 16. These findings are further supported by qualitative analysis showing the Emergence of Objectness, where token attention maps in COMiT-XL show tokens tend to correspond to some semantically meaningful regions in images, such as objects or object parts." 17. The model's latent space is also shown to be well-structured, as nearest neighbors typically share semantics. 18. The authors note that scaling the model from B to L improves both reconstruction and representation, while further scaling from L to XL allocates the additional capacity to aid reconstruction, while diminishing the semantics. 19. The architecture is based on DiT (DiT), utilizing AdaLN layers conditioned on timestamps, and utilizes a vocabulary size of 64000. 12. A key implementation detail is that different modalities (image, message) use separate projections in the AdaLN layers. 13. This paper demonstrates that COMiT's sequential refinement process allows for the progressive ambiguity reduction visible when decoding messages with more aggregated crops, which is inherently compositional. 14. The model's performance on compositional generalization (MSCOCO) and relational reasoning (Visual Genome) is superior to baselines, indicating that the tokenization pipeline is key to structuring messages. 15.

Improvements for AI systems

Based on the provided scientific paper, here are specific improvements for AI systems derived from the COMiT framework, along with a description of what these improved systems can achieve:


),

  1. Improved Semantic Understanding and Object-Centric Reasoning:

The COMiT tokenization approach is designed to induce object-centric tokens through its attentive sequential encoding process. This means that instead of tokens capturing local texture (like traditional methods), they are organized hierarchically, reflecting object-level semantic structures.

  1. Enhanced Compositional Generalization:

By structuring the latent message incrementally via communication-inspired tokenization (iterative observation and refinement), the system learns to represent objects and their relationships as part of a coherent sequence rather than isolated features.

The improved AI system can:

  • Accurately recognize novel object combinations that it has not seen during training (as tested by the MSCOCO compositional generalization benchmark).

  • Reason about relational semantics (e.g., subject-predicate-object) in complex scenes with higher fidelity, as evidenced by its performance on Visual Genome.

  1. More Interpretable Visual Representations:

The framework explicitly enforces structured token sequences through the attentive and sequential tokenization principle. This allows researchers to analyze which tokens correspond to specific objects or parts of objects (as quantified by attention map analysis in Section C).

The improved AI system can:

  • Be used for tasks requiring explainability, such as medical image diagnosis or autonomous driving perception, where the system must justify its decisions by referencing specific visual entities.
  1. Adaptive and Task-Dependent Visual Tokenization:

COMiT's flexibility regarding cropping policies (global vs. local crops) suggests that the model can be adapted to different input modalities or specific reasoning tasks without full retraining at inference time.

The improved AI system can:

  • Perform efficient, on-the-fly adaptation to new visual contexts (e.g., switching between fine details and global scene context) by dynamically selecting the most relevant observation strategy for a given task, potentially accelerating real-time performance in dynamic environments like robotics or video analysis.
  1. Unified Generative and Encoding Pipeline:

The use of a single transformer model for both encoding (via iterative message updates) and decoding (via flow-matching) simplifies architecture and allows the system to learn the necessary reconstruction fidelity simultaneously with semantic organization.

The improved AI system can:

  • Achieve high-fidelity image synthesis that is inherently grounded in structured, meaningful semantic tokens, bridging the gap between generative quality and structural interpretability more effectively than separate encoder/decoder approaches.
  1. Improved Robustness to Training Data Shifts (Domain Generalization):

The application of semantic distillation (SREPA) using frozen self-supervised vision model features (like DINOv2 [CLS] tokens) helps anchor the learned tokens to robust, high-level visual features.

The improved AI system can:

  • Maintain strong performance and semantic integrity when applied to out-of-distribution datasets, making it more reliable in real-world scenarios where training data distribution shifts are common.

Abstract

Discrete image tokenizers provide a sequential interface for vision and multimodal models, but are typically optimized for reconstruction or compression and therefore tend to encode local appearance rather than object-level structure. We introduce COMiT, a communication-inspired framework for learning structured discrete visual representations. COMiT constructs a fixed-length latent message through sequential visual observations: at each step, a transformer processes a localized image crop and updates, refines, and reorganizes the existing token sequence. After several iterations, the resulting message conditions a flow-matching decoder that reconstructs the complete image. The encoder and decoder are implemented within a single transformer and trained end-to-end using flow-matching reconstruction and semantic representation-alignment objectives. COMiT substantially improves compositional generalization and relational reasoning over prior methods. Our experiments show that, while semantic alignment helps ground the representation, attentive sequential tokenization is critical for inducing more interpretable, object-centric token structures.

Sources

Related papers