Structure over Pixels: Learning Variable-Length Visual Programs
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Structure over Pixels".
Jane: Discrete visual tokenizers (DVTs) translate images into ordered sequences of discrete codes,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, focusing on the title and who wrote this—"Structure over Pixels: Learning Variable-Length Visual Programs"—it really highlights their focus on structure first. They aren't just trying to get better pixel reconstruction anymore; they want a representation that captures the scene's organization.
Jane: And the authors are from Poznan University of Technology, which tells us this is coming from a solid research group looking at discrete visual tokenizers and how they can be made more flexible.
Lu: What's striking is that they introduce an auxiliary length head that predicts the active prefix length in a single forward pass, which fundamentally changes how we think about image encoding.
Meng: That prediction mechanism sounds powerful, but I wonder if the quality of that initial prediction really dictates the overall performance or if it just adds another layer of complexity to train.
Lalam: I think that length head is key because it solves the problem where fixed-budget tokenizers just cut off information arbitrarily when a scene is complex.
The paper's summary: Tom: So, summarizing what they actually did in "Structure over Pixels: Learning Variable-Length Visual Programs," they built an architecture called STROP that uses a frozen DINOv3 encoder to get patch features and then maps those into a sequence of tokens through a vector quantization bottleneck.
Jane: The paper explains that STROP forms structural scene representations by learning this variable length program, which means the sequence length L changes based on the image content, not some arbitrary setting.
Lu: They achieved this by bypassing pixel-level reconstruction gradients and using frozen DINOv3 features for supervision, which is a clever way to guide the learning toward structural quality instead of just making sure every pixel matches perfectly.
Meng: Bypassing those specific gradients sounds like they are deliberately focusing the model on higher-level concepts rather than getting bogged down in low-level texture matching, which makes sense for compression goals.
Lalam: The summary emphasizes that the program length grows with scene complexity, which is a big deal because it means a simple sky and a crowded street will result in sequences of different lengths.
The paper's improvements: Tom: What really stands out about the proposed improvements is their four-phase training curriculum, which they use to guide the learning process from random truncation all the way to self-predicted lengths.
Jane: That curriculum seems very systematic; Phase one forces important content to be at the start, and later phases refine this by using oracle estimations derived from rate–feature–distortion probes against DINOv3 features.
Lu: The training objectives are also quite layered, combining latent alignment loss that preserves teacher feature geometry with commitment loss for vector quantization regularization and a diversity loss to prevent codebook collapse.
Meng: Those losses sound like they are working hard to keep the learned codes meaningful and not just random noise in the embedding space, which is important when you're trying to get practical results.
Lalam: The combination of preserving teacher geometry with explicit regularization against codebook collapse suggests they are building a very robust system that resists degradation as it learns these variable programs.
Conclusion: Tom: To wrap up, the paper concludes that STROP preserves substantially more task-relevant content than tokenizers trained with pixel-leaning reconstruction objectives, and they show that program length tracks scene complexity, with stronger correlations observed on CLEVR where longer programs are allocated to richer scenes.
Jane: That means the final result is a representation of the scene that is inherently better at handling tasks because it captures the structure more effectively than what we see in simpler tokenization methods.
Lu: The analysis also shows that they can expose coarse layout information to a generalist LLM more readily than fixed attribute bindings like color or material, which opens up new ways for language models to understand visual scenes.
Meng: I'm interested in the practical implication of this for compression; if you can decouple reconstruction loss from pixel fidelity, they demonstrated ratios up to thirteen thousand times, suggesting a major path toward efficient vision systems.
Lalam: I think the most significant impact is how this variable-length approach could enable AI systems to dynamically allocate computational resources based on the scene's complexity during inference.
Institute of Computing Science, Poznan University of Technology
cs.CV, cs.LG
Submitted: 2026-05-26
Updated: 2026-09-30
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: Discrete visual tokenizers (DVTs) translate images into ordered sequences of discrete codes, but existing adaptive tokenizers often struggle to learn a continuous per-image sequence length coupled to
Key concepts
- Structural Scene Representations
- This refers to learning a way to represent an image that focuses on its underlying structure rather than just individual pixels. STROP aims to create these representations by mapping visual features into variable-length sequences, allowing the model to capture scene complexity naturally through program length.
- Variable-Length Visual Programs
- Instead of using a fixed number of tokens for every image, STROP generates programs whose length grows as the scene becomes more complex. This means richer scenes get longer descriptions, providing a more principled way to encode visual information than fixed-budget tokenizers.
- Auxiliary Length Head
- This is a specific part of the STROP architecture that predicts how long the generated visual program should be. It works by taking pooled features from the program generator and using a sigmoid function to output an estimate of the active prefix length in one step during inference.
Terminology
Summary
Discrete visual tokenizers (DVTs) translate images into ordered sequences of discrete codes, but existing adaptive tokenizers often struggle to learn a continuous per-image sequence length coupled to both the model and scene. This paper proposes STROP, a novel DVT architecture that forms structural scene representations while simultaneously learning the variable length of an image's visual program. By bypassing pixel-level reconstruction gradients and utilizing frozen DINOv3 features for supervision, STROP enables the generation of Variable-length visual programs
where program length grows with scene complexity, offering a more principled structural representation than fixed-budget tokenizers.
Proposed Architecture
STROP is a DVT architecture comprising five stages: a frozen visual encoder producing patch-wise features, a program generator that maps those features into a sequence of tokens, vector quantization, an interpreter that expands the program back to a spatial feature grid, and finally, a convolutional decoder. The core innovation lies in the auxiliary length head which predicts the active prefix length in a single forward pass: Lˆ = Kσhϕ(pool(Ze))
. This allows for variable-length programs at inference without autoregressive sampling.
Training Curriculum
The model is trained using a four-phase curriculum supervised by local rate–distortion probes against frozen DINOv3 features. The phases are designed to guide the learning process:
-
Phase 1: Random truncation, which
forces informative content to be front-loaded into earlier tokens.
-
Phase 2: Oracle estimation, where targets are estimated using reconstruction error and local slope of the rate–feature–distortion curve, leading to a corrected target length L˜i.
-
Phase 3: Supervised training, where the length head is trained against these oracle targets.
-
Phase 4: Handoff to predicted lengths, which
smoothly transition[s] from random to self-predicted truncations.
Key Training Objectives
The full training objective is a composite loss function: L = Llat + Lcommit + Ldiv + Llen,
where the terms are defined as follows:
Latent alignment (Llat)
Given interpreted patches Fˆ and frozen DINO patches F⋆, the loss is calculated using a cosine term to preserve teacher-feature geometry and an MSE term: Llat = 1 − (1/P)PP p=1 cos(Fˆp, F⋆p) + MSE(F, F ˆ ⋆)
.
Commitment loss (Lcommit)
This standard VQ regularization penalty keeps prequantized embeddings close to their selected codebook entries: Lcommit. The standard VQ regularization penalty λq∥Zp − sg[Zq]∥2 keeps prequantized embeddings close to their selected codebook entries.
Diversity loss (Ldiv)
A utilization regularizer penalizes nonuniform code usage and prevents codebook collapse
: Ldiv. A utilization regularizer (λdiv=0.3), warmed up over 40k steps, penalizes nonuniform code usage and helps prevent codebook collapse.
Evaluation and Findings
STROP is evaluated through structural diagnostics, including program structure analysis where token erasure maps are compared against semantic masks to identify compositional handles over the scene.
Downstream probes test how much task information is preserved:
-
Program structure analysis reveals that
Active pairs are compared against random active pairs and inactive-token controls
to find localized, reusable effects. -
Code reuse and unsupervised readouts compare results from DINO patches, interpreted STROP patches, erasure-attribution vectors, and code-to-class mappings to test for non-random program structure.
-
Supervised downstream probes (e.g., mIoU on segmentation) show that the
interpreted STROP field consistently outperforms both adaptive tokenizer baselines on the tasks that test representational content.
The paper concludes that STROP preserves substantially more task-relevant content than tokenizers trained with pixel-leaning reconstruction objectives,
suggesting that pairing a DINO-aligned latent objective with a variable-length program interface is superior for structural representation. Furthermore, analysis shows Program length tracks scene complexity,
with stronger correlations observed on CLEVR where longer programs are allocated to richer scenes. Finally, the codebook structure exposes coarse layout to a generalist LLM more readily than they expose fixed attribute bindings such as color or material.
Limitations and Future Work
The authors note that the readability of the raw program
remains a limitation, as direct probes on quantized tokens trail interpreter-mediated probes. Future work is suggested in three directions: scaling the codebook and program length, developing program-native probing protocols
to isolate prefix order encoding, and using grammar inference over learned sequences to make the codebook a usable symbolic interface for language models.
Improvements for AI systems
Here are specific improvements for AI systems based on the STROP architecture described in the paper:
-
Mend structural ambiguity in visual representations by learning variable-length visual programs that correlate directly with scene complexity (object count/scene richness).
-
Enable robust, interpretable analysis of learned latent representations by allowing program tokens to act as
semantic handles
over scenes, enabling researchers to inspect code reuse and compositional structure directly from the token sequence. -
Improve downstream task performance (segmentation and depth estimation) by using an interpreter that reconstructs spatial features from the variable-length program, which is shown to consistently outperform both frozen teacher features and off-the-shelf adaptive tokenizers like FlexTok.
-
Achieve higher compression ratios in visual tokenization by decoupling reconstruction loss from pixel fidelity; the system learns structural representations guided by DINOv3 features rather than solely texture, allowing for extreme compression (up to 13,000x ratio demonstrated).
-
Develop an end-to-end training pipeline that autonomously determines the optimal sequence length for any given image through a curriculum learning approach, moving beyond fixed codebook budgets to adapt computation dynamically based on scene content.
-
Create a more transparent and usable symbolic interface for Large Language Models (LLMs) by producing discrete codes that encode coarse spatial layout information, enabling LLMs to recover object counts and positions accurately, even if they cannot reliably decode fine-grained attributes like color or material without external priors.
-
Establish a principled method for analyzing compositional structure in visual data by examining pairwise token synergies, which reveal how specific pairs of program tokens induce localized semantic effects aligned with coherent scene regions.
Abstract
Discrete visual tokenizers map images to ordered sequences of tokens, providing a natural representation for structural scene descriptions. Most use a fixed sequence length, while adaptive methods often require post-hoc search or choose among a small set of rates that control the length. We propose STROP, a discrete tokenizer that learns both a visual program and its image-dependent active length. A length head is trained with a four-phase curriculum using local rate-distortion probes against frozen DINOv3 features, then predicts the active prefix in a single forward pass. At a matched rate of about 250 nominal bits per crop, the adaptive model improves segmentation over a separately trained fixed-length baseline on four benchmarks (by 1.6 - 3.1 mIoU), and it also beats a fixed K = 32 baseline that uses more bits. STROP programs also yield higher segmentation mIoU than FlexTok, One-D-Piece, and ALIT at similar or higher rates, under the same readout architecture and training protocol. STROP therefore learns useful per-image sequence lengths without post-hoc search or a predefined set of compression rates.
Sources
- Illiterate DALL-E Learns to Compose
- Bridging the Gap to Real-World Object-Centric Learning
- Neural Programmer-Interpreters
- From Principles to Applications: A Comprehensive Survey of Discrete Tokenizers in Generation, Comprehension, Recommendation, and Information Retrieval
- An Image is Worth 32 Tokens for Reconstruction and Generation
- FlexTok: Resampling Images into 1D Token Sequences of Flexible Length
- One-D-Piece: Image Tokenizer Meets Quality-Controllable Compression
- ElasticTok: Adaptive Tokenization for Image and Video
- Adaptive Length Image Tokenization via Recurrent Allocation
- CAT: Content-Adaptive Image Tokenization
- InfoTok: Adaptive Discrete Video Tokenizer via Information-Theoretic Compression
- BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers
- Representation Alignment for Generation: Training Diffusion Transformers Is Easier Than You Think
- Masked Autoencoders Are Effective Tokenizers for Diffusion Models
- DINOv3
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
- Finite Scalar Quantization: VQ-VAE Made Simple
- Restructuring Vector Quantization with the Rotation Trick
- Spherical Leech Quantization for Visual Tokenization and Generation
- SoftVQ-VAE: Efficient 1-Dimensional Continuous Tokenizer
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models