Learned Image Compression for Vision-Language-Action Models
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Learned Image Compression for Vision-Language-Action Models".
Jane: The gist: SPARC is a learned image compression framework tailored for VLA systems that adaptively allocates bitrate across spatial regions according to their contribution to downstream control,
Tom: First, who's behind it and why it matters.
Title and authors: Tom: The main problem they are addressing is that vision language action models rely on high-frequency multi-camera observations, which creates a major bottleneck for real-time control in environments with limited bandwidth or when robots are distributed.
Jane: They point out that existing image and video codecs focus on generic visual fidelity rather than the actual performance of the AI policy they're trying to support.
Lu: The central observation here is that the importance of visual information changes depending on both the camera view and different spatial regions within an image, which means uniform bitrate allocation across everything is inefficient.
Meng: So, where one part of the picture—like the gripper or a target object—matters way more than another region, SPARC tries to figure that out automatically.
Lalam: SPARC employs a learned framework that performs task-aware bitrate allocation directly within the latent space of a neural image codec.
Tom: They build it on top of a pretrained neural image codec and add this lightweight temporal mask selector that predicts spatial latent masks from temporal context, allowing for adaptive allocation across both camera streams and regions.
Jane: The mechanism involves taking a history of hyperprior features and using them to generate these binary masks, which are then used to selectively mask the latent entries before they are stored or transmitted.
Lu: They train this entire system end-to-end to optimize downstream action quality under a specific communication budget, which is pretty ambitious.
Meng: To do that, they go through two training phases: first warming up the decoder while freezing everything and disabling masking, and then joint optimization with a tilted rate loss to keep things stable.
Tom: That tilted rate loss is key because it prevents the masking policy from becoming too biased toward regions that naturally have high local bitrate.
Jane: It’s designed to stop the system from over-suppressing task-critical visual patterns during the learning process, which is a really clever way to stabilize things.
The paper's summary: Lu: One major improvement they bring is moving beyond fixed compression rates, which is a weakness in many prior learned compression methods that are only trained for one rate at a time.
Tom: They introduce this variable-rate compression method where a single model is trained to support multiple rates simultaneously, which is much more flexible for real-world scenarios.
Jane: They also introduced the specific tilted rate loss function, which they use to stabilize the training process by managing the trade-off between data size and action quality during optimization.
Meng: That loss function prevents the entropy-based objectives from over-suppressing those statistically rare yet task-critical visual patterns that we know are important for control.
Lalam: This stabilization is what allows them to achieve better bitrate success tradeoffs than conventional codecs, meaning they can get more useful information out of the same amount of data sent.
Tom: So, in summary, the improvement isn't just a better compression algorithm; it’s a smarter allocation strategy that is trained specifically to serve the downstream control policy.
Jane: It really shifts the focus from just making the image look good to ensuring that every bit sent contributes meaningfully to the robot’s ability to perform its actions.
Lu: They show this leads directly to superior bitrate-success tradeoffs across various benchmarks, which is a direct comparison against previous work in variable-rate compression.
The paper's improvements: Tom: That's right. The main implication is that we can deploy these vision language action models in bandwidth-constrained or distributed settings because SPARC allows them to work effectively with less communication.
Jane: It really means that visual communication isn't just about sending raw pixels; it’s about intelligently prioritizing the information the downstream AI policy needs for its specific task.
Lu: The future work they mention is incorporating language grounding, which would let the compression system adaptively reflect the language instructions to improve performance even further.
Meng: From an engineering standpoint, I’m still focused on how this translates to actual hardware constraints in remote deployment scenarios, because their real-world experiments didn't reflect extremely busy channel constraints as much as we hoped.
Lalam: I think the idea of adapting based on language instructions is really cool, because it opens up a new layer of control over how the system interprets what it’s seeing.
Tom: This paper, "Learned Image Compression for Vision-Language-Action Models," shows that SPARC is better than other learned codecs and traditional codecs when you are constrained by bitrate.
Jane: It’s a powerful way to think about how to make visual communication efficient for complex robotic systems in the real world.
Conclusion: Tom: So we've been talking about SPARC, which is this learned image compression framework for VLA models that adaptively allocates bitrate based on what matters for control performance and task relevance in the latent space.
Jane: Exactly. We saw how it tackles the problem that existing codecs just compress pictures generally, but they don't care if one part of a picture is critical for the robot to function right.
Lu: It’s really clever because SPARC uses this temporal mask selector to look at the context and decide exactly which parts of the image—which spatial regions—should get more bits.
Meng: From an engineering view, it’s about making sure we aren't wasting bandwidth sending high-frequency details that the robot doesn't actually need for its current action.
Lalam: It shows how we can build systems where the compression itself is aware of the AI task, which could make our future AI applications much more resource efficient.
Tom: And the results are pretty strong; SPARC consistently beats other codecs on control performance even when they're given the same limited bitrate budget.
Jane: That means for anyone building a vision language action model, this is a way to get better action quality without needing more bandwidth overall.
Lu: The mechanism with that tilted rate loss is interesting; it stops the system from being too aggressive in suppressing details that might be important for the task, which is something many of these compression methods struggle with.
Meng: I see how that stabilization helps; if the training keeps pushing those critical patterns through, you end up with a more robust compression method.
Tom: It does. So we're talking about SPARC: Learned Image Compression for Vision-Language-Action Models, which is really pushing the boundaries on how we communicate visual information to robots.
Jane: It’s a solid piece of work because it proves that task awareness in compression directly translates to better real-world control performance.
Lu: It opens up a lot of possibilities for future work, especially if they can connect this spatial allocation idea with the language instructions themselves.
Meng: I’m curious to see if we can apply this kind of adaptive bitrate control to other types of data streams, not just images and video.
Lalam: It makes me think about how we design our own models; maybe the compression layer should be as smart as the model itself for efficiency.
POSTECH
cs.CV, cs.AI
Submitted: 2026-06-15
Updated: 2026-10-07
Importance score: 90/100
The gist: The gist: SPARC is a learned image compression framework tailored for VLA systems that adaptively allocates bitrate across spatial regions according to their contribution to downstream control,
Key concepts
- VLA Models
- Vision-Language-Action models combine vision, language understanding, and action planning to enable robots to understand visual inputs and perform tasks. They often require high-frequency multi-camera observations, making efficient visual communication crucial for real-time control in bandwidth-limited environments.
- Task-Aware Bitrate Allocation
- This is the core idea where the compression system learns how much data (bitrate) to keep for different parts of an image. Instead of allocating bits uniformly, SPARC learns that certain spatial regions, like a gripper or a target object, are more important for robot control and thus receive more bits.
- Temporal Mask Selector
- This is a lightweight component within SPARC that looks at the history of visual features to decide which parts of the image's latent representation should be kept. It generates a spatial mask that selectively keeps or discards information based on what is relevant for the current task.
- Tilted Rate Loss
- This is a specific training loss function used during training to stabilize SPARC. It prevents the masking policy from being overly biased toward regions with high local bit requirements, ensuring that important visual patterns are not aggressively suppressed during compression.
Terminology
Summary
The gist: SPARC is a learned image compression framework tailored for VLA systems that adaptively allocates bitrate across spatial regions according to their contribution to downstream control, consistently achieving stronger control performance than conventional codecs under the same bitrate budget.
Motivation and Problem
Vision-language-action (VLA) models increasingly rely on high-frequency multi-camera observations, making visual communication a major bottleneck for real-time robotic control in bandwidth-constrained or distributed deployment settings <ref:2606.16253#pg2> Existing image and video codecs, however, are designed to preserve generic visual fidelity rather than the control performance of downstream VLA policies <ref:2606.16253#pg2>. The importance of visual information varies substantially across both camera views and spatial regions within an image <ref:2606.16253#pg2>. Uniform bitrate allocation across cameras and image regions is therefore fundamentally inefficient <ref:2606.16253#pg2>.
SPARC Architecture
SPARC employs a learned image compression framework that performs task-aware bitrate allocation directly in the latent space of a neural image codec <ref:2606.16253#pg2>. The architecture is built on top of a pretrained neural image codec, incorporating a lightweight temporal mask selector that predicts spatial latent masks from temporal context, enabling adaptive bitrate allocation across both camera streams and image regions <ref:2606.16253#pg2>.
The key components of SPARC include:
-
A Transformer Encoder w/ FiLM modulation to reshape the input image <ref:2606.16253#pg5>.
-
A Temporal mask selector Sω which takes a history of hyperprior features (Ht,i) and generates a spatial binary mask (mt,i) to selectively mask the latent <ref:2606.16253#pg6>.
-
The application of elementwise masking mt,i ⊙ yt,i so that only the unmasked entries are stored and transmitted <ref:2606.16253#pg6>.
Training Strategy
SPARC is trained end-to-end to directly optimize downstream action quality under a communication budget <ref:2606.16253#pg2>. The training involves two phases:
-
Phase 1: Decoder warm-up, where the decoder is trained for W steps while freezing all other modules and disabling masking, reducing the objective to the action distortion loss E[l(at, aˆt)] <ref:2606.16253#pg5>.
-
Phase 2: Joint optimization with a tilted rate loss to stabilize training <ref:2606.16253#pg2>.
The tilted rate loss is defined as l(α)rate(y; m) = (1/N Σ PN j=1 mj b α j) · (Σ PN j=1 bj), where b is the local bit vector and α ∈ [0, 1] controls the trade-off <ref:2606.16253#pg6>. This loss prevents the masking policy from being overly biased toward regions with high local bitrate by using an exponentiated version of the local bit vector b(α) <ref:2606.16253#pg6>.
Experimental Validation
Experiments on diverse robotic benchmarks, including RoboCasa365, VLABench, and LIBERO, show that SPARC consistently achieves stronger control performance than conventional image/video codecs and recent learned compression methods under the same bitrate budget <ref:2606.16253#pg3>. Across all settings, SPARC consistently achieves stronger bitrate-success tradeoffs than conventional image/video codecs and recent learned compression methods <ref:2606.16253#pg2>.
Qualitative results demonstrate that SPARC prioritizes task-critical regions, such as the gripper and target objects, to facilitate robotic control <ref:2606.16253#pg4>. SPARC allocates more bits to task-relevant regions (i.e., near the microwave and the gripper’s end-effector) <ref:2606.16253#pg5>. Furthermore, SPARC is trained not to mask important latent features, which explains why the performance of SPARC is at least similar to or better than that of MS-ILLM on these tasks <ref:2606.16253#pg10>.
Performance and Analysis
SPARC consistently achieves higher simulation success rates under comparable bpp budgets across diverse VLA models and benchmarks <ref:2606.16253#pg5>. On the RoboCasa365 benchmarks, SPARC shows strong performance on π0 and π0.5 <ref:2606.16253#pg5>. In real-world remote-control settings, SPARC substantially improves the bitrate-success tradeoff <ref:2606.16253#pg2>. The framework reduces the size of latents to enable the reduction of latency during entropy encoding and entropy decoding <ref:2606.16253#pg2>.
The analysis of key components shows that SPARC clearly achieves the highest success rate at higher bpps where the task performance is meaningful <ref:2606.16253#pg10>. The effect of scaling a on the bit-success rate trade-off curve indicates that selecting a moderate value of a (0.4 ≤ a ≤ 0.8) improves the performance of VLA models by preventing the temporal mask selector from aggressively masking high-frequency details to which a high proportion of bits is allocated <ref:2606.16253#pg6>.
The work concludes that SPARC’s performance gain is greater than other learned codecs and traditional codecs <ref:2606.16253#pg10>. The framework reduces the size of latents to enable the reduction of latency during entropy encoding and entropy decoding <ref:2606.16253#pg2>. SPARC is instead optimized for reconstructing task-relevant details for downstream VLA models <ref:2606.16253#pg9>. The future work would incorporate this point to adaptively reflect the language instruction <ref:2606.16253#pg10>. SPARC is capable of allocating more bits to task-centric regions to improve the performance of VLA models <ref:2606.16253#pg9>. The paper has explored how to eliminate unnecessary details, while preserving important information for VLA models <ref:2606.16253#pg9>. The absence of language grounding is one limitation of this paper <ref:2606.16253#pg10>. Due to spatiotemporal constraints, our setting does not reflect extremely busy channel constraints in the real-world experiments <ref:2606.16253#pg10>.
--- Page 1 ---
Learned Image Compression for Vision-Language-Action Models Hyeonjun Kim POSTECH kim.hyeonjun@postech.ac.kr Jegwang Ryu POSTECH jegwang.ryu@postech.ac.kr Sangbeom Ha POSTECH sangbeomha@postech.ac.kr Junhyeok Lee Soongsil University wnsx0000@gmail.com Jun-Hyuk Kim Chung-Ang University junhyukkim@cau.ac.kr Hyemin Ahn POSTECH hmahn@postech.ac.kr Jaeho Lee POSTECH jaeho.lee@postech.ac.kr Figure 1: SPARC is a neural image compression framework for communication-efficient VLA deployment, which adaptively allocates bitrate across spatial regions according to their contribution to downstream control (brighter regions in the bit allocation map indicates higher importance) <ref:2606.16253#pg2>. Abstract: Vision-language-action (VLA) models increasingly rely on highfrequency multi-camera observations, making visual communication a major bottleneck for real-time robotic control in bandwidth-constrained or distributed deployment settings <ref:2606.16253#pg2>. Existing image and video codecs, however, are designed to preserve generic visual fidelity rather than the control performance of downstream VLA policies <ref:2606.16253#pg2>. In this work, we introduce SPARC (SPatially Adaptive Rate Control), a learned image compression framework tailored for VLA-driven robots <ref:2606.16253#pg2>. Our key observation is that the importance of visual information varies substantially across both camera views and spatial regions within an image <ref:2606.16253#pg2>. Based on this observation, SPARC employs a lightweight temporal mask selector that adaptively allocates bitrate over latent representations according to task relevance while leveraging temporal context <ref:2606.16253#pg2>. We further introduce a tilted rate loss that stabilizes training by reducing the tendency of entropy-based objectives to over-suppress task-critical visual patterns <ref:2606.16253#pg2>. Experiments on diverse robotic benchmarks, including RoboCasa365, VLABench, and LIBERO, show that SPARC consistently achieves stronger control performance than conventional image/video codecs and recent learned compression methods under the same bitrate budget <ref:2606.16253#pg2>. We additionally demonstrate real-world deployment benefits in remote-control settings, where our method substantially improves the bitrate-success tradeoff <ref:2606.
Improvements for AI systems
-
Bold header: Task-Aware Bit Allocation in Latent Space. The system will
employ a lightweight temporal mask selector that adaptively allocates bitrate over latent representations according to task relevance while leveraging temporal context,
allowing for dynamic, non-uniform bit allocation across camera views and spatial regions based on the current VLA task. -
Bold header: Stable Training via Tilted Rate Loss. The introduction of
a tilted rate loss that prevents entropybased optimization from aggressively suppressing statistically rare yet task-critical visual patterns
will ensure thattask-critical visual patterns
are retained, thereby stabilizing training and preventing the over-suppression of important information during learning. -
Bold header: Superior Bitrate-Success Tradeoff. The improved system will
consistently achieve stronger control performance than conventional image/video codecs and recent learned compression methods under the same bitrate budget,
enabling VLA policies to function effectively in bandwidth-constrained or distributed deployment settings while maintaining high accuracy on benchmarks like RoboCasa365. -
Bold header: Reduced End-to-End Latency in Deployment. By optimizing the objective as
the aggregated total rate over the cameras, instead of applying the same rate individually,
and leveraging compression to reduce latent size, the system will result inlower end-to-end latency in practical deployments.
Sources
- $\pi_{0.5}$: a Vision-Language-Action Model with Open-World Generalization
- GR00T N1: An Open Foundation Model for Generalist Humanoid Robots
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models