EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control

arXiv:2605.21862 · cs.RO, cs.AI · Submitted 2026-05-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control".

Jane: The paper was written by Chushan Zhang, Ruihan Lu, Jinguang Tong, Xuesong Li, Yikai Wang et al. from Australian National University and The University of Queensland and Beijing Normal University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everybody. Today we are looking at a fresh arXiv paper with a pretty dense title: "EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control." Jane, I’ll be honest, I read that title three times before I had any idea what it meant.

Jane: Ha, same here, Tom. But once you unpack it, it’s actually a really intuitive idea. So imagine a robot that’s doing a task like opening a microwave. It looks at the scene, plans a chunk of actions, and then executes them. The problem is, while it’s moving, the scene changes — the door swings open, the gripper gets closer — but the robot’s plan was made from the old view.

Tom: Right, so it’s planning blind, essentially. It’s like driving with a map from last week.

Jane: Exactly. And that’s what this paper tackles. The authors argue that the robot should carry a "scene belief" — a mental model of what the scene looks like after its own actions — and pass that belief to the next planning step. They call that an action-updated scene prior.

Tom: And that’s the "evolving" part. The scene belief evolves as the robot acts. The paper is from a team at ANU, University of Queensland, and Beijing Normal University, and they’ve built this into a vision-language-action model, or VLA, which is the kind of model that takes camera images and language instructions and outputs robot controls.

Jane: And the key trick is where they put this scene belief. They put it inside the action decoder, which is the part of the model that generates the motor commands. So the decoder doesn’t just output actions — it also outputs an updated scene state, which becomes the input for the next chunk.

Tom: That’s the "inside the action decoder" part of the title. It’s a neat architectural choice, because normally the action decoder is just a one-way street. Here, it’s a loop.

Jane: And that loop is what lets the robot correct itself. When the next camera frame arrives, the model compares its predicted scene state with the real observation and fixes any drift. So you get persistence, action-updating, and correction — all three.

Tom: So this isn’t just a memory buffer of past frames. It’s a living model of the scene that the robot itself is changing.

Jane: Precisely. And that’s the core contribution. They call it a recurrent scene prefix, and it’s what makes the whole thing work.

Tom: I’m already curious how they train something like that. That’s got to be tricky, because you need supervision for the scene state, not just the actions.

Jane: Oh, absolutely. And that’s exactly what we’re going to dig into next — the training setup, the geometric anchoring, and how they keep it from collapsing into a generic image feature.

Tom: Good, because I want to know how they teach the model to predict the future scene without an extra online world model.

Jane: Stay tuned, because that’s the clever part.

Summary: Tom: So we’re back, and we’ve got the title unpacked. Now let’s talk about what the paper actually does, because the summary is dense. Jane, walk us through the core mechanism.

Jane: Sure. So the model, EvoScene-VLA, takes multi-view images — head camera, left wrist, right wrist — plus a language instruction like "clean the sink." It builds a prefix of tokens that goes into the vision-language model. And in that prefix, there are two special groups of "scene slots."

Tom: The observation slots and the prior slots, right?

Jane: Exactly. The observation slots read the current image. The prior slots inherit the scene state from the previous chunk — the one the action decoder wrote. The VLM then refines both against the new image, and the result is a corrected scene representation.

Tom: And then the action decoder takes over. This is where the co-denoising happens.

Jane: Yes. The action decoder doesn’t just denoise the action chunk. It also denoises a matched scene chunk in the same flow-matching pass. So at inference, you get both the motor commands and an updated scene state. The scene token at the executed step gets fed back as the prior for the next chunk.

Tom: So the action decoder is doing double duty — it’s both the motor planner and the scene updater.

Jane: Right. And that’s what lets them drop the auxiliary predictor at deployment. There’s no separate world model running online. The action decoder itself produces the next prior.

Tom: But training that must be a nightmare. How do you supervise the scene chunk if there’s no ground-truth scene state?

Jane: That’s where the two training-only modules come in. First, there’s the Geometric Anchor, which has two levels. The local level supervises each observation slot with cross-view masked depth reconstruction — so it forces the slot to infer depth for a masked view using the other views. The global level distills a frozen three dee foundation model into the scene representation.

Tom: So you’re grounding the scene slots in real geometry, not just letting them encode vague appearance.

Jane: Exactly. And then there’s the Scene Predictor, which takes the current scene representation and the action sequence and predicts future scene latents at key frames. Those future latents are supervised against the three dee foundation model’s features on the actual future frames.

Tom: So the Scene Predictor is like a teacher that shows the action decoder what the future scene should look like.

Jane: Yes, and then the flow-matching loss distills those targets into the action decoder’s scene branch. During training, the action decoder learns to produce scene states that match what the Scene Predictor says the future should be. At inference, the Scene Predictor is gone, and the action decoder does it on its own.

Tom: And the results? I saw the numbers — thirty-one RoboTwin tasks, success rate up from eighty-seven point two to eighty-nine point one in fixed evaluation, and from eighty-six point one to eighty-eight point five in randomized.

Jane: And the randomized gain is bigger, which makes sense. When the initial layout varies, the robot can’t rely on memorized positions, so the scene prior becomes more valuable. It’s correcting for perception errors that compound across chunks.

Tom: That’s a solid improvement, but I’m wondering about the real-world part. They tested on a Galaxea R1-Lite robot doing cleaning tasks, right?

Jane: Yes, wiping mirrors, cleaning sinks, cutting boards. Average success went from thirty-seven point three to forty-two point zero percent. Not huge, but consistent, and those are tasks where the robot’s own actions change the surface state — so exactly the scenario the method targets.

Tom: So the summary is: persistent scene state, action-updated, observation-corrected, and it works in both sim and real.

Jane: That’s the paper in a nutshell. And the ablations show each piece contributes — the global anchor, the local depth, and the recurrent prior.

Tom: I want to dig into those ablations and the design choices next, because there’s a question about whether this generalizes beyond cleaning and microwaves.

Jane: Good, let’s get into that.

Improvements: Tom: So we’ve covered what the paper does. Now let’s talk about what it improves and why that matters. Jane, you mentioned the ablations — what did they actually show?

Jane: They ran ablations on a five-task subset of RoboTwin. The baseline LingBot-VLA with depth supervision got eighty-seven point eight clean and eighty-four point six randomized. Adding the global anchor and scene predictor brought it to eighty-nine point three and eighty-six point two. Then adding the local depth anchor pushed it to ninety point one and eighty-six point five. And finally, adding the recurrent prior at inference — actually carrying the scene state across chunks — brought it to ninety point eight and eighty-seven point eight.

Tom: So each piece adds a bit. The recurrent prior alone is worth about a point and a half.

Jane: Right. And that’s the key improvement over prior work. Spatial VLAs improve geometry in a single frame. Temporal VLAs keep past observations. But neither maintains a scene state that’s updated by the robot’s own actions and then corrected against new observations.

Tom: So the improvement isn’t just a better perception module. It’s a different loop — the action decoder writes the next scene state.

Jane: Exactly. And that’s why they call it "evolving scene beliefs." The belief is updated by actions, not just by new pixels.

Tom: Now, Lu, you’re the researcher here. What do you think is the most exciting implication of this design?

Lu: I think the exciting part is that they’ve shown the action decoder can be a general-purpose scene updater. It’s not a separate world model bolted on the side. It’s the same network that produces motor commands, and it’s producing a latent scene state as a byproduct. That suggests we can get world-model-like behavior without the cost of a separate model at inference.

Jane: And that’s a big deal for deployment, because running an extra world model online is expensive.

Meng: Yeah, and I want to ask about that from an engineering standpoint. The paper says they discard the Scene Predictor and Geometric Anchor at inference. So what’s the actual compute overhead of the recurrent scene prefix?

Jane: The overhead is just the extra scene tokens in the prefix and the co-denoising in the action decoder. The scene chunk is denoised alongside the action chunk in the same flow-matching pass, so it’s not a separate sampling loop. They report running on a single RTX four thousand ninety for real-robot inference.

Meng: That’s reasonable. But I’m curious about the chunk length. They use fifty-step action chunks. If you make chunks longer, the scene changes more, so the prior should be more valuable. But the Scene Predictor targets also get further into the future, so they’re harder to predict. Is there a sweet spot?

Jane: The paper actually flags that as an open question. They say longer chunks may increase the value of recurrence, but they also push the targets further out. So the net effect is unknown.

Tom: So it’s a trade-off they haven’t fully explored yet.

Jane: Right. And there’s another limitation — the scene state is latent, so it’s not directly interpretable. You can’t look at it and say "oh, the robot thinks the cup is here." You have to judge it through behavior.

Lu: But that’s also an opportunity. If the scene state is latent, you could potentially use the mismatch between the observation slots and the prior slots as an uncertainty signal. If the prior and the observation disagree a lot, that means the robot’s model of the scene is off, and maybe it should replan.

Meng: That’s a really practical idea. You could use that mismatch to trigger re-observation or adaptive chunk execution.

Jane: And the paper mentions that as future work too. So the improvements here are solid, but they also open up a bunch of directions.

Tom: So the improvements aren’t just numbers — they’re a new way of thinking about what the action decoder can do.

Jane: Exactly. And that’s the big picture. The action decoder isn’t just a motor command generator anymore. It’s also a scene state writer.

Tom: I love that framing. Let’s wrap up with the conclusion and what this means for the field.

Conclusion: Tom: Alright, we’re at the end of our time with "EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control." Jane, give us the final summary.

Jane: So the core idea is that chunked robot control policies need a persistent scene state that’s updated by the robot’s own actions and corrected against new observations. EvoScene-VLA implements that as a recurrent scene prefix, with the action decoder co-denoising actions and scene tokens together. At inference, there’s no separate world model — the action decoder itself writes the next prior.

Tom: And the training uses two clever tricks — the Geometric Anchor for grounding in three dee structure, and the Scene Predictor for future-scene targets. Both are discarded at deployment.

Jane: Right. And the results show consistent gains on thirty-one RoboTwin tasks, with bigger gains under randomized conditions, and on a real Galaxea R1-Lite robot doing cleaning tasks. The ablations show each component contributes.

Lu: I think the biggest implication is that we don’t need a separate world model to get action-aware scene understanding. The action decoder can be the scene updater. That’s a shift in how we think about VLA architectures.

Meng: And from a practical standpoint, the overhead is modest — just extra tokens and co-denoising in the same pass. That makes it deployable on a single GPU.

Lalam: And I’d add that this has cultural implications too. As robots move into homes and workplaces, they need to understand not just what they see, but how their own actions change the world. A robot that can track its own effect on the scene is a robot that can be trusted with longer, more complex tasks — wiping a counter, tidying a room, assisting a person over time. That’s the kind of capability that makes household robots genuinely useful rather than just impressive demos.

Tom: That’s a great way to put it. So we’re saying goodbye to EvoScene-VLA, but I think this idea of the action decoder as a scene writer is going to stick around.

Jane: Absolutely. It’s a small architectural change with a big conceptual shift. And we’ll be watching to see how the chunk-length question and the uncertainty-signal idea play out in follow-up work.

Tom: Thanks for listening, everyone. We’ll see you next time with another paper from the arXiv.

Jane: Take care, and keep building.

Australian National University · The University of Queensland · Beijing Normal University

cs.RO, cs.AI

Submitted: 2026-05-21

Updated: 2026-09-29

Code: https://github.com/OpenGalaxea/GalaxeaVLA

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 82/100

The gist: EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control The paper addresses a fundamental limitation of chunked vision-language-action (VLA) policies: "Chunked

Key concepts

Scene Belief
A mental model of what the scene looks like after a robot performs actions. This belief is updated by the robot's actions and corrected against new visual observations, providing persistence and accuracy across planning chunks.
Action Decoder
The part of the VLA model responsible for generating motor commands. In EvoScene-VLA, this decoder is modified to also output an updated scene state, allowing it to act as both a planner and a scene updater.
Recurrent Scene Prefix
A mechanism where the action decoder outputs a scene token that serves as the prior for the next chunk's planning. This recurrent prefix allows the robot to carry and evolve its understanding of the scene from one action step to the next.

Terminology

Summary

EvoScene-VLA: Evolving Scene Beliefs Inside the Action Decoder for Chunked Robot Control

The paper addresses a fundamental limitation of chunked vision-language-action (VLA) policies: Chunked vision-language-action policies predict multi-step robot controls, but they usually condition scene context on observations rather than on the robot's own actions. The authors argue that actions can change the scene before the next observation arrives: wiping the counter changes the surface state, closing a drawer hides its contents, and lifting a cup leaves the shelf slot below it empty. Within a chunk, the visual context may not reflect such changes at all, and across chunks, the policy still lacks a compact record of how its recent actions transformed the scene, and must re-infer those changes from partial, noisy, or occluded visual evidence.

The paper identifies that existing approaches fall short: Spatial VLAs improve geometric reasoning within a single image through depth supervision, 3D encoding, or multi-view features, but they reason from current observations alone and do not update scene geometry once actions have changed it. Temporal VLAs retain past observations through memory banks, trace prompts, or recurrent states, yet past observations alone do not specify how recent or planned actions should update the scene. Action-conditioned prediction methods anticipate future representations, but consume each prediction inside the current decision and discard it afterwards.

The central claim is: The missing piece, we argue, is an action-updated scene representation. The policy should pass this representation to the next control call as a prior.

The paper specifies three properties a useful scene representation for chunked control must have: "It should persist across chunks, so the policy is not forced to re-infer it from each new observation. It should update under the actions the policy generates, since those actions change the scene. And it should correct against each new observation, so prediction errors do not accumulate. The authors emphasize: Missing any one of these undermines the design: no persistence means no cross-chunk context, no action update means no post-action prior, and no correction means errors compound."

EvoScene-VLA extends LingBot-VLA with a recurrent scene prefix for chunked control. The architecture adds two slot groups to the VLM prefix: observation slots that gather evidence from the current image, and prior slots that inherit the scene state denoised by the action expert in the previous chunk. The action expert co-denoises the next action chunk together with a matched scene chunk in a single flow-matching pass, and the denoised scene token at the executed step is fed back as the prior for the next call.

Two training-only modules supervise this loop: Geometric Anchor grounds scene slots in metric geometry, and Scene Predictor supplies future scene-token targets. At inference, EvoScene-VLA discards both training-only modules and retains only the recurrent scene prefix and action–scene co-denoising.

Let D denote the VLM hidden dimension, V the fixed camera views (head, left wrist, right wrist), and N the number of slots in each observation or prior group. The VLM prefix is augmented with "per-view observation slots s obs(v) ∈ R(N×D), v ∈ 1,..., V, which gather geometric evidence from each camera, and a single set of prior slots s̄ t ∈ R(N×D), which inherit the action-updated scene state from the previous chunk. The prefix is ordered: [x t, s obs(1:V), s̄ t, l]." At the first chunk, s̄ t is initialized from learnable scene embeddings; for later chunks, it carries the recurrent state denoised by the action expert at the previous step.

An asymmetric attention mask routes information: "Image and language tokens ignore scene slots, preserving pretrained visual and linguistic paths. Each observation slot attends to its own view's image tokens and to other slots in its own observation group. Prior slots attend to all observation slots and to themselves, but not to image or language tokens, and no other token attends back to them. This isolation preserves the pretrained image–language pathway intact while giving the scene state a dedicated bottleneck. The VLM output at the prior-slot positions, denoted s p, is taken as the corrected scene representation for the current chunk."

The paper states: The recurrent prefix specifies where the scene state lives but not what it should encode; without targeted supervision, the slots are free to absorb generic image features rather than 3D structure. Geometric Anchor addresses this with two complementary training-only branches.

The local branch grounds each observation slot in per-view geometry while forcing it to integrate evidence from the other views, so that no slot can rely on its own image tokens alone. The method masks one view at a time and requires the head to recover its depth representation from the remaining views and the cross-view representation s p. Observation slots are pooled from a 256-token query bank q tmpl, used as the query set for a lightweight cross-attention head g depth. For each target view i, the VLM image tokens are masked by broadcasting a learned embedding m, and the masked multi-view tokens together with s p are passed to g depth. The loss is: L geo = (1/V) Σ i SmoothL1(f̂ d t,i, f d t,i), where f d t,i = MDT(x t,vi) comes from a frozen Monocular Depth Teacher. The paper notes: Masking the target view removes the shortcut of copying its own VLM features, forcing g depth to aggregate cross-view evidence from the unmasked tokens and from s p.

The global branch grounds the cross-view representation s p in metric 3D by distilling a frozen multi-view 3DFM. A lightweight view-conditioned decoder g 3D takes learnable queries q dec and uses s p as keys and values; a linear projector W proj maps the decoder output to the foundation-model feature space: H t = g 3D(q dec; s p), P t = W proj H t. The loss is: L rep = (1/V) Σ v P t(v) − Z t(v) 1, where Z t = 3DFM(x t). The l1 objective regresses the foundation-model features element-wise, providing dense token-level supervision that is robust to outliers and preserves both direction and magnitude of the target representation.

Scene Predictor produces the future scene-token targets that the action expert later distills into its scene branch. Conditioned on the current scene representation s p and the action sequence a t:t+H, it predicts a sequence of absolute future scene latents at sparse key-frame steps, supervised against features from the 3D foundation model on the corresponding future frames.

The module is a causal Transformer that takes as input the robot state r t, the current scene representation s p, the action sequence a t:t+H, and K key-frame query groups initialized from s p. Under a causal mask, each query group q i attends to r t, s p, the action prefix a t:t+k i up to its target step, and earlier query groups, so that the prediction at step t + k i is conditioned only on actions executed up to that step. The output is a sequence of absolute future scene latents ŝ t+k 1:t+k K at sparse key-frame steps k 1,..., k K ⊆ 1,..., H.

Scene Predictor reuses the view-conditioned decoder (g 3D, W proj): the same operator that grounds s p, also decodes each predicted future latent and matches it to foundation-model features on the corresponding future multi-view frame. The loss is: L pred = (1/(K·V)) Σ i,v P̃ t+k i(v) − Z t+k i(v) 1.

The action expert learns to denoise actions and future scene latents jointly under a single flow-matching vector field. During training, this couples motor and scene targets on a shared denoising schedule. At inference, the same denoising loop produces both the action chunk and the next recurrent prior, replacing Scene Predictor and closing the recurrence loop without any auxiliary online module.

Given the Scene Predictor outputs, they are stacked into future-scene targets z 0 ∈ R(K×N×D). A single flow-matching time τ ∈ [0, 1] is sampled, shared between the action and scene paths, with independent Gaussian noises ϵ a and ϵ s, forming straight-line interpolants: a τ t:t+H = τϵ a + (1−τ)a t:t+H, z τ = τϵ̃ s + (1−τ)z 0, where ϵ̃ s:= σϵ s rescales the scene noise to match the empirical magnitude of z 0.

The action expert receives the suffix [r t, z τ, a τ t:t+H] under a causal suffix mask, attends to the VLM prefix cache, and predicts per-token velocities. The losses are: L sceneFM = v θ(s)(z τ, τ) − (ϵ̃ s − z 0)22 and L actFM = v θ(a)(a τ, τ) − (ϵ a − a t:t+H)22. The paper notes: L actFM is the standard π0.5 action FM loss. L sceneFM distills future scene representations into the action expert.

At deployment: "each chunk runs one VLM forward followed by one Euler-step denoising pass; the robot executes the resulting action chunk, and the scene token at the final key-frame offset k K becomes the prior s̄ t+1 for the next chunk. The VLM call at the next chunk corrects this prior against the new observation, closing the recurrent loop."

The full training objective is: L = L actFM + λ1L geo + λ2L rep + λ3L pred + λ4L sceneFM. The four scene-side losses play three roles: L geo and L rep ground current scene representations in geometry, L pred trains future representations in the same coordinate, and L sceneFM transfers those representations into the action expert. Training is end-to-end in a single stage.

Testbeds. The paper evaluates on the RoboTwin simulated benchmark and the Galaxea R1-Lite real-robot platform. In simulation, the main evaluation uses 31 tasks; for ablations, a 5-task subset (RoboTwin-5Task) is used. Two simulation settings are tested: Clean fixes initial object positions to the training distribution. Rand randomizes initial positions and orientations within the task workspace.

Baselines. The model is fine-tuned from the public LingBot-VLA pretrained checkpoint with the same 50-step action chunk. Comparisons are against three baselines: π0.5, LingBot-VLA, and LingBot-VLA∗ (depth-augmented LingBot-VLA).

On 31 RoboTwin tasks, EvoScene-VLA improves the LingBot-VLA∗ average from 87.2 to 89.1 under Clean (+1.9) and from 86.1 to 88.5 under Rand (+2.4). The gain is larger under Rand than under Clean. The paper explains: "Two factors in Rand plausibly contribute: varied initial layouts make the scene harder to perceive from a single observation, and the resulting per-chunk perception errors are more likely to compound across chunks. The recurrent scene prefix is designed to mitigate both issues."

Qualitative analysis shows that in tasks like Open Microwave, the gripper retracts before reaching the door handle for the baseline, while EvoScene-VLA, whose action expert co-denoises a scene update alongside the action chunk, conditions on this intermediate scene state and continues the motion through to completion. Trajectory plots show EvoScene-VLA produces noticeably smoother paths than LingBot-VLA.

On RoboTwin-5Task, the additive ablations show: LingBot-VLA baseline achieves 87.8 (Clean) and 84.6 (Rand); LingBot-VLA∗ achieves 81.6 (Clean) and 75.8 (Rand). Adding the global anchor (L rep) and Scene Predictor (L pred) jointly reaches 89.3 (Clean) and 86.2 (Rand). Adding the local depth anchor (L geo) reaches 90.1 (Clean) and 86.5 (Rand). Finally, propagating the recurrent prior across chunks at inference, rather than reinitializing s̄ t from learnable embeddings at every chunk, further improves performance to 90.8 (Clean) and 87.8 (Rand).

The real-robot evaluation uses the Galaxea R1-Lite dual-arm platform, training on the indoor-cleaning subset of the Galaxea Open-World Dataset (mirror, sink, and cutting-board cleaning). The subset spans 7 recording sessions, 439 episodes, 48,419 frames, and 1,756 video clips, totaling approximately 9 hours of bimanual demonstrations at 15 fps. The three long-horizon tasks require the robot to track how its tool changes the surface: the surface state evolves during wiping, and subsequent controls must target unwiped regions that can be subtle, occluded, or ambiguous in the current view. For each task, 100 closed-loop rollouts are run with randomized object placements, lighting, and initial robot state.

Results show EvoScene-VLA improves average success from 37.3% to 42.0% over the best baseline (LingBot-VLA∗). Per-task results: Mirror (29% vs. 26-28%), Sink (51% vs. 42-49%), Cutting-board (46% vs. 34-44%). The paper notes: These gains show the same pattern as simulation: the method helps when the robot changes the scene as it acts.

The paper states: "The experiments suggest a specific failure mode in chunked VLA control: the policy does not only need better current-frame geometry or longer observation history; it needs a prior for the scene after its own actions have changed it. The method improves most in settings where this mismatch should matter, including randomized RoboTwin evaluation and real cleaning tasks with evolving surface state. The paper argues: the action decoder is a natural place to update scene state because it already generates the action sequence that will reshape the scene. The resulting prior does not need to be a dense reconstruction. A compact policy-facing latent can help if training grounds it in 3D structure and the next VLM call corrects it with fresh observations."

Limitations acknowledged: "Because the recurrent state is latent, we judge its geometric content through downstream behavior and ablations. Because future scene targets come from future-frame 3D foundation-model features through Scene Predictor, target quality can limit the prior that the action decoder learns. The paper notes: Longer chunks may increase the value of recurrence by creating larger scene changes, but they also push Scene Predictor targets farther into the future, so the net effect of chunk length on prior quality remains an open question. A suggested next step is to use the mismatch between observation slots and prior slots as an uncertainty signal for replanning or adaptive chunk execution."

The paper lists three contributions: (1) We show that an action expert can produce an action-updated scene prior for the next control call. It uses a recurrent scene prefix and action–scene co-denoising, without an auxiliary online predictor at inference. (2) We introduce a two-level Geometric Anchor that combines local depth supervision with a global 3D-foundation-model anchor, sharing one decoder to ground both current and future scene latents. (3) "We report consistent gains on the RoboTwin benchmark and on a Galaxea R1-Lite real-robot platform, with ablations showing cumulative contributions from future-scene supervision, geometric anchoring, and the recurrent prior."

The optimization recipe follows LingBot-VLA: AdamW optimizer, learning rate 1×10−4, effective batch size 256, 20,000 total update steps, bf16 storage with fp32 reductions, trained on 8×A800 GPUs. The VLM backbone is Qwen2.5-VL-3B-Instruct with hidden dimension D=2048, three camera views, 224×224 resolution, 256-token template bank, N=16 slots per group, K=3 key frames, and 50-step action chunk. Loss weights are λ1=0.04, λ2=0.10, λ3=0.10, λ4=0.01. At inference, 10 Euler steps are used with a re-observation period of 1 chunk (50 steps). The frozen teachers are LingBot-Depth for local depth and Pi3 multi-view 3D foundation model for global anchoring.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in an AI system, and what the improved system can do:


  • What to change: Instead of conditioning each action chunk only on the current visual observation, I will add a persistent, action-updated scene-state buffer (a set of latent tokens) that is carried across control calls.

  • How: The VLM prefix will contain two slot groups: (a) observation slots that read the current multi-view image, and (b) prior slots that inherit the scene representation denoised by the action expert in the previous chunk. An asymmetric attention mask ensures the prior slots only receive information from observation slots (not directly from image/language tokens), preserving the pretrained VLM pathway.

  • Benefit: The policy will no longer re-infer scene changes from scratch after each action chunk. It will start each new chunk with a prior that already reflects the robot’s recent actions, reducing errors from occlusion, contact, and object motion.

  • What to change: The action decoder will not only predict the next action chunk but also a matched scene-token chunk in the same denoising loop.

  • How: During training, the action expert receives both the action interpolant and a scene interpolant (derived from future scene latents). It predicts velocities for both. At inference, the same Euler-step denoising produces both the action chunk and the updated scene token. The scene token at the executed step becomes the prior for the next chunk.

  • Benefit: Each predicted action is grounded in a scene state that is consistent with the action sequence itself. This eliminates the “action-scene mismatch” failure mode where the policy plans a chunk based on a stale observation and then commits to a target that is no longer valid.

  • What to change: I will add two training-only supervision branches to ground the scene tokens in 3D geometry:

  • Local Anchor: Cross-view masked depth reconstruction. Mask one camera view at a time and require the model to reconstruct that view’s depth representation from the other views plus the cross-view scene representation. Supervise with a frozen monocular depth teacher.

  • Global Anchor: Distill a frozen 3D foundation model (e.g., Pi3) into the aggregated scene representation using a lightweight view-conditioned decoder and an l1 loss.

  • Benefit: The scene tokens will encode metric geometry rather than generic appearance. This makes the prior more useful for tasks requiring spatial reasoning (e.g., reaching into a drawer, stacking blocks, wiping a surface).

  • What to change: I will add a causal Transformer (Scene Predictor) that takes the current scene representation, robot state, and action sequence, and predicts future scene latents at sparse key-frame steps. These predictions are supervised against future-frame 3D foundation-model features.

  • How: The Scene Predictor is used only during training. Its outputs become the targets for the action expert’s scene branch via the flow-matching loss. At inference, the Scene Predictor is discarded; the action expert itself acts as the scene updater.

  • Benefit: The action expert learns to write a scene prior that is consistent with future observations, not just with the current frame. This improves the quality of the recurrent state and prevents drift across long horizons.

  • What to change: The global anchor decoder (g3D + Wproj) will be reused for both the current scene representation (for Lrep) and the predicted future scene latents (for Lpred).

  • Benefit: This forces the current and future scene representations to live in the same geometric coordinate space, making the recurrent prior more coherent and easier to correct against new observations.

  • What to change: In the joint action–scene denoising, I will sample a single flow-matching time τ for both action and scene paths, and rescale the scene noise to match the empirical magnitude of the scene targets.

  • Benefit: This prevents the scene branch from dominating or being ignored during training, and ensures the action and scene updates are temporally aligned.

  • What to change: At deployment, I will keep only the VLM, the recurrent scene prefix, and the action expert. No Scene Predictor, no Geometric Anchor, no depth teacher, no 3D foundation model.

  • How: The action expert co-denoises actions and scene tokens. The scene token at the final key-frame offset is written back as the prior for the next chunk. The next VLM call corrects this prior against the new observation.

  • Benefit: The deployed system is lightweight, with no extra inference-time memory or prediction modules, yet it carries an action-updated scene prior across chunks.

  1. Execute long-horizon manipulation tasks with fewer mid-chunk failures.

Example: In “Open Microwave,” the system will not retract the gripper prematurely because it will have a scene prior that reflects the door’s current position after the first part of the chunk.

  1. Handle tasks where the robot’s own actions change the scene state.

Example: Wiping a counter, closing a drawer, lifting a cup, or stacking blocks. The system will track how its actions transform the scene and plan subsequent actions accordingly.

  1. Maintain spatial awareness under occlusion and partial visibility.

Because the prior is updated by actions and corrected by new observations, the system can reason about objects it cannot currently see (e.g., inside a drawer, behind the gripper).

  1. Produce smoother, more consistent end-effector trajectories.

The co-denosing of actions and scene tokens reduces abrupt corrections that arise from planning each chunk from a stale observation.

  1. Improve success rates under randomized initial conditions.

The paper reports a larger gain under randomized evaluation (+2.4%) than under fixed evaluation (+1.9%), suggesting the recurrent prior makes the policy more robust to pose and layout variation.

  1. Deploy on real robots without extra inference-time compute.

The system uses only the VLM and action expert at deployment. No world model, no future-frame predictor, no depth model, no 3D foundation model is needed online.

  1. Generalize to real-world cleaning and manipulation tasks.

On the Galaxea R1-Lite platform, the system improved average success from 37.3% to 42.0% across three cleaning tasks, showing the benefit transfers from simulation to physical robots.

If you implement these changes, the resulting AI system will be a chunked vision-language-action policy that not only plans actions but also maintains a self-correcting, action-updated scene prior. This will make it significantly more reliable for real-world manipulation tasks where the robot’s own actions continuously reshape the environment.

Abstract

Chunked vision-language-action (VLA) policies predict multi-step robot controls, conditioning each update on the current visual observation alone. Yet robot actions cause contact, occlusion, and object motion, and the geometry that later decisions depend on can change before the next visual update arrives. Spatial VLAs improve current-frame geometry. Temporal VLAs aggregate past frames. Neither maintains an action-updated scene prior across chunks. We argue for a persistent action-updated scene state across control calls, and introduce EvoScene-VLA. Its recurrent scene prefix carries a geometry-aware scene state across chunks. At each vision-language model (VLM) call, the VLM combines scene information from the current observation with the action-updated prior from the previous chunk; the action decoder outputs both the next action chunk and a compact scene update. This update becomes the next prior, which the VLM corrects against the new observation when the next call arrives. Each control call therefore starts from a scene prior that reflects both recent actions and fresh visual evidence. During training, Scene Predictor supplies future scene-token targets, and Geometric Anchor aligns scene slots with frozen depth and 3D teachers. We discard both modules at deployment. On 31 RoboTwin tasks, EvoScene-VLA raises average success from 87.2% to 89.1% in fixed evaluation and from 86.1% to 88.5% in randomized evaluation. On the Galaxea R1-Lite real robot, EvoScene-VLA outperforms all baselines.

Sources

Related papers