G0.5: One Autoregressive Stream for Robot Reasoning and Action

arXiv:2608.11739 · cs.RO, cs.AI · Submitted 2026-08-12 · Read on arXiv

Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zhang, Jianning Cui, Zijie Zhao, Haoyu Zhang, Kaiming Xu, Haodong Yang, Bowen Zhang, Jiahui Niu, Shaoting Zhu, Shiduo Zhang, Hang Zhao

Galaxea

cs.RO, cs.AI

Submitted: 2026-08-12

Updated: 2026-08-13

Code: https://github.com/Stanford-ILIAD/openvla-mini

Project page: https://opengalaxea.github.io/G05

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 95/100

The gist: The paper introduces G0.5, a pretrained autoregressive Vision-Language-Action (VLA) model in which "a single transformer decoder emits reasoning and action tokens under a single objective." The

Terminology

Summary

The paper introduces G0.5, a pretrained autoregressive Vision-Language-Action (VLA) model in which a single transformer decoder emits reasoning and action tokens under a single objective. The authors argue against the prevailing VLM-as-encoder recipe that couples a pretrained VLM with a separately trained flow-matching action expert, which makes the VLM a context encoder rather than a decision-maker. Instead, they advocate focusing on the VLM backbone: a unified model with a single set of weights that generates both reasoning and actions within a single autoregressive token stream.

The paper states: We therefore return to the autoregressive formulation, but remove the source of its original inefficiency: excessive action tokenization. The key insight is that once reasoning and action share the same autoregressive stream, chain-of-thought can be trained as a native component of control.

The paper introduces a learning-based action codec that maps continuous action sequences from embodiments with different degrees of freedom, control frequencies, and morphologies into a shared token vocabulary. Unlike FAST, which applies a fixed DCT-based pipeline separately to each embodiment, this codec is learned end-to-end and cross-embodiment by design.

The approach decomposes "each robot into independent motion parts (e.g., left control, right control, lower body), and pad each part to a shared maximum dimensionality before training a residual vector quantization (RVQ) model over the grouped actions. A temporal contrastive objective improves token consistency across temporally adjacent motions."

The unified action space is 27-dimensional, partitioned as: (9) (1) (9) (1) (7). The model predicts only the motion parts that are actively involved in the current behavior, enabling sparse action prediction during inference, where inactive parts remain stationary without requiring additional token generation.

The model is trained "to optionally perform intermediate reasoning before action prediction across four self-describing targets—task decomposition (Subtask:), key-object localization (BBox:), motion planning (Trace:), and action hints (ActionHint:). Any subset can be emitted at each step, and we draw from 8 curated combinations (including a no-CoT baseline) per training step, all supervised within the same next-token objective."

The paper emphasizes: "Unlike CoTVLA, DualCoT-VLA, and related approaches that attach reasoning modules to VLM-as-encoder backbones, our CoT tokens share the decoder, context, and objective with the action tokens. Reasoning and action are therefore not separate stages, but coupled phases of one generative process."

The model follows pi0.7 and MEM by inserting factorized spatial and temporal attention modules every four layers within the Vision Transformer. This separable mechanism efficiently fuses historical context by sequentially mixing information across time steps and spatial patches. To bound latency, we discard all historical tokens at the final layer, and stochastically drop all historical frames during training to prevent overfitting.

The model is initialized from Qwen3.5 2B, a pretrained vision-language model. At inference, given multi-view RGB observations, an embodiment identifier, a natural-language task instruction, and a proprioceptive state, the model autoregressively generates a structured output that concludes in a sequence of discrete action codes.

Training uses the standard next-token cross-entropy loss, computed only over the generative segment. The paper stresses: this single loss jointly supervises CoT generation and action generation: there is no auxiliary regression objective or expert distillation in pre-training.

Pre-training covers 14 embodiments across diverse real-world and simulated robot ontologies. DROID data are excluded from foundation pre-training. The data mixture includes web VQA, embodied VQA, and in-house annotations generated by the autolabeling pipeline using multimodal model APIs such as Gemini 3 and Doubao Seed 2.0 Pro for language, SAM3 tracking for visual grounding, and forward kinematics for 2D end-effector traces.

After post-training on DROID data, G0.5 is deployed on a held-out Franka Research 3 platform with unseen environment and objects. G0.5 achieves 82.5% average success rate, outperforming pi0.5-DROID (57.5%) by 25.0 percentage points and MolmoAct2-DROID (52.0%) by 30.5 percentage points. G0.5 outperforms pi0.5 on all 10 tasks.

G0.5 achieves 87.3% average success rate, the highest among compared methods including pi0, pi0-FAST, pi0.5, GR00T-N1.5, StarVLA-GR00T, RoboBrain2.5-8B, MemoryVLA, EO-1, and Xiaomi-Robotics-0.

G0.5 achieves 93.3% average success rate (93.7% clean, 92.8% randomized), outperforming all baselines including Fast-WAM (91.8%) and LingBot-VA (92.2%).

G0.5 achieves 98.9% average success rate, with the strongest performance on the challenging Long suite (98.6%). This surpasses all listed baselines including Xiaomi-Robotics-0 (98.7%), LingBot-VA (98.5%), and Cosmos Policy (98.5%).

On 50 long-horizon household mobile manipulation tasks, G0.5 (4 epochs) outperforms the first-place solution by +20.4% using only a single checkpoint, whereas the competition winner relies on a set of four distinct checkpoints. Specifically:

  • G0.5 (1 epoch): 0.2904 Task Success Score

  • G0.5 (4 epochs): 0.3136

  • pi0.5 (4 epochs): 0.2626

  • RLC (1st place, 4 checkpoints): 0.2605

The paper notes: With only a single epoch of post-training, G0.5 already surpasses pi0.5 trained for four epochs by +10.6%. G0.5 leads on 29 out of 50 evaluated household tasks (58%) while pi0.5 leads on only 15 (30%).

Across six task-embodiment settings, G0.5 achieves 76.7% average success rate, compared with 53.3% for pi0.5 and 24.4% for GR00T-N1.7. G0.5 also achieves an average process score of 129.2, compared with 105.2 for pi0.5 and 68.9 for GR00T-N1.7. G0.5 achieves the highest success rate in five out of six settings.

G0.5 demonstrates strong zero-shot language following (65.6% language following, 59.4% task success). With 50H post-training, G0.5 reaches language following rates of 84.4% and task success rates of 75.0%, outperforming pi0.5 by 15.6 percentage points in language following and 9.4 percentage points in task success.

The paper finds "AR > FM across multiple dimensions: Converges faster, to a better optimum," Generalizes zero-shot, and Token likelihoods plug straight into RL. In GRPO fine-tuning experiments, the AR policy converges substantially faster, reaches a higher final success rate, and trains more stably, with lower variance, than the FM policy.

On single-stage PP Bench, CoT brings essentially no change to language following. However, on five-stage long-horizon tasks (Air Fryer and Cook Bacon), CoT lets the policy ground the relevant object and execute each sub-step more reliably, lifting AR's progress score from 2.4 to 3.8 on Air Fryer and from 1.5 to 3.4 on Bacon.

The paper reports preliminary qualitative indications that per-stage instruction wording such as adverbial qualifiers, spatial cues, or near-synonymous verbs visibly shifts policy behavior without retraining. For example, adding adverbial or spatial qualifiers (e.g., 'push it in hard', 'vertically') tended to bring the executed motion closer to the intended sub-goal.

The paper attributes long-horizon performance gains to factorized temporal attention in the vision encoder, noting that even in standard pick-and-place episodes, the robot frequently moves its base between grasp and place locations, causing consecutive frames from each camera to exhibit large visual changes.

The paper notes: "Drawer-insertion and semi-transparent cabinet tasks remain weak across both G0.5 variants, pointing to a sensing limit not closed by AR alone; and our visual memory captures only seconds of history, leaving long-horizon memory open to research. Lower-body actuation is represented in the unified action space, but is not evaluated separately in this work."

The paper concludes: "the path forward for VLA models is to let the VLM be what it was pretrained to be—an autoregressive reasoner that now also acts, remembers, and adapts in-context—rather than to design ever more sophisticated action experts on top of an underutilized backbone. The authors state: We hope this work re-establishes autoregressive modeling as a foundation for VLA and that the pretrained backbone we release provides a useful starting point for future work."

Improvements for AI systems

Improvement 1: Unified Autoregressive Decision-Maker with Native Chain-of-Thought

  • What to build: Replace the common VLM-as-encoder + flow-matching action head architecture with a single transformer decoder that autoregressively emits interleaved reasoning tokens (subtask, bounding box, motion trace, action hint) and discrete action codes under one next-token loss.

  • What the improved system can do: Reason and act in a single generative pass, eliminating the need for separate action experts or auxiliary regression losses. It can decompose long-horizon tasks into sub-steps, localize objects, plan motion traces, and execute actions with the same weights, enabling faster convergence, better zero-shot generalization, and direct integration with RL (e.g., GRPO) for stable policy improvement.

Improvement 2: Cross-Embodiment Learned Action Codec with Sparse Prediction

  • What to build: A residual vector quantization (RVQ) action codec trained end-to-end across embodiments, with per-part padding (left/right control, grippers, lower body) and a temporal contrastive objective for token consistency. Predict only active motion parts during inference.

  • What the improved system can do: Transfer skills across robots with different degrees of freedom, control frequencies, and morphologies without per-robot retraining. It reduces token generation by skipping inactive parts, enabling efficient real-time control on mobile manipulators and humanoids while maintaining high task success (e.g., 82.5% on DROID zero-shot).

Improvement 3: Factorized Visual Memory with Stochastic Frame Dropping

  • What to build: Insert factorized spatial-temporal attention modules every four layers in the vision transformer, discard historical tokens at the final layer, and stochastically drop all historical frames during training.

  • What the improved system can do: Retain short-term visual context (seconds) to handle large viewpoint changes (e.g., base movement during pick-and-place) without overfitting or latency blowup. This improves long-horizon task success (e.g., +20.4% over BEHAVIOR Challenge winner) and enables robust performance in dynamic, partially observable environments.

Improvement 4: Prompt-Driven Behavioral Steering via In-Context Reasoning

  • What to build: Leverage the autoregressive CoT stream to condition on per-stage instruction wording (adverbial qualifiers, spatial cues, near-synonyms) without any retraining or fine-tuning.

  • What the improved system can do: Adjust execution style (e.g., push it in hard vs. push gently, vertically vs. horizontally) on the fly, enabling users to steer robot behavior through natural language at inference time. This is useful for personalized assistance, safety constraints, and adapting to user preferences in real-world deployment.

Improvement 5: Single-Loss Joint Training for Reasoning and Action

  • What to build: Train the model with standard next-token cross-entropy over both CoT and action tokens, with 8 curated CoT combinations (including no-CoT) sampled per training step.

  • What the improved system can do: Eliminate the need for separate reasoning and action training pipelines, reducing engineering complexity and data requirements. It enables the model to learn when reasoning is beneficial (e.g., long-horizon tasks) and when to skip it (e.g., simple pick-and-place), improving sample efficiency and generalization across task complexities.

Improved System Capabilities Summary:

The resulting AI system can:

  • Operate across 14+ embodiments with a single pretrained backbone, zero-shot on unseen robots and environments.

  • Perform long-horizon mobile manipulation (e.g., 50-task household challenges) with 58% task leadership and +20% score over prior best.

  • Follow natural language instructions with fine-grained behavioral control (e.g., speed, orientation, force) via prompt steering.

  • Learn new tasks from as little as 50 hours of post-training, achieving 84.4% language following and 75% task success on pick-and-place.

  • Integrate seamlessly with reinforcement learning for iterative improvement, converging faster and more stably than flow-matching alternatives.

Abstract

The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert. This makes the VLM a context encoder rather than a decision-maker. We introduce G0.5, a pretrained autoregressive VLA in which a single transformer decoder emits reasoning and action tokens under a single objective. Three components make this tractable at foundation-model scale: a learnable cross-embodiment action tokenizer that maps heterogeneous robot actions into a shared vocabulary; a native chain-of-thought stream interleaving task decomposition, object grounding, and action hints with action tokens; and a visual memory module that injects multi-second history through the vision encoder. Because reasoning and action share a single set of weights, the pretrained VLM's capabilities carry over to physical behavior: the model follows instructions closely, and prompts directly steer action granularity, task horizon, and out-of-distribution scene handling without further training. Pretrained on a large collection of robot datasets together with VQA samples, G0.5 surpasses state-of-the-art models across 7 independent regimes: real-world fine-tuning on R1lite and R1pro robots (76.7% vs. 53.3% for pi 0.5 and 24.4% for GR00T-N1.7), the 2025 BEHAVIOR Challenge on 50 long-horizon household mobile manipulation tasks using a generalist policy (31.4% vs. 26.3% for pi 0.5 and 26.1% for the challenge winner), DROID post-training followed by zero-shot transfer to an unseen environment and objects (82.5%), a language-following Pick-and-Place benchmark, LIBERO (98.9%), RoboTwin 2.0 (93.3%), and SimplerEnv-Bridge (87.3%).

Sources

Related papers