RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

arXiv:2608.09853 · cs.RO, cs.CV, cs.LG · Submitted 2026-08-10 · Read on arXiv

DAMO Academy, Alibaba Group · Hupan Lab

cs.RO, cs.CV, cs.LG

Submitted: 2026-08-10

Updated: 2026-09-29

Comments: 23 pages, 5 figures

Code: https://github.com/alibaba-damo-academy/RynnValuehttps:

Project page: https://alibaba-damo-academy.github.io/RynnValue.github.iohttps://github.com/alibaba-damo-academy/RynnValuehttps://huggingface.co/collections/Alibaba-DAMO-Academy/rynnvaluehttps://www.modelscope.cn/collections/DAMO_Academy/RynnValue

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance Summary This paper introduces RynnValue, an open-source value foundation model for robotic manipulation that replaces

Terminology

Summary

RynnValue: Scaling Robotic Value Foundation Models with Temporal Distance

Summary

This paper introduces RynnValue, an open-source value foundation model for robotic manipulation that replaces trajectory-internal progress with temporal distance as its supervision target. The authors argue that general-purpose reward models are increasingly the bottleneck for scaling robot learning, yet existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. RynnValue instead estimates the directed cost-to-go from an observation to the language-specified goal, using temporal distance as its scaling target.

Core Approach

RynnValue adopts temporal distance as its learning objective, defined as the directed temporal cost from the current observation to the language-specified goal. Under a minimum-time objective, this corresponds to the hitting-time cost-to-go, yielding clear directionality and task conditioning. Because temporal-distance labels can be derived directly from timestamps once a completion cutoff is identified, RynnValue scales to over 7,000 hours and roughly 3M instruction-conditioned clips without preference or progress annotations. The training corpus spans diverse embodiments, viewpoints, and task families, including real-world, simulated, and egocentric trajectories.

Architecture

RynnValue is built on RynnBrain and further pretrained on large-scale robot data. Given an instruction l, embodiment metadata m, and a sequence of K observations (with K=8 in the main setting), the model jointly predicts absolute temporal distance for each observation and relative temporal displacement between consecutive observations. Key architectural elements include:

  • Grouped temporal queries: Each temporal prediction uses a group of N=8 repeated query tokens, allowing richer temporal representations. The multimodal sequence is: x = [m, l, I t1, V1, I t2, R1, V2,..., I tK, R K-1, VK, p ver], where V i are absolute-value query groups, R i are relative-value query groups, and p ver is a verification prompt.

  • Dual distributional heads: Two specialized heads produce continuous absolute and relative temporal estimates. Targets are discretized into 256 symlog-spaced bins over [0,512] seconds (absolute) and [-256,256] seconds (relative), trained with two-hot targets over adjacent bins.

  • Value-isolation attention: This prevents absolute and relative temporal queries associated with different observations from attending to one another, while queries belonging to the same observation remain mutually visible. Context tokens are also prevented from attending to temporal-query tokens, ensuring each temporal estimate is grounded in the task instruction and its available visual context.

  • Natural-language supervision: The model retains autoregressive language capability, generating video descriptions, task-matching judgments, and success judgments after processing the visual sequence.

Training Recipe

To make temporal-value learning reliable at scale, RynnValue incorporates complementary shortcut-suppression designs:

  1. Random temporal sampling: Observations are sampled at irregular timestamps, breaking the correspondence between sequence position and temporal distance.

  2. Temporal-order shuffling: Half of the training sequences are independently sampled without temporal sorting, while the remainder follow a forward-biased temporal walk with occasional backward transitions (rewind probability 0.3). Relative temporal targets can therefore be either positive or negative.

  3. Value-isolation attention: Prevents extrapolation from other value-query representations.

  4. Instruction-mismatch augmentation: For 10% of training samples, the original instruction is replaced with one from a different trajectory, teaching the model to detect instruction-video mismatches. The absolute temporal-distance loss is masked for these samples, while the relative loss is retained.

The joint training objective combines three cross-entropy losses: absolute temporal-distance loss (L abs), relative temporal-distance loss (L rel), and natural-language loss (L lang with λ lang=2).

Data Preparation

The training mixture includes data from AgiBot, EgoDex, Galaxea Open-World, InternData-A1, Open X-Embodiment, RDT, RoboCOIN, RoboMIND, RoboTwin, and Soft-FOLD. Before subtask expansion, the corpus contains 1.67M original episodes, converted into over 3M instruction-conditioned trajectory segments after subtask segmentation and cutoff relabeling. A source-aware curation pipeline addresses annotation issues including non-English task annotations, placeholders, data-quality metadata, and pure-motion instructions, retaining 83.35% of trajectory units while preserving 98.99% of unique instructions.

Benchmark Results

On the RBM-EVAL-OOD test suite (976 trajectories from six out-of-distribution datasets), RynnValue-8B achieves an average Kendall's τa of 0.675, surpassing the fully preference-supervised state of the art (Robometer at 0.655) and more than doubling a progress-only counterpart (0.292). RynnValue-4B reaches 0.670. Among methods trained without explicit trajectory-level preference supervision, RynnValue variants attain the best results on all six datasets, improving the strongest preference-free prior average from 0.502 to 0.670 (4B) and 0.675 (8B).

In instruction-trajectory alignment analysis, RynnValue produces the clearest diagonal structure and achieves the highest normalized diagonal margin of 0.79, outperforming the strongest baseline at 0.67.

Ablation Study

The full model achieves the highest average Kendall's τa of 0.675. Removing temporal-order shuffling causes the largest degradation (0.189), followed by uniform sampling (0.379), removing value-isolation attention (0.482), removing natural-language supervision (0.537), and removing relative temporal-distance supervision (0.627).

Scaling Analysis

Task-diversity scaling reduces error monotonically across the entire range, while episode-volume scaling saturates almost immediately. This demonstrates that the heterogeneous data recipe contributes not merely additional samples, but the diversity required for general-purpose temporal-value learning.

Real-World Policy Learning

RynnValue is evaluated on four real-world manipulation tasks using a dual-arm Franka robot: Bread Basket Placement, Steak Serving with a Spatula, Box-in-Drawer Placement, and Bimanual Box Transfer. Converted into dense rewards via potential-based shaping, RynnValue raises:

  • Online RL: average success from 52.5% (Robometer) to 72.5%, with sparse rewards at 48.8%

  • Offline RL: average success from 63.8% (Robometer) to 82.5%, with SFT at 23.8%

RynnValue consistently achieves the highest success rate across all four tasks in both online and offline RL, with particularly pronounced improvements on Steak Serving with a Spatula (+30 percentage points online) and Bimanual Box Transfer (+35 percentage points online). It also provides a strong balance between task success and execution efficiency, achieving 100% offline success on Bread Basket Placement using 16.8 action chunks on average.

Conclusion

The authors conclude that temporal distance is both a scalable supervision target for value foundation models and a practical reward interface for generalist robot policies. Future work aims to extend RynnValue toward broader temporal horizons, richer value semantics (incorporating task-specific costs such as energy, safety, or precision), and more diverse embodiments including dexterous hands and mobile manipulation settings.

Improvements for AI systems

Improvements to AI Systems:

  1. Temporal-Distance Reward Modeling for Generalist Policies
  • Replace task-specific reward functions or preference-based reward models with a unified temporal-distance value model that estimates directed cost-to-go toward language-specified goals.

  • The improved AI system can provide dense, scalable reward signals for any manipulation task without manual reward engineering, enabling faster and more stable reinforcement learning across diverse embodiments and unseen tasks.

  1. Shortcut-Suppression Training for Robust Temporal Estimation
  • Incorporate random temporal sampling, temporal-order shuffling, value-isolation attention, and instruction-mismatch augmentation into the training of any sequence-based value or reward model.

  • The improved AI system can avoid learning spurious correlations between sequence position and temporal progress, making its reward estimates reliable even with irregular, shuffled, or mismatched observation sequences—critical for real-world deployment where data is noisy and non-sequential.

  1. Dual-Head Distributional Temporal Prediction
  • Use grouped temporal queries with dual distributional heads (absolute and relative) and symlog-discretized targets to predict both global and local temporal displacements.

  • The improved AI system can output fine-grained, uncertainty-aware temporal estimates (e.g., 3.2 seconds to goal and 0.5 seconds between frames), enabling smoother reward shaping, better action chunking, and more precise progress tracking in long-horizon tasks.

  1. Cross-Embodiment and Cross-Viewpoint Generalization via Diverse Data Mixture
  • Train on a heterogeneous corpus spanning real, simulated, egocentric, and multi-embodiment trajectories (7,000+ hours, 3M clips) with source-aware curation.

  • The improved AI system can transfer value estimation across different robot morphologies, camera angles, and task families without retraining, reducing the need for embodiment-specific reward models and accelerating deployment on new hardware.

  1. Natural-Language Supervision for Value-Grounded Reasoning
  • Retain autoregressive language generation alongside temporal prediction, training the model to produce video descriptions, task-matching judgments, and success judgments.

  • The improved AI system can explain its reward decisions in natural language, detect instruction-video mismatches, and self-verify task completion—enabling interpretable and auditable reward signals for human oversight and debugging.

  1. Task-Diversity-Driven Scaling over Volume
  • Prioritize task diversity over episode volume during data collection, as scaling analysis shows diversity monotonically reduces error while volume saturates.

  • The improved AI system can achieve better generalization with fewer but more varied demonstrations, making it more sample-efficient and cost-effective for real-world robotic learning.

  1. Potential-Based Reward Shaping for Offline and Online RL
  • Convert temporal-distance predictions into dense rewards via potential-based shaping (e.g., difference in predicted temporal distance).

  • The improved AI system can boost both online RL (from 52.5% to 72.5% success) and offline RL (from 63.8% to 82.5% success) on real-world manipulation tasks, while also improving execution efficiency (fewer action chunks) and enabling sparse-reward tasks to be solved reliably.

  1. Out-of-Distribution Robustness for Reward Evaluation
  • Use temporal-distance supervision to achieve state-of-the-art Kendall's τa (0.675) on OOD benchmarks, outperforming preference-based methods.

  • The improved AI system can rank candidate trajectories for a given instruction more accurately than human-preference-based models, even on unseen datasets—useful for automated data filtering, trajectory selection, and reward model validation.

  1. Bimanual and Dexterous Task Support
  • Extend the value model to handle bimanual manipulation and dexterous hand tasks (e.g., Bimanual Box Transfer, Steak Serving with Spatula) by leveraging temporal distance as a universal interface.

  • The improved AI system can learn complex, coordinated multi-arm policies with dense temporal rewards, achieving high success rates (e.g., 100% offline on Bread Basket Placement) and enabling new capabilities in dexterous and mobile manipulation.

Abstract

General-purpose reward models are increasingly the bottleneck for scaling robot learning, yet the recipe for learning value-related capabilities from large-scale heterogeneous corpora remains underexplored. Existing approaches tie supervision to task-internal anchors such as preferences or normalized progress, none of which transfer cleanly across embodiments and data sources. We introduce RynnValue, an open-source value foundation model for robotic manipulation that replaces these anchors with temporal distance, the directed cost-to-go from an observation to the language-specified goal. Because temporal-distance labels can be derived directly from timestamps, RynnValue scales to over 7,000 hours and roughly 3M instruction-conditioned clips without preference or progress annotations. To make temporal-value learning reliable at scale, we combine random temporal sampling, temporal-order shuffling, and value-isolation attention, suppressing shortcuts that would leave predictions insensitive to failures and regressions. Trained without preference labels, RynnValue attains an average Kendall's tau a of 0.675 on RBM-EVAL-OOD, surpassing the fully preference-supervised state of the art (0.655) and more than doubling a progress-only counterpart (0.292), while generalizing zero-shot to unseen tasks, embodiments, and viewpoints. Converted into dense rewards via potential-based shaping, it raises real-world policy success from 52.5% to 72.5% online and from 63.8% to 82.5% offline. These results establish temporal distance as a scalable supervision target and practical reward interface for generalist robot policies.

Sources

Related papers