STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets

arXiv:2610.00003 · cs.CV, cs.AI, cs.LG, cs.RO · Submitted 2026-05-21 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "STATERA: Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets".

Jane: Inferring hidden physical properties from motion remains challenging for foundation models, especially when dealing with opaque, asymmetric rigid bodies where surface cues are unreliable.

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, to wrap things up today, we've been talking about STATERA’s "Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets." This paper is all about using motion data from video to figure out hidden physical properties of objects without actually seeing the mass distribution.

Jane: That's right, Tom. The authors are essentially showing how a specific AI architecture can use temporal data—the frozen parts of the video—to isolate the invisible forces at play, like rotational momentum, from just the visual mess on the surface.

Lu: What’s really striking about this work is their ability to use these temporal representations to compute velocity derivatives, which lets them encode momentum and inertia directly into a latent space. It’s an elegant way to inject Newtonian mechanics into the system.

Meng: From an engineering standpoint, it means we can design systems that don't just track pixels but actually predict how a heavy object will behave when it tumbles or falls, which is a much more robust foundation for any physical interaction we build.

Lalam: This pushes our culture forward because it shows vision AI moving away from simple pattern matching toward understanding actual physical laws, which helps us create more realistic simulations and models.

Tom: Exactly, Lalam. The performance numbers they shared are really telling; achieving a significant Physics Capture Ratio in zero-shot real-world transfers is a solid result when compared to traditional spatial models.

Jane: It’s not just about the accuracy; it’s about the method, which uses something called phase awareness to guide the prediction away from statistically safe guesses and toward the true physical direction.

Lu: And that focus on directional phase information, rather than a simple spatial centroid, is what allows the model to succeed where others fail with asymmetric payloads.

Meng: I'm thinking about how this translates practically; if we can get this kind of accuracy for grasping unknown objects, it opens up entirely new avenues for robotic manipulation in messy environments.

Lalam: And culturally, it expands what we consider possible for AI vision systems, showing that complex physical reasoning is achievable through motion analysis rather than just static image recognition.

Tom: So, the core message here is that frozen temporal data gives us the necessary tools to disentangle inertia from geometry, setting a new standard for how AI can infer hidden physical properties.

Jane: It really boils down to using motion dynamics as the primary signal, which makes these models much more grounded in the actual physics of the world when they encounter complex shapes.

Lu: This work definitely sets a benchmark for how we can use learned temporal representations to capture continuous physical behavior in visual data, opening up so many creative avenues for future research.

Meng: I'm looking forward to seeing how this specific methodology translates into reliable, real-time grasping solutions that handle the complexity of asymmetric objects in practice.

Conclusion: Tom: So, we're closing out our discussion on STATERA, which is officially titled "Hidden Mass Estimation via Zero-Shot Sim-to-Real Kinematics using Frozen Temporal Tubelets." This paper tackles the challenge of finding the center of mass in objects when you can’t see it directly.

Jane: That's right, Tom; the authors are essentially showing how they use specific temporal features from video to predict where an object's hidden mass is located. They separate the movement physics from just looking at what’s on a screen.

Lu: The core idea is that by freezing certain parts of the temporal data, they can isolate rotational momentum and inertial forces, which is super useful for understanding how things move dynamically in the real world.

Meng: From an engineering standpoint, this means we can build systems that don't rely on perfect visual geometry; they can learn to guess mass distribution based on physical motion patterns. That robustness is what matters most for real-world deployment.

Lalam: I think the implication here is huge because it moves us toward AI that understands the underlying mechanics of physics, which will help create simulations that are far more realistic than anything we've seen before in vision systems.

Tom: Exactly, Lalam; this isn't just about finding a dot on a screen, it’s about modeling actual physical laws using motion cues. The authors really show how separating inertial dynamics from visual geometry is the trick here.

Jane: It’s simple to think of it as teaching an AI to "feel" the momentum of an object instead of just memorizing its shape and color. That phase awareness they introduced is key for that intuitive understanding.

Lu: And that focus on directional phase information, rather than a static spatial guess, is what allows STATERA to succeed with unevenly weighted payloads where traditional methods get confused.

Meng: When I think about the impact on our startup’s work, this suggests we can drastically reduce the amount of manual labeling we need for complex physical interactions because the AI can infer that physics itself.

Lalam: It expands what we consider possible for AI vision systems by showing that complex physical reasoning is achievable through analyzing how things move, not just what they look like in a single frame.

Tom: So, the main message is that frozen temporal data gives us the necessary tools to disentangle inertia from geometry when estimating hidden mass. It sets a new standard for inferring physical properties from video motion.

Jane: It really boils down to using those motion dynamics as the primary signal, making these models much more grounded in how objects actually behave in the physical world when they encounter tricky shapes.

Lu: This work definitely sets a benchmark for using learned temporal representations to capture continuous physical behavior in visual data, opening up so many creative avenues for future research into AI mechanics.

Meng: I'm looking forward to seeing how this specific methodology translates into practical systems that can handle the complexity of asymmetric objects in real-time manipulation tasks.

cs.CV, cs.AI, cs.LG, cs.RO

Submitted: 2026-05-21

Updated: 2026-05-21

Comments: 17 pages, 7 figures, 3 tables. Preprint

Code: https://github.com/Animesh-Varma/STATERA

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 82/100

The gist: Inferring hidden physical properties from motion remains challenging for foundation models, especially when dealing with opaque, asymmetric rigid bodies where surface cues are unreliable.

Key concepts

Center of Mass (CoM)
The CoM is the object's true center of mass, which can be hidden from view. This paper aims to locate this invisible point by analyzing how the object moves and rotates in a video sequence, rather than just looking at its visible surface points.
Frozen Temporal Representations
These are static or partially frozen representations of the video's time dimension. By freezing these, the model can focus on inertial dynamics—like momentum transfer during freefall—instead of being distracted by rapidly changing visual geometry frame-by-frame.
Visual-Kinematic Aliasing
This occurs when standard tracking methods confuse visual cues (like surface points) with actual physical motion. STATERA is designed to prevent this by using temporal derivatives to separate the true inertial dynamics from misleading visual features, leading to a more accurate mass estimation.

Terminology

Summary

Inferring hidden physical properties from motion remains challenging for foundation models, especially when dealing with opaque, asymmetric rigid bodies where surface cues are unreliable. This work introduces STATERA, an architecture that extracts unobservable Center of Mass (CoM) coordinates from raw video by leveraging frozen temporal representations to separate inertial dynamics from visual geometry.

The gist

STATERA successfully isolates the directional phase of the offset Center of Mass, demonstrating that frozen temporal representations can better separate inertial dynamics from visual geometry for hidden-parameter estimation.

Problem Formulation and Motivation

The primary challenge addressed is localizing the Center of Mass (CoM) of an object with an uneven internal mass distribution absent direct visual indicators. Standard spatial foundation models and point-tracking architectures fail because they rely on spatial Euclidean metrics, leading to visual-kinematic aliasing, where they track spurious visual features rather than recovering the true hidden mass. While geometric centroids are a valid heuristic for uniformly dense objects, they fail for asymmetric payloads. The paper highlights that standard surface point-trackers are mathematically ill-posed for internal volumetric tracking because surfaces undergo self-occlusion during tumbling, meaning models conflate visual appearance with physical mass.

STATERA Architecture

STATERA employs a 2.5D multi-task architecture built upon a partially-frozen V-JEPA backbone to process video sequences. The key components of the architecture include:

  1. A partially frozen V-JEPA 2.1 backbone, which observes rotational torque and momentum transfer that occurs during freefall and impact.

  2. A temporal tubelet mixer utilizing a 1D Convolution (kernel = 3) across the temporal dimension of the latent tubelets to compute temporal velocity derivatives across the sequence. This operation is performed in the continuous latent space to preserve smooth dynamics.

  3. A Spatial Preservation Decoder that reshapes patches into their native grid using Stacked ConvTranspose2d layers upscale this feature map to 64×64, preserving local spatial equivariance and enabling sub-pixel acuity.

  4. A bifurcated output structure: a continuous spatial heatmap (Head A) and a regularizing 1D Z-Depth scalar (Head B).

Training Dynamics and Regularization

To train the network to rely on inertial momentum rather than frame-by-frame visual detection, several regularization techniques are employed:

  1. A multi-task loss combining Kullback-Leibler Divergence for the spatial heatmap and Huber Loss for the Z-depth regularizer.

  2. Temporal feature dropout applied after the 1D convolution to prevent discontinuous input spikes while requiring interpolation of latent physical priors.

  3. The use of a Crescent Target, which is formed by intersecting a spatial radial prior with a Von Mises angular phase mask, to enforce rotational phase information and prevent visual-kinematic aliasing.

  4. Curriculum label smoothing via variance decay is applied, where the target systematically shrinks from a broad Gaussian annulus to an impulse distribution over the training curriculum, preventing the model from overfitting to static visual heuristics during ambiguous initial frames.

Evaluation and Key Findings

The performance of STATERA is evaluated using several metrics designed to isolate physical disentanglement:

  1. Normalized Center of Mass Error (N-CoME): Measures spatial Euclidean distance normalized by the object's bounding diagonal, designed to mitigate perspective scaling issues.

  2. Normalized Kinematic Jitter (J˜): Penalizes non-physical, high-frequency acceleration spikes caused by visual aliasing.

  3. Physics Capture Ratio: Tracks the percentage of absolute distance the prediction successfully moves away from the visual centroid toward the true hidden offset mass along the true kinematic vector.

The results show that STATERA-50K-Crescent (Phase-Aware) achieves a substantial 41.0% Physics Capture Ratio in zero-shot real-world transfer, outperforming spatial foundation models like DINOv2 (which achieve only 2.6%). While this model exhibits a Vector Overshoot artifact, its ability to isolate the directional phase of the hidden mass is demonstrated as superior to methods that collapse toward statistically safe centroids, such as STATERA-50K-Sigma. The paper concludes that frozen temporal representations are critical for separating inertial dynamics from visual geometry.

Limitations and Future Directions

Limitations include reliance on a partially frozen V-JEPA backbone due to compute constraints, the potential degradation of performance for objects caught mid-air without collision events, and the Euclidean vector overshoot artifact which suggests predicted CoM may occasionally fall outside physical bounds. Future work aims to investigate training foundational video architectures entirely from scratch on physical dynamics datasets and exploring conditioning the network on static fiducial reference markers to correct the Euclidean magnitude overshoot. The spatial variance of the probability mass is also suggested as a potential confidence calibration metric for rejecting unsafe grasps.

Improvements for AI systems

As a fastidious researcher, I have analyzed the STATERA paper and identified several high-leverage improvements for existing AI systems, particularly in perception, robotics, and autonomous navigation.

Here are the specific improvements that can be made:


  1. Improve the robustness of object manipulation in unstructured environments by integrating a latent kinematic disentanglement module into foundation models.

  2. Enhance the generalization of vision models to infer hidden physical properties (like Center of Mass or internal mass distribution) from raw video, moving beyond surface-level geometry recognition.

  3. Develop more reliable autonomous navigation and grasping systems capable of operating safely with unknown payloads or asymmetric objects, especially in scenarios involving self-occlusion or tumbling motion.

  4. Create a novel training paradigm for foundation models that explicitly penalizes settling-state bias to ensure the model learns true physical dynamics rather than memorizing static resting configurations.

  5. Improve the calibration and uncertainty estimation of monocular vision systems by leveraging latent spatial variance as a proxy for prediction confidence, leading to safer decision-making.

The improved AI systems can do the following:

  1. Autonomous robotic arms could safely grasp and manipulate objects (e.g., tools, containers) with unknown or asymmetric weights and internal mass distributions without requiring explicit sensor measurements of those properties or manual calibration.

  2. Self-driving vehicles could better predict the dynamic behavior of passing obstacles (like tumbling debris or irregularly shaped cargo) by understanding the object's true momentum and inertia, leading to safer collision avoidance maneuvers in complex scenarios.

  3. Vision-based inspection systems (e.g., quality control for manufacturing) could accurately assess internal structural integrity or material composition based on subtle motion cues that reveal hidden mass asymmetries, rather than relying solely on external visual appearance.

  4. Foundation models trained with this methodology will exhibit superior physics awareness, meaning they will not default to the geometric centroid when faced with ambiguous visual data; instead, they will correctly predict the direction of momentum required for physical interaction.

  5. Robotic systems can use the model's inherent uncertainty (spatial variance) to dynamically adjust grasping force or trajectory planning, allowing for safe grasping decisions that reject high-risk predictions where uncertainty suggests an unbounded Euclidean error (overshoot).

Sources

Related papers