Open-CHOIR: Open-World Contact-Aware 4D Hand-Object Interaction Reconstruction

arXiv:2605.20992 · cs.CV · Submitted 2026-05-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Open-CHOIR: Open-World Contact-Aware 4D Hand-Object Interaction Reconstruction".

Jane: Reconstructing 4D hand-object interactions (HOI) from monocular RGB videos in open-world settings remains difficult because existing methods often fail under clutter, occlusion, and unseen object geometries.

Tom: First, who's behind it and why it matters.

Title and authors: Tom: The title itself, Open-CHOIR: Open-World Contact-Aware 4D Hand-Object Interaction Reconstruction, really tells us the core value proposition here is achieving contact awareness in open environments <ref:2605.20992#pg0,Contact-Aware 4D Hand-Object Interaction Reconstruction>. It’s not just about seeing a hand and an object; it’s about understanding the physical interaction happening between them over time.

Jane: Exactly, Tom; they are aiming to provide something that captures the three dee hand motion, the object's shape with its 6D pose trajectory, and crucially, when and where contact actually occurs in these messy scenes <ref:2605.20992#pg0>. They want a complete picture of how a hand engages with an object.

Lu: The authors are pulling together techniques from different areas—from 2D cues to generative models like flow-matching—to tackle this coupled problem; it shows a very broad approach to tackling complex perception tasks <ref:2605.20992#pg1>.

Meng: I'm curious about the implementation details, how they managed to keep all those separate stages coordinated without the entire system collapsing under the complexity of real-world data.

Lalam: It suggests that for AI to truly understand physical interaction, it needs more than just visual recognition; it needs an explicit mechanism to model physical coupling. This moves us toward richer AI systems capable of genuine manipulation understanding.

The paper's summary: Tom: So, the paper summarizes Open-CHOIR by breaking the task down into three main stages to achieve this reconstruction: first, they establish a coarse, contact-agnostic sequence; then, they use a flow-matching model to spatially correct those initial misalignments; and finally, they refine everything through contact-aware optimization.

Jane: That’s the essence of it—they start simple and progressively add physical constraints until the final result is physically consistent across all dimensions. They show how this staged strategy helps diagnose issues more easily than a single monolithic model would.

Lu: The Stage one involves extracting 2D interaction cues, recovering initial three dee geometry using something like SAM-three dee-Objects for the object mesh, and then tracking the motion into a preliminary 4D hand trajectory <ref:2605.20992#pg1>.

Meng: Reconstructing that initial metric scale from an anchor frame seems like a delicate part of Stage one; if that anchor is off, everything downstream will be distorted <ref:2605.20992#pg0>. How robust is that initial scaling recovery?

Lalam: The paper highlights how this process yields reusable interaction primitives, which means we can extract structured data—the hand motion, the object's shape over time—which is super valuable for building better simulators and planning tools.

The paper's improvements: Tom: Their main improvement seems to be explicitly coupling hand and object geometry through this staged pipeline, specifically by using contact as an explicit signal in Stage three to enforce physical consistency across the entire sequence <ref:2605.20992#pg0>.

Jane: They introduce several specific losses in Stage three like the contact loss which pulls active hand vertices toward object anchors, plus a penetration loss to stop things from passing through surfaces <ref:2605.20992#pg0>. This is what really gives it that physical grounding they're aiming for.

Lu: The spatial rectification module in Stage two is particularly innovative; using a flow-matching prior to predict ray-depth corrections is a clever way to rectify the relative placement errors before contact correspondences are even started, which I think significantly cleans up the input for the final optimization <ref:2605.20992#pg1>.

Meng: That flow-matching aspect sounds computationally demanding, so I wonder how they balanced that generative modeling with the need for real-time or near real-time performance when dealing with these complex 4D trajectories <ref:2605.20992#pg0>.

Lalam: The fact that they can model contact evidence as a trainable quantity within that optimization loop is really powerful because it means the hand's position can actively change based on what it’s touching, rather than just being passively placed there initially.

Conclusion: Tom: So, to wrap up Open-CHOIR: Open-World Contact-Aware 4D Hand-Object Interaction Reconstruction, the authors demonstrate a way to achieve high-fidelity traces in open environments by separating the problem into analysis, spatial correction, and contact optimization <ref:2605.20992#pg0,Contact-Aware 4D Hand-Object Interaction Reconstruction>.

Jane: They show that this staged reconstruction strategy yields physically consistent 4D hand–object interactions by using specific losses like penetration loss and contact loss to refine the sequence <ref:2605.20992#pg0>. It’s a solid methodology for tackling complex perception challenges.

Lu: The implication is that we can start mining real-world interaction data—articulated hand motion, object shape with 6D pose over time, and contact evidence—which opens up avenues for scene-aware synthesis and planning applications in unstructured settings <ref:2605.20992#pg0,articulated hand motion, object shape with 6D pose over time, and>.

Meng: I think the practical impact lies in creating a pipeline that produces reusable primitives; if we can extract these structured interactions reliably from messy video, it makes building robust AI agents much more feasible.

Lalam: For me, the real culture shift is seeing AI systems capable of modeling physical contact meaningfully; this moves us toward building systems that don't just see things but truly understand how they interact physically in the world around them.

The Chinese University of Hong Kong · University College London

cs.CV

Submitted: 2026-05-20

Updated: 2026-10-07

Importance score: 86/100

The gist: Reconstructing 4D hand-object interactions (HOI) from monocular RGB videos in open-world settings remains difficult because existing methods often fail under clutter, occlusion, and unseen object

Key concepts

Stage 1 (Open-World HOI Analysis)
The first stage establishes a basic sequence by extracting 2D cues like bounding boxes and joints. It reconstructs an initial 3D object mesh and tracks the hand motion to create a rough, contact-agnostic trajectory, setting up the foundation for refinement.
Stage 2 (Generative HOI Spatial Rectification)
This stage corrects depth misalignments between the hand and object. It uses a flow-matching model to predict corrections along the camera ray, mapping noisy depth offsets to a rectified relationship that is geometrically plausible before contact points are identified.
Contact-Aware Optimization (Stage 3)
The final stage refines the sequence by minimizing five losses. These include contact loss to pull hand vertices toward object surfaces, penetration loss to prevent hands from passing through objects, and temporal loss to ensure smooth, physically consistent motion.

Terminology

Summary

Reconstructing 4D hand-object interactions (HOI) from monocular RGB videos in open-world settings remains difficult because existing methods often fail under clutter, occlusion, and unseen object geometries. This paper presents CHOIR, a Contact-aware HOI Reconstruction framework that uses contact as an explicit coupling signal to recover articulated hand motion, object shape with 6D pose over time, and when/where contact occurs.

How it works

CHOIR employs a staged reconstruction strategy consisting of three main stages: Stage 1 (Open-World HOI Analysis), Stage 2 (Generative HOI Spatial Rectification), and Stage 3 (Contact-Aware Optimization). This separation allows the pipeline to be easier to diagnose, as each stage has a clear input, output, and failure mode before the final joint refinement.

Stage 1 focuses on establishing a coarse, contact-agnostic sequence. This involves:

  1. Extracting 2D interaction cues, such as the hand bounding box, chirality, and 2D joints, along with the interacting-object bounding box.

  2. Recovering 3D geometry using methods like SAM-3D-Objects to reconstruct a canonical object mesh and an initial metric scale from an anchor frame.

  3. Tracking the motion using techniques like VIPE for camera estimates and Dyn-HaMR to stabilize the initial MANO hand estimates into a 4D hand trajectory.

  4. Performing Coarse isolated sequence fitting by optimizing MANO parameters using losses such as Lh2D, Lhdepth, and terms that suppress invalid poses (Lh anat).

How it works

Stage 2 addresses the issue where initial estimates misalign due to wrong relative depth by performing a spatial correction. This stage uses a flow-matching prior to predict ray-depth corrections for interaction frames, aiming to rectify hand-object relative placement and refine the geometry into a contact-plausible relation.

  1. The model is conditioned on the current 3D hand joints, viewing direction, object scale, and the camera ray.

  2. A flow-matching model predicts a scalar correction along the viewing ray to map a noisy depth offset to a rectified offset consistent with the object geometry.

  3. Initial contact correspondences are then read out from this rectified geometry as hand contact vertices and barycentric object anchors, which serve as initialization evidence for optimization, rather than final outputs.

How it works

Stage 3 refines the sequence into a physically consistent 4D HOI trace using a contact-aware joint optimization. This stage minimizes five weighted objectives:

  1. Contact loss (Lcon), which pulls active contact vertices toward object-surface anchors, using a soft correspondence term with weights applied to top-K candidate anchors.

  2. Penetration loss (Lpen), which is one-sided and pushes hand vertices that lie inside the object back toward the surface, defined by the signed penetration depth.

  3. Silhouette loss (Lsil) to align rendered silhouettes to amodal masks.

  4. Anchor loss (Lanc) for 2D hand joint, anatomy, pose, and object pose consistency.

  5. Temporal loss (Ltemp) to enforce smooth motion across the sequence using first- and second-order finite differences on hand joints and object centers.

How it works

The framework is trained using a large-scale synthetic dataset called GraspPair, which contains simulated grasps over diverse object shapes of randomized scales and initial placements. This data is pruned using a physics engine (PyBullet) to retain only stable under gravity and perturbations configurations, yielding contact-valid target grasps that supervise the flow-matching model.

How it works

The evaluation involves comparing CHOIR against state-of-the-art baselines on controlled HO3D videos and challenging in-the-wild benchmarks like TASTE-Rob and self-captured videos. Performance is assessed using both metric accuracy (e.g., root-relative MPJPE for hand joints, Chamfer Distance for object reconstruction) and reference-free metrics such as physical plausibility (H-O Dist. and Pen. Ratio), video alignment (mIoU between rendered silhouettes and tracked 2D masks), and temporal smoothness (Acch and Acco). Ablation studies confirm that key components, such as the Stage 2 rectification module, are crucial for improving hand-object alignment before contact correspondences are constructed. CHOIR achieves superior performance in these metrics across both controlled and real-world datasets.

Improvements for AI systems

As a fastidious researcher, I have analyzed the CHOIR framework presented in this paper. The core innovation lies in explicitly coupling hand and object geometry through a staged reconstruction pipeline: open-world analysis, generative spatial rectification, and contact-aware optimization.

Here are the specific improvements that can be made to AI systems using the CHOIR methodology:


The improved AI system will be capable of performing high-fidelity, physically grounded 4D interaction reconstruction from monocular video in open-world settings. It moves beyond simple pose estimation or object tracking by recovering a complete, temporally consistent interaction trace (articulated hand motion + object shape + 6D pose trajectory + explicit contact evidence).

The specific improvements and capabilities are:

  1. [Hand-Object Interaction Reconstruction in Open-World Scenes]: The system can reconstruct the full 4D HOI sequence from a single monocular RGB video, even when the object is unseen or the scene is cluttered.

  2. [Robust Physical Plausibility Enforcement]: By incorporating contact constraints (via Stage 3 optimization) and penetration losses, the system ensures that reconstructed hand-object poses are physically valid—preventing interpenetration artifacts and floating hands—which standard methods often fail to achieve.

  3. [Contact-Driven Spatial Correction]: The generative HOI spatial rectification module (Stage 2) explicitly predicts ray-depth corrections using a flow-matching prior to rectify initial misalignments between the hand and object geometry before contact correspondences are established, ensuring contact search is based on rectified geometry rather than noisy estimates.

  4. [Scalable Interaction Mining]: The system produces reusable 4D interaction primitives (articulated hand motion, object shape with 6D pose over time, and when/where contact occurs). This enables scalable mining of real-world interactions for downstream tasks like scene-aware synthesis and planning.

  5. [Phase-Aware Interaction Modeling]: The framework can detect and model different interaction phases (pre-static, approach, manipulation, release) across the 4D trajectory, allowing for more nuanced understanding of the interaction lifecycle.

  6. [Contact Evidence Integration]: The system explicitly models contact evidence as a trainable quantity within a joint optimization loop (Stage 3), allowing active hand vertices to vary during interaction based on real-time physical constraints, rather than relying on static initial correspondences.

In summary, this system transforms ambiguous visual data into structured, physically consistent interaction traces, which is critical for applications requiring true scene understanding and planning in unstructured environments.

Sources

Related papers