Reconstructing Humans and Objects in Interaction using Large Reconstruction Models

arXiv:2608.27407 · cs.CV · Submitted 2026-08-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "Reconstructing Humans and Objects in Interaction using Large Reconstruction Models".

Tom: Estimation of Human-Object Interactions in 3D (3D HOI) is a fundamental problem in 3D computer vision with applications in AR/VR, robotics, and embodied AI.

Jane: First, who's behind it and why it matters.

Paper summary: Tom: Welcome back to the radio show, everyone! Today we’re digging into a paper that’s making waves in the field of three dee computer vision <ref:2608.27407#pg0>. We’ve got a deep dive into "Reconstructing Humans and Objects in Interaction using Large Reconstruction Models." Jane, you were looking at this abstract earlier; what is the main idea here?

Jane: Well, Tom, it looks like the core thesis revolves around tackling the notoriously difficult problem of three dee human-object interaction estimation <ref:2608.27407#pg0>. The authors claim that they can recover these interactions in three dee using Large Reconstruction Models instead of relying on traditional methods that depend heavily on contact information or reprojection constraints <ref:2608.27407#pg0>.

Lu: That’s fascinating, Jane. It sounds like they’re shifting the focus from trying to solve depth ambiguities directly to interpreting a richer geometric structure provided by these models <ref:2608.27407#pg0>. The idea that LRMs offer a "powerful geometric scaffold" is what really caught my attention; it simplifies things significantly.

Meng: From an engineering standpoint, simplifying the input from a complex reconstruction problem to interpreting a mesh sounds promising for scalability, but I’m wondering if that mesh interpretation step introduces new failure modes when dealing with highly occluded scenes.

Lalam: From my perspective as a vision model, the concept of using an LRM output as a direct scaffold is very compelling because it bypasses some of the noise inherent in single-image reconstruction <ref:2608.27407#pg1>. It suggests that the visual understanding provided by these large models is inherently good at preserving relative spatial arrangements, which is exactly what we need for interaction modeling.

Tom: Exactly, Lalam! So, to clarify for our listeners, the paper introduces a framework called MILO that uses an LRM to predict a combined human-object mesh from just one image. Jane, can you elaborate on what MILO actually claims it achieves?

Jane: The authors claim MILO works by taking that predicted mesh and then segmenting it into separate human and object parts. They then fit a specific parametric body model, the SMPL-H, to the human component using some keypoints derived from rendering that mesh through multiple virtual views <ref:2608.27407#pg1>.

Lu: And that whole process is supported by a one-way robust Chamfer loss used during fitting, which helps improve reconstruction in areas where things are occluded by comparing observed three dee points against the visible predicted vertices <ref:2608.27407#pg2>. That level of detail in the optimization strategy is quite sophisticated.

Meng: When they talk about fitting SMPL-H, are they just using those keypoints to get a pose, or is it more involved than just adjusting orientation and translation? I need to know the complexity for practical deployment.

Paper summary: Jane: It’s actually two stages for the SMPL-H parameters. First, they optimize the root fitting by minimizing an objective L rf which aligns the joints to those estimated three dee keypoints <ref:2608.27407#pg0>. Then, there's a pose fitting stage that optimizes all the parameters—shape, body pose, and hand pose—while using priors like VPoser and MANO <ref:2608.27407#pg2>.

Lalam: The use of those external priors like VPoser adds a layer of regularization that helps ensure the resulting human mesh looks anatomically plausible even when the input data is ambiguous <ref:2608.27407#pg1>. It’s about leveraging existing knowledge to guide the reconstruction when the visual evidence is sparse.

Tom: That sounds like a solid pipeline, Jane, but I want to circle back to the main claim: what makes this approach different from what was done before? What's the big win they’re pointing toward?

Jane: The primary contribution is that MILO demonstrates state-of-the-art performance on several established datasets like InterCap <ref:2608.27407#pg2>, HODome <ref:2608.27407#pg1>, and IMHD eighty. Crucially, they achieve this without needing ground truth contact information, which is a big deal because those datasets are often limited in that regard.

Lu: That lack of reliance on contact supervision is what opens up so many possibilities for real-world applications in AR and robotics where true contact labels are incredibly hard to get <ref:2608.27407#pg1>. It suggests the visual cues embedded within LRMs are sufficient for inferring these complex interactions.

Meng: So, if we look at the practical impact, does this mean we can finally build systems that understand physical relationships between objects and people without needing tedious manual labeling of every single point of contact? I’m curious about the development cost there.

Lalam: It could significantly improve how embodied AI agents perceive their environment. If an agent can robustly infer where a human's hand is relative to an object based on visual arrangement alone, its manipulation tasks become much more intuitive and less dependent on perfect pre-training data <ref:2608.27407#pg1>.

Tom: It sounds like the authors are really showing that the geometric scaffold from the LRM is robust enough to handle these complex scene interpretations. So, if we look at the overall presentation of "Reconstructing Humans and Objects in Interaction using Large Reconstruction Models," what’s their final message to us?

Jane: The authors emphasize three main contributions: first, proving how powerful LRMs are for this kind of reconstruction. Second, designing a method that fits a parametric human model to that LRM mesh. And third, achieving state-of-the-art results across multiple benchmarks using weaker information than previous work <ref:2608.27407#pg2>.

Paper summary: Lu: I think the most important conceptual point is reframing the problem entirely; they aren't trying to solve depth ambiguities with traditional reprojection objectives anymore, they are interpreting a mesh instead <ref:2608.27407#pg0>. That fundamental shift in approach is where the real potential lies for future research directions.

Meng: I agree that reframing the problem is smart because it makes the reliance on external shape repositories less critical for every single instance, as long as the LRM can provide a decent initial geometry. We need to see if this holds up when we move from controlled environments to, say, a busy street scene.

Lalam: If we can generalize this interpretation of the mesh structure across different object categories and human poses, it could fundamentally improve how complex scenes are understood by language models that integrate visual data <ref:2608.27407#pg1>. It’s about building a more physically informed world representation.

Tom: That’s a huge vision for embodied AI! So, wrapping up this segment on "Reconstructing Humans and Objects in Interaction using Large Reconstruction Models," the title itself points to the core function: reconstructing those interactions in three dimensions. What does that mean for us as listeners?

Jane: In simple terms, it means we are moving closer to systems that can accurately model not just where a person is, but precisely how they are interacting with physical items around them in a detailed three dee space <ref:2608.27407#pg0>. This moves us past just seeing an object and starting to understand the spatial relationship between two entities.

Lu: It’s about achieving that level of spatial reasoning that is currently very difficult for purely visual systems to maintain consistently <ref:2608.27407#pg1>. The paper shows a path where powerful generative models can provide the necessary geometric context for this reasoning.

Meng: From a practical impact view, this means fewer errors in robotic grasping tasks or virtual reality simulations where accurate interaction is key. We’re talking about more reliable physical interaction modeling in AI applications.

Lalam: It could enhance the cultural understanding of embodied AI by making those systems feel more grounded and capable of nuanced spatial awareness rather than just reacting to isolated visual features <ref:2608.27407#pg1>.

Tom: That’s a great way to put it, Lalam, grounding the systems. So we’ve covered the overview of MILO and what the authors claim regarding their performance metrics across InterCap, HODome, and IMHD <ref:2608.27407#pg2>. Now we move on to what this actually means for the future in our next part of the discussion.

Jane: That’s right, Tom. We’re looking at how these findings translate into tangible progress for robotics and embodied AI in the coming years, and that's what we explore next.

Conclusion: Tom: So, to wrap up this paper, we've seen how they use Large Reconstruction Models to predict human and object interactions in three dee space from just one image. Jane, what do you think about the title itself?

Jane: I think it tells us exactly what the core achievement is—they are reconstructing those complex physical relationships in three dimensions. It’s a clear statement of purpose for the entire study.

Lu: From a theoretical standpoint, that framing shifts our focus away from trying to solve depth problems directly and toward interpreting the rich geometric structure an LRM provides. That's where the real creative potential lies for future research in scene understanding.

Meng: I’m interested in what this means practically; how does this translate when we move these concepts from a lab setting into something that actually runs reliably on a robot?

Lalam: For me, the implication is huge for culture; if AI can build such nuanced spatial understandings, it could fundamentally improve how we perceive and interact with our physical world in virtual environments.

Tom: That's an interesting transition from the technical details to the real-world impact. So, what are the main authors of this work that we should know?

Jane: The paper features a team of researchers who have been working extensively on large-scale reconstruction and three dee modeling for quite some time. Their collective expertise really shines through in how they've structured this approach.

Lu: I remember reading their earlier work on generative models, and it’s clear that the foundation they built there is what makes this specific application possible right now.

Meng: I’m more focused on the practical side—the robustness of these methods when things get messy in a real-world scenario where you can't guarantee perfect input quality.

Lalam: My perspective is that their ability to handle weaker information without needing explicit contact supervision opens up pathways for building more intuitive and physically grounded AI systems.

Tom: That’s the big picture—moving toward systems that understand the *how* of physical interaction rather than just *what* is there. We're heading into a look at how this impacts robotics next.

Agniv Chatterjee, Georgios Pavlakos

University of Texas at Austin

cs.CV

Submitted: 2026-08-27

Updated: 2026-08-27

Comments: Accepted at ECCV 2026. Project Page: https://ac5113.github.io/MILO

Journal ref: Computer Vision - ECCV 2026, Lecture Notes in Computer Science, vol 17037, pp. 37-57, Springer, Cham (2026)

DOI: 10.1007/978-3-032-37467-7_3

Project page: https://ac5113.github.io/MILO

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 86/100

The gist: Estimation of Human-Object Interactions in 3D (3D HOI) is a fundamental problem in 3D computer vision with applications in AR/VR, robotics, and embodied AI.

Key concepts

Large Reconstruction Models (LRMs)
These models are used as a starting point to generate a detailed 3D mesh from an input image. They provide a 'holistic' geometric structure that captures the overall arrangement and proximity of humans and objects in the scene, acting as a powerful scaffold for reconstruction.
SMPL-H Model
This is a parametric body model used to represent the human shape. The method fits this model to the 3D keypoints derived from the LRM mesh. It allows researchers to capture both the specific pose and shape of the human while ensuring anatomical plausibility through regularization.
3D Keypoint Estimation
This process involves rendering a 60-view LRM mesh, then using prediction models (ViTPose and HaMeR) to estimate 2D keypoints. These are then robustly triangulated into 3D points. This sparse, confidence-weighted set is essential for accurately fitting the human body model.
One-way Robust Chamfer Loss
This loss function is used during the SMPL-H fitting stage to improve reconstruction quality, especially in areas that are occluded. It compares the observed 3D points from the LRM reconstruction with the predicted vertices of the human mesh, helping to tighten the fit even when parts are hidden.

Terminology

Summary

Estimation of Human-Object Interactions in 3D (3D HOI) is a fundamental problem in 3D computer vision with applications in AR/VR, robotics, and embodied AI.

How it works

MILO is a framework that leverages the visual capabilities of Large Reconstruction Models (LRMs) to recover detailed 3D human-object interactions from a single image. The key observation is that LRMs provide a powerful geometric scaffold that preserves relative human-object arrangement and proximity cues, which simplifies the reconstruction procedure by reframing it as interpreting the LRM mesh.

The process involves several key steps:

  1. Applying an off-the-shelf LRM, such as Hunyuan3D-2.0, to the input RGB image to obtain a holistic human-object mesh. This output includes an alpha channel representing combined segmentation maps of the human and object.

  2. Interpreting this mesh by segmenting it into human and object components.

  3. Fitting a parametric body model, specifically SMPL-H, to the human part using predicted 3D keypoints obtained by rendering the LRM mesh from multiple virtual views and running ViTPose and HaMeR.

  4. Optionally aligning an object template to the object part if one is available, establishing semantic correspondences between it and the LRM mesh.

Key Technical Components

The method utilizes several technical components to achieve accurate reconstruction:

)&3D Keypoint Estimation:

** The LRM mesh is rendered from a set of virtual views (60 viewpoints). ViTPose predicts 2D body keypoints, and HaMeR predicts 2D hand keypoints for each view. These are robustly triangulated to obtain 3D keypoints on the LRM reconstruction. A reprojection-error threshold is used to select the hypothesis with the largest consensus set. The final set of body and hand 3D keypoints, denoted as X, is in a dimension of R67x3. This sparse, confidence-weighted set is then used to fit SMPL-H to the LRM mesh. A one-way robust Chamfer loss (Eq. 2) is employed to improve fitting in occluded regions by comparing observed 3D points with visible predicted vertices. The final fitting objective combines root loss, pose fitting losses, and the 3D consistency loss (Eq. 3). This optimization runs for 60 iterations to yield a human mesh tightly aligned with the LRM reconstruction while remaining anatomically plausible.**

)&Human Optimization:

** The SMPL-H parameters are refined in two stages: (i) root fitting, optimizing only global orientation and translation by minimizing the objective L rf, which aligns SMPL-H joints to the estimated 3D keypoints. (ii) pose fitting, which optimizes the full parameter set—shape β, body pose Θ, and hand pose Θh—while regularizing them using priors like VPoser [45], MANO [51], and an L2 shape prior. To improve robustness in occluded regions, an anchoring loss based on the HMR2.0 initialization is added to penalize deviations from the initialized latent body pose (Eq. 1). Furthermore, a visible vertex 3D consistency loss (Eq. 2) is introduced to improve fitting in self-occlusion and object-occlusion cases by comparing observed LRM points with visible predicted vertices.**

Evaluation and Contributions

MILO demonstrates strong performance across multiple benchmarks, including InterCap [20], HODome [78], and IMHD [80]. A key contribution is achieving state-of-the-art performance across multiple benchmarks, all while using weaker information than previous work, i.e., without relying on contact information.

The method's robustness is tested through several ablation experiments:


  1. Point Cloud Segmentation Ablation (Table S.3): The segmentation module is shown to be critical; dropping either visibility coverage or mask quality leads to worse PA-CD, confirming that all elements of the segmentation design are important for obtaining clean object point clouds from the combined LRM mesh and, consequently, high-quality HOI reconstructions.

  2. Template Alignment Ablation (Table S.2): Using ICP on top of the initial weighted Kabsch fit consistently improves alignment across all metrics, indicating that the geometry-aware correspondences provide a reliable coarse similarity transform.

  3. Fitting Stage Ablation (Table 7): Removing the pose fitting or both fitting stages consistently degrades performance, confirming that both stages are necessary for achieving the best overall reconstruction quality.

Conclusion

MILO successfully redefines HOI reconstruction by using LRMs as a geometric scaffold, capturing relative arrangement and proximity cues without requiring explicit contact supervision.

Improvements for AI systems

Here are the specific improvements that can be made to AI systems by implementing the MILO framework, and what those improved systems will be capable of:


  1. The ability to perform high-fidelity, single-image 3D Human-Object Interaction (HOI) reconstruction.

  2. The system can recover a detailed 3D human mesh (using SMPL-H) and a corresponding object geometry from a single RGB image, even in complex in-the-wild settings, without requiring pre-existing ground truth contact annotations or object templates.

  3. The system will accurately model the spatial arrangement and proximity cues between humans and objects by interpreting the geometric scaffold provided by Large Reconstruction Models (LRMs).

  4. The improved system can robustly segment a reconstructed mesh into distinct human and object components using multi-view segmentation, even in regions with ambiguous depth or occlusion.

  5. The system can perform template alignment to precisely position known object models relative to the reconstructed scene geometry, leveraging semantic correspondences between the LRM output and the template mesh.

  6. The system will achieve state-of-the-art reconstruction accuracy across multiple benchmarks (InterCap, HODome, IMHD) when compared against existing methods that rely on contact constraints or explicit object retrieval.

This improved AI system can be used for:

  1. Immersive AR/VR applications where users need to understand the 3D spatial relationship between a person and an object in a single captured frame (e.g., virtual furniture placement, realistic interaction simulation).

  2. Robotics and embodied AI tasks requiring real-time scene understanding, such as grasping or manipulation, by having an accurate 3D model of the human's pose and the surrounding environment immediately available from a camera feed.

  3. Enhanced scene analysis for autonomous systems to infer physical affordances (e.g., understanding how a person is supporting or interacting with an object) based on recovered geometry rather than relying solely on explicit contact sensors or prior knowledge of object shapes.

  4. Creating more physically plausible and coherent 3D scenes from sparse visual input, especially in complex indoor environments where traditional methods struggle with depth ambiguities and occlusions.

Sources

Related papers