Query-Conditioned Articulation Estimation from a Single Image

summary

Video file (mp4)

The gist

QueryArt is a model designed to estimate kinematic parameters of articulated objects from only a single RGB image and a 2D query point, which enables robots to infer object structure before physical

In short

QueryArt estimates kinematic parameters of articulated objects from a single RGB image and a 2D query point. It predicts joint type, motion axis direction, and offset for revolute joints. The model decouples interaction location from movement by predicting parameters in terms of the query point's depth, allowing metric recovery using only one depth measurement.

Key concepts

Query Art
A model that estimates how an object moves (its kinematics) just by looking at one picture and picking a spot on it. It helps robots understand object structure before touching it physically.
Decoupling Interaction and Motion
The core idea is separating *where* the robot interacts from *how* the selected point moves. The model predicts movement parameters relative to the query point's depth, making the target shape consistent regardless of uniform scaling.
Articulation Parameters
These are three main things the model predicts: 1) The joint type (static, revolute, or prismatic), 2) The direction of motion (a unit vector), and 3) For revolute joints, the offset of that axis from the query point.
Invariance to Scaling
The formulation is designed so that predictions are in units related to the query point's depth. This means the model is robust; it can correctly identify joint structure even if the object appears larger or smaller in a single image.

Terminology used across episodes

This episode discusses

The paper

Query-Conditioned Articulation Estimation from a Single Image · Read on arXiv

Socially Intelligent Robotics Lab, Institute for Artificial Intelligence, University of Stuttgart

Enabling robots to estimate the kinematic parameters of articulated objects unlocks a wide range of capabilities for interaction and manipulation. The estimation has to happen from the information the robot currently observes, often just a single RGB image of an object it has never seen before. Current single-image approaches couple articulation part segmentation with articulation estimation, making their predictions vulnerable to missed detections and incorrect part associations, and they regress metric 3D geometry that a single view fixes only up to scale. We present QueryArt, a model that estimates articulation parameters from a single RGB image, a 2D query point, and camera intrinsics. QueryArt is trained to estimate the 3D articulation geometry relative to the queried point and in units of its depth, which keeps its target identifiable from the image alone. A single depth measurement at the query point then supplies the scale and recovers the metric parameters. We train QueryArt on a curated mixture of synthetic and real-world articulation datasets. We evaluate QueryArt on several benchmarks and compare it against recent baselines. QueryArt outperforms recent baselines on most articulation metrics, including on out-of-distribution data. To demonstrate the model's capabilities in real-world settings, we evaluate QueryArt on a mobile manipulator across 57 manipulation trials spanning 16 object parts and five viewpoint classes, achieving a 70.2% success rate. We provide code and videos at: https://abwerby.github.io/queryart/

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "Query-Conditioned Articulation Estimation from a Single Image".

Dev: QueryArt is a model designed to estimate kinematic parameters of articulated objects from only a single RGB image and a 2D query point,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: We're starting by looking at the title, "Query-Conditioned Articulation Estimation from a Single Image," Abdelrhman Werby and Fabio Scaparro developed this work, which suggests they’ve managed to get the robot to figure out how an object is put together just from one picture and a spot on it.

Dev: It sounds like the core idea is decoupling where you interact from how the point moves, which simplifies things because it doesn't need complex segmentation beforehand, Rosa.

Taro: I see that decoupling as a way to isolate the kinematic prediction errors from any potential localization errors that might happen when trying to pinpoint an object in a cluttered scene.

Rosa: Right; so instead of having one system fail if the part segmentation is off, this framework treats articulation estimation as a regression of normalized geometry relative to the query point, which keeps the target identifiable from just the image.

Dev: That normalization step is key because it means we don't have to guess how big or small the object is just by looking at its appearance in one photo.

The paper's summary: Rosa: The authors summarize QueryArt as a model that takes an RGB image, a 2D query point, and camera intrinsics to estimate three things: the joint type, the three dee motion axis direction, and for revolute joints, the offset of that axis relative to where we queried.

Dev: I'm focusing on how they achieved this by using a frozen DINOv3 ViT-B/sixteen encoder and fusing features across three scales—S, S/two and S/four—which sounds like they are getting both broad context and fine detail simultaneously.

Taro: The mention of the query embedding being built from a Fourier encoding of the pixel location plus a hand-crafted calibration vector is interesting because it shows they are combining visual data with explicit camera knowledge to pinpoint exactly where we're looking.

Rosa: That combination allows the model to construct a Query embedding that feeds into learned tokens, TYPE and LINE, which then use deformable cross-attention across those feature pyramid scales to make their predictions.

Dev: So the mechanism is essentially using a powerful visual foundation model for context and then feeding that context into lightweight heads designed specifically for these articulation parameters.

The paper's improvements: Rosa: The authors highlight several key improvements, including using an outer-product embedding for the revolute axis direction vector to ensure it's invariant to reversing that direction, which is a neat trick for consistency.

Dev: And they use a specific target definition for the offset prediction, = I three - rev rev v r, where that projection guarantees the offset vector is orthogonal to the axis direction and points toward the nearest point on that line.

Taro: That geometric constraint, ensuring orthogonality between the predicted axis and the offset vector, makes sense for physical consistency; it prevents nonsensical predictions where an axis direction would be perpendicular to its own offset.

Rosa: They also emphasize that they train the model with a composite loss function involving type classification cross-entropy and unsigned cosine loss for the axis direction, along with Smooth L1 loss for the depth-normalized offset target r*.

Dev: The crucial part of their training is that the target r* is specifically designed to be invariant to reversing a star direction or jointly scaling the metric scene and query depth, which means the model learns a representation that holds up better under scale changes.

Conclusion: Rosa: To wrap up, QueryArt shows a way to estimate full three dee kinematic parameters from just one image and a query point by training on both synthetic and real-world data, achieving success rates around seventy percent on mobile manipulator tasks.

Dev: The real implication here is that if we can reliably get these predictions in real-time with low latency, we could have robots autonomously opening furniture or interacting with complex objects in cluttered environments without needing extensive prior exploration.

Taro: I think the ability to infer structure before physical interaction, even under uncertain conditions, is what's most significant for autonomy; it moves us closer to systems that can react intelligently when things don't go exactly as planned.

Rosa: Exactly; we’ve seen strong performance on out-of-distribution data too, suggesting good generalization capabilities for novel object instances.

Dev: From an engineering standpoint, the challenge remains ensuring this inference loop is fast enough to be practical for continuous interaction, especially considering the complexity of the Transformer architecture they use.

Taro: We’ll have to see how researchers address those latency concerns in future work; that’s where we need more rigor if we want this technology to move out of the lab and into truly autonomous operation.

More episodes

← Home