Query-Conditioned Articulation Estimation from a Single Image
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.
Rosa: Today's paper: "Query-Conditioned Articulation Estimation from a Single Image".
Dev: QueryArt is a model designed to estimate kinematic parameters of articulated objects from only a single RGB image and a 2D query point,
Rosa: First, who's behind it and why it matters.
Title and authors: Rosa: We're starting by looking at the title, "Query-Conditioned Articulation Estimation from a Single Image," Abdelrhman Werby and Fabio Scaparro developed this work, which suggests they’ve managed to get the robot to figure out how an object is put together just from one picture and a spot on it.
Dev: It sounds like the core idea is decoupling where you interact from how the point moves, which simplifies things because it doesn't need complex segmentation beforehand, Rosa.
Taro: I see that decoupling as a way to isolate the kinematic prediction errors from any potential localization errors that might happen when trying to pinpoint an object in a cluttered scene.
Rosa: Right; so instead of having one system fail if the part segmentation is off, this framework treats articulation estimation as a regression of normalized geometry relative to the query point, which keeps the target identifiable from just the image.
Dev: That normalization step is key because it means we don't have to guess how big or small the object is just by looking at its appearance in one photo.
The paper's summary: Rosa: The authors summarize QueryArt as a model that takes an RGB image, a 2D query point, and camera intrinsics to estimate three things: the joint type, the three dee motion axis direction, and for revolute joints, the offset of that axis relative to where we queried.
Dev: I'm focusing on how they achieved this by using a frozen DINOv3 ViT-B/sixteen encoder and fusing features across three scales—S, S/two and S/four—which sounds like they are getting both broad context and fine detail simultaneously.
Taro: The mention of the query embedding being built from a Fourier encoding of the pixel location plus a hand-crafted calibration vector is interesting because it shows they are combining visual data with explicit camera knowledge to pinpoint exactly where we're looking.
Rosa: That combination allows the model to construct a Query embedding that feeds into learned tokens, TYPE and LINE, which then use deformable cross-attention across those feature pyramid scales to make their predictions.
Dev: So the mechanism is essentially using a powerful visual foundation model for context and then feeding that context into lightweight heads designed specifically for these articulation parameters.
The paper's improvements: Rosa: The authors highlight several key improvements, including using an outer-product embedding for the revolute axis direction vector to ensure it's invariant to reversing that direction, which is a neat trick for consistency.
Dev: And they use a specific target definition for the offset prediction, = I three - rev rev v r, where that projection guarantees the offset vector is orthogonal to the axis direction and points toward the nearest point on that line.
Taro: That geometric constraint, ensuring orthogonality between the predicted axis and the offset vector, makes sense for physical consistency; it prevents nonsensical predictions where an axis direction would be perpendicular to its own offset.
Rosa: They also emphasize that they train the model with a composite loss function involving type classification cross-entropy and unsigned cosine loss for the axis direction, along with Smooth L1 loss for the depth-normalized offset target r*.
Dev: The crucial part of their training is that the target r* is specifically designed to be invariant to reversing a star direction or jointly scaling the metric scene and query depth, which means the model learns a representation that holds up better under scale changes.
Conclusion: Rosa: To wrap up, QueryArt shows a way to estimate full three dee kinematic parameters from just one image and a query point by training on both synthetic and real-world data, achieving success rates around seventy percent on mobile manipulator tasks.
Dev: The real implication here is that if we can reliably get these predictions in real-time with low latency, we could have robots autonomously opening furniture or interacting with complex objects in cluttered environments without needing extensive prior exploration.
Taro: I think the ability to infer structure before physical interaction, even under uncertain conditions, is what's most significant for autonomy; it moves us closer to systems that can react intelligently when things don't go exactly as planned.
Rosa: Exactly; we’ve seen strong performance on out-of-distribution data too, suggesting good generalization capabilities for novel object instances.
Dev: From an engineering standpoint, the challenge remains ensuring this inference loop is fast enough to be practical for continuous interaction, especially considering the complexity of the Transformer architecture they use.
Taro: We’ll have to see how researchers address those latency concerns in future work; that’s where we need more rigor if we want this technology to move out of the lab and into truly autonomous operation.
Socially Intelligent Robotics Lab, Institute for Artificial Intelligence, University of Stuttgart
cs.RO
Submitted: 2026-10-01
Updated: 2026-10-02
Comments: Code and video are available at https://abwerby.github.io/queryart/
Project page: https://abwerby.github.io/queryart
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 87/100
The gist: QueryArt is a model designed to estimate kinematic parameters of articulated objects from only a single RGB image and a 2D query point, which enables robots to infer object structure before physical
Key concepts
- Query Art
- A model that estimates how an object moves (its kinematics) just by looking at one picture and picking a spot on it. It helps robots understand object structure before touching it physically.
- Decoupling Interaction and Motion
- The core idea is separating *where* the robot interacts from *how* the selected point moves. The model predicts movement parameters relative to the query point's depth, making the target shape consistent regardless of uniform scaling.
- Articulation Parameters
- These are three main things the model predicts: 1) The joint type (static, revolute, or prismatic), 2) The direction of motion (a unit vector), and 3) For revolute joints, the offset of that axis from the query point.
- Invariance to Scaling
- The formulation is designed so that predictions are in units related to the query point's depth. This means the model is robust; it can correctly identify joint structure even if the object appears larger or smaller in a single image.
Terminology
Summary
QueryArt is a model designed to estimate kinematic parameters of articulated objects from only a single RGB image and a 2D query point, which enables robots to infer object structure before physical interaction. The core contribution is decoupling where to interact from how the selected point moves by predicting articulation parameters in units of the query point's depth, thus making the target invariant to uniform scaling and allowing for metric recovery using a single depth measurement.
How it works
The model estimates three primary articulation parameters: the joint type (static, revolute, or prismatic), the 3D motion axis direction, and for revolute joints, the 3D offset of that axis relative to the query point. The formulation is designed to predict a joint type from a single calibrated RGB image and a pixel coordinate query. For moving parts, it predicts a unit direction vector (axis) and an offset vector.
Model Architecture
QueryArt employs an architecture that combines high-resolution frozen visual foundation features with lightweight prediction heads. Specifically, it utilizes a frozen DINOv3 ViT-B/16 image encoder to produce 768-channel patch features. These features are fused using a three-scale feature pyramid, downsampling the fused features to resolutions S, S/2, and S/4. A query embedding is constructed by combining three signals: a Fourier positional encoding of the normalized pixel query, a bilinear sample from the final DINOv3 block feature at the query location (a 768-channel descriptor), and a hand-crafted calibration vector derived from camera intrinsics. This combined signal forms the Query embedding, which is then added to two randomly initialized learned task tokens: a TYPE token and a LINE token. These tokens pass through two decoder layers where they undergo deformable cross-attention to the three-scale image pyramid, followed by self-attention between the TYPE and LINE tokens, and finally a token-wise feedforward network.
Articulation Heads
The model utilizes four distinct output heads. A single MLP classifier reads the TYPE token to predict the joint type, denoted as p(j I, q, K). Two independent MLPs read the LINE token to produce raw vectors for revolute and prismatic directions, vˆrev and vˆpris, respectively. These raw vectors are normalized to yield aˆrev and aˆpris. For revolute joints specifically, the fourth head predicts the axis offset relative to the query point. To ensure this prediction is invariant to axis sign, an outer-product embedding ψ(a) is used for the axis direction vector aˆrev. A separate MLP projects both the LINE token and ψ(aˆrev) into 64 dimensions, which are summed and passed through a linear layer to produce a raw vector vr. The final offset prediction is then enforced via the equation: ˆr = I3 − aˆrevaˆ⊤ rev vr, where I3 is the identity matrix. This projection guarantees that the resulting offset vector is orthogonal to the predicted axis direction and points from the unprojected query point to the nearest point on the predicted axis line.
Loss Functions
The model is trained using a composite loss function L = Ltype + 2Lrev axis + 2Lpris axis + 2Loffset. The joint type prediction is supervised using a three-class cross-entropy loss (Ltype). For each moving type, an unsigned cosine loss supervises the axis direction vector (Laxis). For the revolute joint, the normalized offset target r∗ is defined as r∗ = I3 − a∗a∗⊤ o∗ zq − p˜. This target is designed to be invariant to reversing a star direction, shifting o star along its axis, and jointly scaling the metric scene and query depth. The predicted offset ˆr is supervised using Smooth L1 loss applied to the Euclidean prediction error: Loffset = SmoothL1β(∥ˆr − r∗∥2), with β = 0.02. Eligibility for geometric supervision depends on ground-truth labels, not the predicted type, ensuring that incorrect type predictions do not suppress geometric learning.
Evaluation and Results
QueryArt was trained on a curated mixture of synthetic and real-world articulation datasets, including Articulate3D [9], SceneFun3D [38], MultiScan [37], PartNet-Mobility [39], and HSSD [40]. Evaluation metrics include axis error in degrees, type recall in %, depth error (dh), RMSE in meters, and the overall success rate (S). On held-out splits of four training datasets and two excluded out-of-distribution datasets (Arti4D [10] and HOI! [41]), QueryArt outperformed recent single-image baselines on most articulation metrics. In real-robot experiments on a mobile manipulator, QueryArt achieved a 70.
Improvements for AI systems
Here are the specific improvements for AI systems based on the QueryArt paper, along with what those improved systems can achieve:
-
The core improvement is a shift from coupled
detection and segmentation
methods to a decoupledquery-conditioned regression
formulation for articulation estimation. This decouples localization errors from kinematic prediction errors. -
This system can estimate the full 3D kinematic parameters (joint type, motion axis direction, and query-relative axis location) of an articulated object solely from a single RGB image and a 2D query point, without requiring prior scene exploration or multi-view observations.
-
The system utilizes a normalized coordinate space representation for the articulation geometry relative to the query point's depth.
-
This normalization makes the target invariant to uniform scaling of the scene, meaning the model doesn't need to infer absolute metric size from appearance alone (like traditional methods that regress raw 3D geometry).
-
The system explicitly predicts a direction vector for both revolute and prismatic joints, along with an offset vector.
-
For revolute joints, the predicted offset is expressed in units of the query point's depth, allowing for deterministic lifting of the prediction to metric 3D space once the actual depth is measured at that query point (a single measurement suffices).
-
The architecture leverages a frozen foundation model encoder (DINOv3 ViT-B/16) combined with specialized, lightweight articulation-specific heads and a feature pyramid structure.
-
This combination allows the system to effectively combine high-resolution global semantic context (from the foundation model) with fine-grained local detail necessary for precise joint localization and classification.
-
The training objective is multi-task: it supervises the joint type classification, the axis direction using an unsigned cosine loss, and crucially, uses a Smooth L1 loss on a depth-normalized offset target that is invariant to axis reversal and global scaling.
-
This rigorous geometric supervision ensures that the predicted kinematic parameters are physically consistent (e.g., ensuring the predicted offset vector is orthogonal to the predicted axis direction), leading to highly accurate and reliable predictions, even on out-of-distribution data (like unseen furniture).
-
The resulting system can drive real-world manipulation tasks by generating analytic motion trajectories directly from its predictions.
-
Specifically, a robot equipped with this system can autonomously open previously unseen furniture (drawers, cabinets) or interact with objects in cluttered environments by first querying the model at a grasp point to determine the required joint parameters and motion path before executing physical interaction.
-
The system's performance is robust across various data modalities (synthetic and real-world) and viewpoint classes, demonstrating strong generalization capabilities to out-of-distribution object instances without requiring further retraining or fine-tuning.
Abstract
Enabling robots to estimate the kinematic parameters of articulated objects unlocks a wide range of capabilities for interaction and manipulation. The estimation has to happen from the information the robot currently observes, often just a single RGB image of an object it has never seen before. Current single-image approaches couple articulation part segmentation with articulation estimation, making their predictions vulnerable to missed detections and incorrect part associations, and they regress metric 3D geometry that a single view fixes only up to scale. We present QueryArt, a model that estimates articulation parameters from a single RGB image, a 2D query point, and camera intrinsics. QueryArt is trained to estimate the 3D articulation geometry relative to the queried point and in units of its depth, which keeps its target identifiable from the image alone. A single depth measurement at the query point then supplies the scale and recovers the metric parameters. We train QueryArt on a curated mixture of synthetic and real-world articulation datasets. We evaluate QueryArt on several benchmarks and compare it against recent baselines. QueryArt outperforms recent baselines on most articulation metrics, including on out-of-distribution data. To demonstrate the model's capabilities in real-world settings, we evaluate QueryArt on a mobile manipulator across 57 manipulation trials spanning 16 object parts and five viewpoint classes, achieving a 70.2% success rate. We provide code and videos at: https://abwerby.github.io/queryart/
Sources
- MolmoMotion: Forecasting Point Trajectories in 3D with Language Instruction
- Articulated 3D Scene Graphs for Open-World Mobile Manipulation
- DINOv3
- ScrewSplat: An End-to-End Method for Articulated Object Recognition
- Hoi! - A Multimodal Dataset for Force-Grounded, Cross-View Articulated Manipulation
Related papers
- FMT x: An Efficient and Asymptotically Optimal Extension of the Fast Marching Tree for Dynamic Replanning
- MPCFormer: A physics-informed data-driven approach for explainable socially-aware autonomous driving
- RoboLab: A High-Fidelity Simulation Benchmark for Analysis of Task Generalist Policies
- HRDexDB: A 4D Dexterous Grasping Dataset Across Human and Multiple Robot Embodiments
- APT: Action Expert Pretraining Improves Instruction Generalization of Vision-Language-Action Policies
- Fine-tuning is Not Enough: A Parallel Framework for Collaborative Imitation and Reinforcement Learning in End-to-end Autonomous Driving