Robotic Ultrasound Makes CBCT Alive
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Robotic Ultrasound Makes CBCT Alive".
Tom: Intraoperative Cone Beam Computed Tomography (CBCT) provides essential 3D anatomical context, but its static nature fails to monitor soft-tissue deformations caused by respiration or surgical manipulation, leading to navigation discrepancies.
Jane: First, who's behind it and why it matters.
Paper summary: Tom: Okay, we’ve got a good overview of what the paper is aiming for; we learned that "Robotic Ultrasound Makes CBCT Alive" tackles the problem of static imaging failing to track soft-tissue movement during surgery. Now, Jane, can you elaborate on how they actually get from that initial problem statement to their proposed solution?
Jane: Certainly. The paper sets up the premise that while CBCT is great for planning because it provides three dee context, its fixed nature causes navigation discrepancies when respiration or manipulation deforms the tissue <ref:2603.10220#pg0>. Their solution involves a deformation-aware CBCT updating framework that uses robotic ultrasound as a dynamic proxy to infer tissue motion and then apply those inferred motions to update the static CBCT slices in real time.
Lu: What’s really compelling is the initial step they take—starting with calibration-initialized alignment using LC2-based rigid refinement, which establishes this accurate multimodal correspondence before moving on to the more complex non-rigid deformation estimation.
Meng: I see that in Figure one(a) they show this rigid calibration between the robotic ultrasound and CBCT <ref:2603.10220#pg0>. How precise is that initial relationship, and what happens when that initial alignment isn't perfect?
Jane: They state that this initial calibration achieves a mean alignment error of approximately one to two millimeters, but they acknowledge that residual misalignment still persists without image-based refinement, which can be clinically significant in procedures requiring high precision.
Lalam: That residual error is where the real power of the subsequent steps comes into play; it’s not just about getting close initially, it's about correcting those tricky small errors using more sophisticated methods later on.
Tom: Right, and that leads us to the core of their dynamic approach. Lu, can you walk us through what they introduce to capture that non-rigid motion? What is USCorUNet doing here?
Lu: They introduce USCorUNet for this purpose; it’s a lightweight network trained with optical flow-guided supervision to learn "deformation-aware correlation representations," which allows it to produce accurate, real-time dense deformation field estimation from ultrasound streams.
Meng: A lightweight network that outputs a dense deformation field—that sounds computationally intense for real-time use. How are they keeping the inference fast enough for intraoperative use?
Jane: The architecture consists of a ResUNet-style context encoder-decoder combined with a shared-weight correlation encoder that builds local correlation volumes, and these are fused with context features before being decoded into dense fields at one eighth resolution.
Lalam: If this network can successfully estimate those dense fields in real time, it opens up possibilities for AI to assist in predicting tissue deformation before the manipulation even happens, which is a huge leap for proactive surgical guidance.
Tom: So, if USCorUNet estimates the motion, what’s the next critical step? How does that inferred motion actually get applied back onto the static CBCT slice?
Jane: They then use this inferred deformation field to update the CBCT reference, which produces deformation-consistent visualizations without needing repeated radiation exposure. Specifically for probe-induced motion, they correct non-uniform convex compression using a Gaussian profile where drobot denotes displacement magnitude.
Lu: That specific correction mechanism, D geoy = Drawy - P(x), which corrects the vertical deformation component with that Gaussian profile, shows a very targeted approach to handling probe interaction.
Meng: So they're essentially taking the dynamic information from ultrasound and mathematically translating it into geometric corrections for the static image. That translation step seems like where most of the engineering complexity lies.
Lalam: This process of transferring deformation fields is incredibly powerful because it allows us to visualize what the tissue *would* look like under different conditions, enhancing our understanding of surgical maneuvers through augmented reality tools.
Tom: It’s a very clever way to create dynamic visualization from static data without having to re-scan the patient repeatedly. This sets up the next big part of our discussion: what does this all mean when we look at the authors and their overall conclusions?
Conclusion: Tom: We’ve covered a lot about how this framework works, from calibration to USCorUNet, and now we’re heading into the final thoughts. Tom and Jane will discuss the title of "Robotic Ultrasound Makes CBCT Alive" and what the implications of this work are for future medical applications.
Jane: They'll summarize how they achieved accurate multimodal correspondence through their combination of calibration, LC2 refinement, USCorUNet for deformation estimation, and ultrasound-informed CBCT slice updating to achieve dynamic refinement of static guidance during robotic ultrasound-assisted interventions.
Lu: I think the authors really highlight that they’ve established a robust pipeline that integrates these different components into one cohesive system for intraoperative imaging. It shows how you can build complex systems by layering simpler, well-established techniques effectively.
Meng: From a practical perspective, what is the main limitation the authors themselves point out regarding this approach? What does it stop working at?
Jane: They do mention that ultrasound captures only deformation-dependent observations without providing a consistent global reference for comprehensive anatomical context. This means while it gives dynamic information, it doesn't replace the full anatomical context of CBCT.
Lalam: That distinction is important because it sets realistic expectations; this method doesn't magically give us perfect, complete anatomy; it provides a way to dynamically correct the errors in our static data using real-time input.
Tom: So, to wrap up on the broader implications, Lu, what is your take on how this paper might influence the future of AI in surgical planning and execution?
Lu: I see this as establishing a new standard for accurate multimodal correspondence in surgical guidance where the system can dynamically adapt its static priors based on real-time biological feedback from ultrasound. This could significantly enhance the precision of robotic interventions guided by intraoperative imaging.
Meng: I think for me, it means we need to focus our engineering efforts on creating systems that can handle this kind of real-time data fusion efficiently, moving beyond just one modality to truly integrate them seamlessly in the operative environment.
Lalam: For culture, this advance in AI capabilities suggests that future medical AI will be much more sophisticated at modeling physical interaction, leading to better training data generation and safer robotic systems overall.
Tom: It sounds like a really solid piece of work that moves us closer to truly dynamic intraoperative imaging by connecting ultrasound and CBCT in a meaningful way. That’s all for this discussion today.
Technical University of Munich Center for Computer-Aided Medical Procedures and Augmented Reality Center for Machine Learning Munich Center for Machine Learning The University of Hong Kong
cs.CV, cs.AI, cs.RO
Submitted: 2026-03-10
Updated: 2026-03-10
Comments: 10 pages, 4 figures
DOI: 10.1007/978-3-032-38236-8_42
Code: https://github.com/anonymous-codebase/us-cbct-demo
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 82/100
The gist: Intraoperative Cone Beam Computed Tomography (CBCT) provides essential 3D anatomical context, but its static nature fails to monitor soft-tissue deformations caused by respiration or surgical
Key concepts
- Deformation-aware CBCT updating framework
- This is a system designed to make static 3D CT scans (CBCT) dynamic. It uses real-time ultrasound data to detect how tissues move due to breathing or surgical tools. By inferring this motion, the system updates the static CT slices instantly, ensuring the guidance remains accurate even when anatomy changes.
- USCorUNet
- This is a specific deep learning model used to estimate tissue deformation from ultrasound streams. It analyzes local correlation volumes within the ultrasound data and learns how these features change as tissues move. It is trained using multiple loss functions to ensure the estimated deformations are both accurate and physically plausible.
- Optical Flow-guided supervision
- This is a training technique used for USCorUNet. Instead of just looking at static images, the network is guided by optical flow—a mathematical representation of how pixels move between two slightly different frames. This helps the AI learn to predict dense deformation fields accurately in real time.
- Deformation transfer
- This is the process where the motion information learned by USCorUNet is applied back to the CBCT data. It corrects any warping or compression in the CT slices based on ultrasound observations. This results in new visualizations that accurately reflect the current, deformed state of human tissue.
Terminology
Summary
Intraoperative Cone Beam Computed Tomography (CBCT) provides essential 3D anatomical context, but its static nature fails to monitor soft-tissue deformations caused by respiration or surgical manipulation, leading to navigation discrepancies. This paper proposes a deformation-aware CBCT updating framework that leverages robotic ultrasound as a dynamic proxy to infer tissue motion and update static CBCT slices in real time, enabling dynamic refinement of static guidance during robotic interventions.
The gist
We propose a deformation-aware CBCT updating framework that leverages robotic ultrasound as a dynamic proxy to infer tissue motion and update static CBCT slices in real time.
Framework Components and Workflow
The system is structured around four main modules:
-
Calibration-based rigid initialization: This establishes the initial spatial relationship between the CBCT and ultrasound data, computed as a rigid transformation where
C TU = C TR(UTR)−1
. This is refined by LC2 to correct residual errors, modeling the relationship via a local linear approximation: f(xi) = αpi + βgi + γ. -
Non-rigid deformation estimation: This is performed by USCorUNet, a lightweight network trained with
optical flow-guided supervision
to learndeformation-aware correlation representations,
enablingaccurate, real-time dense deformation field estimation from ultrasound streams.
-
Deformation transfer for CBCT slice updating: The inferred deformation field is applied to update the CBCT reference, producing
deformationconsistent visualizations without repeated radiation exposure.
USCorUNet Architecture and Training
The USCorUNet architecture consists of a preprocessing module feeding into a ResUNet-style context encoder-decoder combined with a shared-weight correlation encoder. It builds local correlation volumes,
such as CV01(x, ∆) = √1/C ⟨f0(x), f1(x + ∆)⟩, which are fused with context features and decoded into dense fields at 1/8 resolution.
The network is trained using a bidirectional objective combining three losses:
Lflow:
Optical flow distillation (an l1 loss to the optical flows).
:Lphoto:
Confidence-weighted photometric consistency (a Charbonnier penalty on post-warp intensity residuals using random-walk confidence maps).
:Lreg:
Regularization, which combines edge-aware smoothness
and a Jacobian-based folding penalty.
Deformation Field Application
The estimated deformation field is applied to update the CBCT slice in real time. Specifically for probe-induced motion, non-uniform convex compression is corrected using a Gaussian profile P(x) = drobot · exp(−(x − cx)2/2σprobe2, where drobot denotes displacement magnitude. This corrects the vertical deformation component: Dgeoy = Drawy − P(x). To align the local ultrasound field with the larger CBCT ROI, Euclidean Distance Transform (EDT)-based spatial weighting is used, and the corrected field is padded to Dpad and scaled by W(x, y) = exp(−D(x, y)/σsmooth), resulting in the final field Dfinal.
Validation and Results
The approach was validated through deformation estimation and ultrasound-guided CBCT updating experiments on four datasets: (A) in vivo forearm/upperarm ultrasound, (B) a pork-tissue gel phantom, (C) a chicken/pork gel phantom, and (D) an abdominal phantom. The results demonstrate real-time endto-end CBCT slice updating and physically plausible deformation estimation,
achieving a favorable accuracy–efficiency trade-off against RAFT-based and classical baselines.
Quantitative comparisons show that USCorUNet slightly outperforms RAFT while reducing runtime by 5× compared to LC2-FFD, which is significantly slower. The fine-tuned models further show improvements in FB consistency and physical plausibility
for both probe-induced and externally induced motion.
Conclusion
The proposed framework successfully integrates calibration, LC2 refinement, USCorUNet for deformation estimation, and ultrasound-informed CBCT slice updating to achieve dynamic refinement of static CBCT guidance during robotic ultrasound-assisted interventions. The method establishes accurate multimodal correspondence
and provides a robust pipeline for intraoperative imaging.
--- Page 10 ---
References
-
Bi, Y., Jiang, Z., Duelmer, F., Huang, D., Navab, N.: Machine learning in robotic ultrasound imaging: Challenges and perspectives. Annual Review of Control, Robotics, and Autonomous Systems 7 (2024)
-
Bruhn, A., Weickert, J., Schnörr, C.: Lucas/kanade meets horn/schunck: Combining local and global optic flow methods. International journal of computer vision 61(3), 211–231 (2005)
-
Curiale, A.H.
Improvements for AI systems
Here are the specific improvements and capabilities derived from this scientific paper, framed for enhancing AI systems:
The proposed framework, centered on deformation-aware Cone Beam Computed Tomography (CBCT) updating using robotic ultrasound as a dynamic proxy, offers several high-impact improvements for AI systems in medical image guidance and robotic interventions.
-
Predictive Dynamic Image Reconstruction:
-
Real-time Probe/Motion Correction:
-
Robust Multimodal Correspondence Learning:
-
Physically Plausible Deformation Modeling:
-
Predictive Dynamic Image Reconstruction
The system can perform continuous, real-time 3D volumetric image reconstruction that dynamically adjusts to intraoperative physical changes (e.g., respiration, probe pressure). It moves beyond static imaging by inferring and applying a deformation field to update a static CBCT slice into a deformation-consistent visualization.
- Real-time Probe/Motion Correction
The AI system can automatically correct for non-uniform convex compression induced by surgical tools (probes) or external patient motion. Specifically, it uses the estimated deformation field to derive corrective profiles (like the Gaussian profile P(x)) that precisely counteract these artifacts in real time, ensuring guidance remains accurate during dynamic procedures.
- Robust Multimodal Correspondence Learning
The system can establish highly accurate spatial relationships between modalities (ultrasound and CBCT) even when there is significant non-rigid deformation. This is achieved through a pipeline:
-
Rigid Calibration Initialization (using LC2 refinement for residual error correction).
-
Deformation Field Estimation using the USCorUNet network, which learns
deformation-aware correlation representations
from ultrasound streams, enabling dense, real-time estimation of tissue motion fields between consecutive frames. -
Physically Plausible Deformation Modeling
The AI system's deformation estimation is constrained by a multi-objective loss function (L = λflowLflow + λphotoLphoto + λregLreg). This ensures that the inferred deformations are not just mathematically plausible but also physically plausible
by incorporating:
-
Optical Flow Distillation (for motion fidelity).
-
Confidence-Weighted Photometric Consistency (to ensure anatomical faithfulness based on image appearance).
-
Regularization terms (edge-aware smoothness and Jacobian-based folding penalty) to enforce biomechanical constraints, preventing the generation of
biomechanically implausible deformation fields.
This improved AI system can be used to create an autonomous closed-loop surgical guidance platform where a static pre-operative CBCT is continuously refined by live robotic ultrasound data, providing a highly accurate, dynamically updated 3D anatomical context for interventions (e.g., needle insertion, ablation) with reduced radiation exposure and minimal navigation error.
Sources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models