Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth".
Jane: The paper was written by Sho MITARAI, Hikaru KAYO, Hisashi OZAKI, Yuichiro IMAI and Megumi NAKAO from Kyoto University and Rakuwakai Otowa Hospital.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Alright, listeners, we are back with another paper that has genuinely got me excited. Today we're looking at "Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth." Jane, this one is all about getting two very different three dee images of your mouth to line up perfectly.
Jane: And Tom, I have to say, the title is a mouthful, but the problem is so relatable if you've ever been to a dentist. You've got a CT scan that shows the bone inside, and you've got an intraoral scan that shows the surface of your teeth. They're both of the same mouth, but they don't naturally overlap because one is a solid block and the other is just a shell.
Tom: Exactly. And the team behind this is from Kyoto University and Rakuwakai Otowa Hospital in Japan. Sho Mitarai, Hikaru Kayo, Hisashi Ozaki, Yuichiro Imai, and Megumi Nakao. They're tackling this problem of merging the inside and the outside of the jaw.
Jane: What I love about this paper is that they admit something that a lot of research glosses over. When you scan a patient with a CT and then scan them with an intraoral scanner, you don't actually know the true transformation between the two. You don't know exactly how they should line up because they're captured at different times, in different positions.
Tom: Right, so how do you even test if your registration is good if you don't know the answer? That's the clever part. They built a "pseudo-IOS" — a fake intraoral scan — generated directly from the CT data. Since both the fake scan and the CT come from the same coordinate frame, they know the exact ground truth transformation.
Jane: It's like having the answer key to a test you're writing yourself. You can finally measure how good your algorithm is with real numbers instead of just eyeballing it. And that's what makes this paper so valuable for the field of digital dentistry.
Tom: So before we get into the nitty-gritty of how they actually align these clouds, I want to flag the bigger picture. This isn't just about making pretty three dee models. This is about surgical planning, about guiding implants, about reconstructing jaws after trauma. Getting this registration right has real clinical consequences.
Jane: Absolutely, Tom. And the fact that they've built this evaluation framework means that future researchers can actually compare methods fairly. That's a huge step forward. But let's not get ahead of ourselves — next we need to talk about what the paper actually found when they ran their experiments.
Tom: Good point. We've got the setup, now let's look at the results. Stick around.
Summary of the Paper: Tom: Welcome back. We're still on "Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth." Jane, we set the stage with the problem, but what did they actually do to solve it?
Jane: So they used a two-stage approach. First, they use something called GeDi, which is a learning-based three dee local descriptor. It looks at small patches of the point cloud and creates a signature for each point that's invariant to rotation and scale. This lets them find matching points between the CT and the pseudo-IOS without any good initial guess.
Tom: And that's the "global registration" part. It gets you roughly in the right neighborhood. But rough isn't good enough for surgery. So they follow it up with the classic Iterative Closest Point, or ICP, algorithm. That's the fine-tuning step that minimizes the distance between the two surfaces.
Jane: And the results were pretty striking. When they ran ICP alone from a perturbed starting position, it failed spectacularly. We're talking a mean absolute error of over twenty-five millimeters. That's not just off by a little — that's aligning the jaw to the wrong part of the head.
Tom: But with GeDi plus ICP, they got the mean absolute error down to between zero point five five and zero point six nine millimeters across all conditions. That's submillimeter accuracy. For context, that's about the width of a human hair or a thin slice of paper.
Jane: And it wasn't just accurate, it was consistent. The standard deviation was tiny, which means it wasn't just getting lucky on some cases and failing on others. It was reliably finding the right alignment.
Tom: They also tested with simulated metal artifacts. You know, when you have metal fillings or crowns, they create streaks in the CT scan that distort the image. They simulated having one to eight metal teeth, and while the artifacts did degrade performance for GeDi alone, the full GeDi+ICP pipeline stayed robust.
Jane: That's the key finding for me. The combination of a global descriptor-based approach with local refinement is what makes this work. Neither stage alone was sufficient, but together they nail it.
Tom: And they proved it statistically with a three-way repeated-measures ANOVA. The method effect was significant, the artifact effect was significant, but the absolute impact on their proposed method was small.
Jane: So the summary is: this is a robust, accurate, and reliable method for a clinically critical task. And they built the evaluation framework to prove it. Now, let's talk about what this means for actual practice.
Tom: Right, because a paper can have great numbers, but the real question is whether it changes how dentists and surgeons work. That's our next segment.
Improvements Suggested: Tom: We're back on the show, still discussing "Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth." Jane, we've covered the results, but what improvements does this paper suggest over the current state of the art?
Jane: The biggest improvement is the evaluation framework itself. Before this, if you wanted to test CT-to-IOS registration, you had to rely on manual alignment as your reference, or use surrogate metrics. That means your "ground truth" was itself an estimate, so you could never be sure if your algorithm was actually good or if your reference was just wrong.
Tom: So by generating the pseudo-IOS from the CT, they eliminate that uncertainty entirely. The transformation is known by construction. It's a controlled experiment instead of a guess.
Jane: Exactly. And that's not just a technical detail. It means that when they report a mean absolute error of zero point five five millimeters, that number is trustworthy. It's not contaminated by the error in the reference alignment.
Tom: The other improvement is methodological. They show that using a domain-generalizable descriptor like GeDi is the right choice for this multimodal problem. GeDi was trained on three deeMatch, which is a dataset of indoor scenes, not dental data. Yet it generalizes well to this completely different domain.
Jane: That's the beauty of it. You don't need to train a new model for every clinical application. You can take a pretrained descriptor and apply it to a new problem, as long as you combine it with a good refinement step like ICP.
Tom: And they also addressed the metal artifact problem head-on. Instead of just hoping the algorithm would be robust, they simulated artifacts in a controlled way and measured exactly how much they hurt. That's the kind of rigorous testing that clinical adoption requires.
Jane: Right. Because in the real world, patients have fillings, crowns, implants. If your registration method falls apart when there's a metal crown, it's not useful clinically. Their method held up, with the artifact-induced error increase being statistically detectable but practically small.
Tom: So the improvements here are threefold: a trustworthy evaluation framework, a validated coarse-to-fine pipeline, and a clear understanding of how metal artifacts affect performance. That's a solid contribution.
Jane: And it opens the door for future work. Now that we have this framework, researchers can test other descriptors, other refinement methods, other artifact simulation techniques. The field can move forward on solid ground.
Tom: Let's take a quick break, and when we come back, we'll dig into the actual first page of the paper and the details of their method.
First Page Discussion: Tom: Welcome back to the show. We're diving into the first page of "Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth." Jane, what stands out to you when you look at the opening?
Jane: The abstract really sets the stage. It frames the problem as a registration challenge where the two modalities share only a limited region in common. The CT shows the whole bone structure, but the intraoral scan only captures the visible surfaces of the teeth. So you have a solid volume and a hollow shell, and they only overlap in a thin region.
Tom: And that's what makes it hard. If you had two full three dee models of the same object, ICP would work fine. But here, the source cloud is missing the interior and the deep gingival side. It's like trying to fit a glove onto a hand when you can only see the palm.
Jane: That's a great analogy. And the abstract also mentions the "true correspondence is generally unknown" problem, which we've already talked about. But it also introduces their solution: the pseudo-IOS evaluation framework.
Tom: Right, and then it previews the method: GeDi for global alignment, ICP for refinement. And it mentions the metal artifact simulation. So the first page is essentially a roadmap for the whole paper.
Jane: One thing I appreciate is how they acknowledge the limitations of previous work. They cite that conventional studies relied on manual alignment or surrogate metrics, which conflates estimation error with reference error. That's a subtle but important point.
Tom: It's the difference between measuring your height with a ruler that's already bent versus measuring with a laser. The bent ruler gives you a number, but you don't know how wrong it is.
Jane: And they also mention the clinical motivation: integrating internal bone structure with high-resolution dental surface geometry is essential for surgical design and treatment simulation. This isn't just an academic exercise.
Tom: The first page also lists their contributions clearly. First, the evaluation framework. Second, the GeDi+ICP method. Third, the analysis of metal artifacts. It's a clean, well-structured paper.
Jane: And I think that clarity is a strength. You know exactly what they're claiming and how they're going to prove it. That's good science communication.
Tom: So we've covered the abstract, the motivation, the contributions. Next we need to wrap up with our final thoughts on the whole paper.
Conclusion: Tom: Alright, we're at the end of our discussion on "Coarse-to-Fine Registration of Jawbone CT and Intraoral Scan Data Using GeDi and ICP with Pseudo-IOS Ground Truth." Jane, let's pull it all together.
Jane: This paper gives us a trustworthy way to evaluate CT-to-IOS registration, and a method that actually works. The pseudo-IOS framework means we can finally measure accuracy against a known ground truth, and the GeDi+ICP pipeline delivers submillimeter errors consistently.
Tom: And that's not just for clean data. Even with simulated metal artifacts, the method stayed accurate, with mean absolute error staying between zero point five five and zero point six nine millimeters. That's a huge win for clinical applicability.
Jane: The implications go beyond dentistry. Any field that needs to register a solid volumetric scan with a surface scan — think orthopedics, or even industrial inspection — could benefit from this approach.
Tom: And the evaluation framework is reusable. Other researchers can adopt it to test their own registration methods, which could accelerate progress in the whole field.
Jane: There are limitations, of course. The dataset was small, only seven patients. And the artifacts were simulated, not real. But the framework is sound, and the results are promising.
Tom: So our verdict is that this is a solid, rigorous paper that solves a real problem and gives the community the tools to build on it. We're saying goodbye to this one, but we're excited to see what comes next.
Jane: Thanks for joining us, listeners. We'll be back soon with another paper to break down. Until then, keep your teeth clean and your point clouds registered.
Tom: See you next time.
Sho MITARAI, Hikaru KAYO, Hisashi OZAKI, Yuichiro IMAI, Megumi NAKAO
Kyoto University · Rakuwakai Otowa Hospital
eess.IV, cs.AI, cs.CV
Submitted: 2026-08-03
Updated: 2026-08-11
Comments: 9 pages, 2 figures, 1 table
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 49/100
Key concepts
- GeDi
- A learning-based three dee local descriptor used in the registration process. It examines small patches of a point cloud to create a signature for each point, which is invariant to rotation and scale. This allows the method to find matching points between the CT and intraoral scan without needing an initial good guess.
- ICP
- Iterative Closest Point algorithm, used as the second stage of registration. It refines the alignment by minimizing the distance between two surfaces after a rough initial alignment has been achieved by GeDi. This fine-tuning step is necessary for achieving high accuracy.
- Pseudo-IOS Ground Truth
- A fake intraoral scan generated directly from CT data. Because both the fake scan and the CT originate from the same coordinate frame, this method provides an exact ground truth transformation, allowing researchers to accurately measure how well their registration algorithm performs.
- Metal Artifacts
- Distortions in CT scans caused by metal objects like fillings or crowns. The paper tested how these artifacts affect performance, showing that while they degrade performance for GeDi alone, the full GeDi+ICP pipeline remained robust against simulated metal artifacts.
Terminology
Summary
Summary
This paper addresses the problem of registering jawbone CT data with intraoral scanner (IOS) data, which is essential for integrating internal bone structure with high-resolution dental surface geometry in digital dentistry and oral surgery. The authors identify two fundamental challenges: (1) the two modalities share only a limited common region around the dentition, making registration difficult, and (2) because CT and IOS data are acquired separately, their true correspondence is generally unknown, which has prevented rigorous quantitative evaluation of registration accuracy.
To address these issues, the authors propose a pseudo-IOS evaluation framework in which a point cloud emulating an intraoral scan is generated from CT data within the same coordinate frame, so that the transformation between them is known by construction and can serve as a true ground truth.
They also propose a coarse-to-fine registration method that combines a domain-generalizable local descriptor (GeDi) for initialization-independent global alignment with the iterative closest point (ICP) algorithm for local refinement,
and evaluate the influence of metal artifacts on registration.
The main contributions are stated as: "(1) an evaluation framework that enables quantitative assessment of multimodal registration against a true ground truth, and (2) a proposed GeDi+ICP method, whose accuracy and robustness to metal artifacts are demonstrated within this framework."
Problem Setting: The registration is performed independently for each arch (mandible L, maxilla U). The pseudo-IOS cloud serves as the source, and the CT-derived cloud as the target. The authors seek a rigid transformation g = (R, t) to map the source onto the target. A key property is that the pseudo-IOS cloud is generated from the CT cloud within the same normalized coordinate frame, so the ground-truth transformation g* is known by construction.
Registration Method: The proposed method estimates the rigid transformation in a coarse-to-fine manner. First, GeDi provides an initialization-independent global alignment through descriptor-based correspondence matching and RANSAC. Points are randomly sampled from both clouds, and at each sampled point, GeDi extracts and canonicalizes a local patch, encoding its multiscale geometry using a PointNet++ backbone to obtain a rotation- and scale-invariant descriptor. Candidate correspondences are established through nearest-neighbor matching in descriptor space. Because such matching produces outliers, the rigid transformation is estimated by RANSAC, which randomly samples minimal subsets of correspondences, computes candidate transformations, and counts inliers within a distance threshold. The transformation supported by the largest number of inliers is retained and re-estimated from its inliers. Taking this global alignment as the initial position, ICP alternately establishes nearest-point correspondences between the transformed source and the target, updating the rigid transformation until convergence.
Dataset Construction: For each artifact-free CT volume, a dental-region point cloud was extracted by thresholding CT values at 300 HU to retain hard tissue. TotalSegmentator was used to segment the maxillary and mandibular dentition, and oriented bounding boxes fitted to the segmentation were used to crop the dental region. All clouds were mapped to a common normalized coordinate frame. The pseudo-IOS source cloud was generated by visibility-based surface extraction, placing virtual viewpoints on a hemisphere about the occlusal axis (estimated via principal component analysis), with polar angles Θ = 0°, 30°, 55°, M = 16 azimuth angles, and ρ = 1.5, giving 33 viewpoints total. The hidden point removal operator of Katz et al. was used to determine visible points, and the union of visible points across all viewpoints formed the pseudo-IOS cloud, which is hollow and open on the deep (gingival/alveolar-bone) side. Metal artifacts were simulated following the procedure of Nakao et al. to reproduce streak and dark-band artifacts that dental metals induce in reconstructed CT images.
Experiments: The dataset consisted of jawbone CT volumes from seven patients diagnosed with jaw deformity, with pseudo-IOS point clouds generated independently for the mandible and maxilla. Initial misalignments were applied to the source cloud (random rotations within ±10° of each axis, translations within ±50 mm along each axis). Registration was repeated 10 times under every artifact condition. Nine artifact conditions were generated by designating n m = 0, 1,..., 8 teeth as metal regions. Three registration conditions were compared: ICP alone, GeDi alone, and GeDi+ICP (proposed). Evaluation metrics were mean absolute error (MAE), mean distance (MD), Hausdorff distance (HD), and computation time.
Results: GeDi+ICP achieved the smallest error across all conditions, with MAE ranging from 0.55 to 0.69 mm and HD remaining near or below 1.1 mm. GeDi alone was less accurate (MAE 1.09–3.64 mm) and markedly more sensitive to metal artifacts: under the artifact condition, its MAE for the mandible increased to 3.64 mm, whereas that for GeDi+ICP remained at 0.69 mm. GeDi+ICP also showed consistently smaller standard deviation (e.g., 0.06 mm vs. 1.16 mm for the artifact-affected mandible). ICP alone produced very large errors under every condition (overall MAE 25.89 mm, HD 20.19 mm), indicating that from the perturbed initial position it frequently converged to local minima. The computation time for GeDi+ICP was approximately 10 seconds.
A three-way repeated-measures ANOVA showed significant main effects of method (F(1,6) = 58.82, p <.01), artifact (F(1,6) = 35.62, p <.01), and jaw (F(1,6) = 20.44, p <.01). The method × artifact, method × jaw, artifact × jaw, and method × artifact × jaw interactions were also significant. The condition means indicated that these interactions were primarily associated with the marked artifact-related increase in GeDi error for the mandible, whereas GeDi+ICP showed comparatively small changes across jaw and artifact conditions.
Discussion: The authors note that the performance of the proposed method arises from combining two stages with different strengths: GeDi estimates the global position from local-descriptor correspondences without relying on the initial position, while ICP corrects residual discrepancies once an approximately correct global position is obtained. Metal artifacts had a statistically significant effect on registration accuracy, but their absolute influence was substantially smaller for GeDi+ICP than for GeDi alone, suggesting that artifact-induced surface distortion primarily disrupts descriptor correspondences during global alignment, and that ICP can compensate for part of this error. The authors acknowledge limitations including the small dataset, simulated rather than real metal artifacts, and that the pseudo-IOS framework does not reproduce the full domain gap between independently acquired CT and IOS data (lacking scanner noise, scanning-path deformation, soft-tissue interference, and partial occlusion). They conclude that validation using real CT–IOS pairs, larger cohorts, and real metal artifacts is required.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems:
-
Improvement: Implement a framework that generates paired point clouds (pseudo-IOS source and CT-derived target) within the same coordinate frame, ensuring known ground-truth transformations for quantitative evaluation.
-
What the improved AI system can do: Automatically generate synthetic intraoral scan data from CT volumes, enabling rigorous, reproducible accuracy assessment of any registration algorithm without manual alignment or surrogate metrics. This eliminates the conflation of estimation error with reference error.
-
Improvement: Integrate a two-stage pipeline: (a) GeDi-based global alignment using rotation- and scale-invariant local descriptors with RANSAC for initialization-independent pose estimation, followed by (b) point-to-point ICP refinement.
-
What the improved AI system can do: Achieve submillimeter mean absolute error (0.55–0.69 mm) across all jaw and artifact conditions, even from perturbed initial positions (±10° rotation, ±50 mm translation). It avoids local minima that cause ICP alone to fail (MAE > 25 mm) and outperforms GeDi alone (MAE 1.09–3.64 mm), with significantly lower variance (SD 0.06 mm vs. 1.16 mm under artifacts).
-
Improvement: Design the pipeline so that artifact-induced descriptor mismatches during global alignment are corrected by ICP refinement, which exploits broader surface overlap after initialization.
-
What the improved AI system can do: Maintain registration accuracy (MAE 0.69 mm for artifact-affected mandible) even when up to 8 teeth contain simulated metal artifacts, whereas GeDi alone degrades to 3.64 mm. The system statistically confirms (three-way repeated-measures ANOVA, p <.01) that artifact impact is significantly reduced, making it clinically viable for patients with dental restorations.
-
Improvement: Implement the hidden point removal operator with multi-viewpoint aggregation (33 virtual viewpoints on a hemisphere) to generate hollow, open-surface point clouds that mimic intraoral scanner coverage.
-
What the improved AI system can do: Produce realistic pseudo-IOS data that captures only scannable surfaces (occlusal, buccal, lingual) while excluding interior and deep gingival regions, enabling accurate evaluation of partial-overlap registration scenarios typical of clinical use.
-
Improvement: Integrate a systematic artifact generation protocol (0–8 metal teeth in predefined order) to create graded difficulty conditions.
-
What the improved AI system can do: Quantitatively assess registration degradation as a function of artifact severity, allowing developers to identify failure thresholds and optimize algorithms for specific clinical populations (e.g., patients with multiple restorations).
-
Improvement: Use GeDi descriptors pretrained on 3DMatch, which generalize across domains unseen during training, with 32-dimensional embeddings and local reference frame canonicalization.
-
What the improved AI system can do: Establish reliable correspondences between heterogeneous modalities (CT vs. pseudo-IOS) without fine-tuning on dental data, demonstrating cross-domain robustness that is critical for real-world deployment where training data is scarce.
-
Improvement: Implement a three-way repeated-measures ANOVA (method × artifact × jaw) with case as the repeated factor to rigorously compare registration methods.
-
What the improved AI system can do: Provide statistically significant evidence (F(1,6) = 58.82, p <.01 for method) that the proposed pipeline outperforms alternatives, enabling researchers to make data-driven decisions about algorithm selection rather than relying on anecdotal visual inspection.
-
Improvement: Optimize the pipeline to complete registration in approximately 10 seconds (9.14–10.05 s) on an NVIDIA RTX A6000 GPU.
-
What the improved AI system can do: Integrate into clinical workflows where near-real-time registration is required, balancing accuracy (submillimeter MAE) with practical usability, unlike slower global optimization methods.
-
Improvement: Leverage the framework's property that both clouds share the same coordinate frame, enabling direct MAE computation over known point correspondences.
-
What the improved AI system can do: Provide per-point error maps (as shown in Fig. 2) that visualize spatial distribution of misalignment, aiding clinicians in identifying regions of poor registration (e.g., artifact-affected areas) for targeted intervention.
-
Improvement: Implement separate registration pipelines for each arch with sign-flipped occlusal axes and independent normalization.
-
What the improved AI system can do: Handle the geometric differences between mandible (upward-facing occlusal surface) and maxilla (downward-facing), maintaining consistent accuracy (MAE 0.55–0.69 mm) across both arches, which is essential for full-arch treatment planning.
Abstract
In digital dentistry and oral surgery, the registration of jawbone CT and intraoral scanner (IOS) data is essential for integrating internal bone structure with high-resolution dental surface geometry. However, this registration is challenging because the two modalities share only a limited region in common, and their true correspondence is generally unknown. This uncertainty has prevented rigorous quantitative evaluation of registration accuracy. In this study, we propose a pseudoIOS evaluation framework in which a point cloud emulating an intraoral scan is generated from CT data within the same coordinate frame, so that the transformation between them is known by construction and can serve as a true ground truth. Using this framework, we propose a coarse-to-fine registration method that combines a domain-generalizable local descriptor (GeDi) for initialization-independent global alignment with the iterative closest point (ICP) algorithm for local refinement, and also evaluate the influence of metal artifacts on registration. In experiments on seven cases, ICP alone frequently converged to local minima from a perturbed initial position, whereas GeDi + ICP maintained submillimeter mean absolute error (MAE) across all evaluated jaw and artifact conditions (0.55-0.69 mm). A three-way repeated-measures analysis confirmed that GeDi + ICP was significantly more accurate than GeDi alone. Metal artifacts had a statistically detectable overall effect, but their absolute impact on GeDi + ICP was small.
Related papers
- Revisiting Integration of Image and Metadata for DICOM Series Classification: Cross-Attention and Dictionary Learning
- VesselSDF: Distance Field Priors for Vascular Network Reconstruction
- cSVR: Convolutional Slice-to-Volume Reconstruction
- NAIMA: Semantics Aware RGB Guided Depth Super-Resolution
- AneumoBench: A Source-Linked Benchmark for Synthetic-Geometry Transfer in Aneurysm CFD
- RETO: A Rotary-Enhanced Transformer Operator for High-Fidelity Prediction of Automotive Aerodynamics