Unsupervised Point Cloud Registration with Self-Distillation

arXiv:2409.07558 · cs.CV, cs.LG, cs.RO · Submitted 2026-08-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Unsupervised Point Cloud Registration with Self-Distillation".

Jane: The paper was written by Christian Löwens, Thorben Funke, André Wagner and Alexandru Paul Condurache from Bosch Research and University of Lübeck and Robert Bosch GmbH.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Alright, welcome back to the show, everyone. Today we are diving into a paper that has the whole registration community buzzing — it's called "Unsupervised Point Cloud Registration with Self-Distillation." Jane, I have to say, just that title alone is a mouthful, but it's so packed with meaning.

Jane: It really is, Tom. And let's break it down for our listeners who might not be deep in the weeds of computer vision. Point cloud registration is basically the problem of taking two three dee scans — think of a room scanned by a robot, or a street scanned by a self-driving car — and figuring out how to rotate and shift one scan so it perfectly lines up with the other.

Tom: Right, and the "unsupervised" part is the magic word here. Normally, to teach a neural network to do this, you need ground truth — you need to know the exact correct alignment for thousands of pairs of scans, which is incredibly expensive to collect. You'd need special survey equipment, manual labeling, the works.

Jane: Exactly. And that's where this paper from Bosch Research and the University of Lübeck comes in. They've figured out a way to train a network to do this alignment without any of those labels. They use a clever trick called self-distillation, where you have two networks — a teacher and a student — and the teacher helps train the student, but the teacher itself is just a slowly-updated copy of the student.

Tom: It's like the student is learning from its own past self, which is a wild concept. And the team — Christian Löwens, Thorben Funke, André Wagner, Alexandru Paul Condurache — they've shown this works not just on indoor RGB-D scans, but also on automotive radar, which is a totally different beast. That's huge for scalability.

Jane: It really is. Because in the real world, especially in autonomous driving, you have millions of cars driving around collecting data every day. If you can train your registration models on that unlabeled data instead of requiring a professional survey crew, you can improve your systems at a fraction of the cost.

Tom: And that's the promise here. We're going to dig into how they actually pull this off, the architecture, the training tricks, and why it beats the previous state of the art. Stick around, because this gets really interesting.

Jane: Absolutely, Tom. Let's get into the mechanics of it.

Summary: Tom: So Jane, we've set the stage. Now let's talk about what this paper actually does. The core idea, as I mentioned, is self-distillation. But the way they've structured it is really elegant. They have a teacher network that takes the original point clouds, extracts features, and uses a classic robust solver — RANSAC — to estimate the alignment.

Jane: Right, and RANSAC is that old-school algorithm that randomly samples a few points, guesses a transformation, and checks how many other points agree with it. It's not learning-based, but it's very robust to outliers. So the teacher uses this to generate what they call "pseudo labels" — essentially its best guess at the correct alignment.

Tom: Then the student network gets the same point clouds, but with a twist — they're augmented, meaning they're randomly rotated. The student has to predict the same correspondences that the teacher found, but from these harder, rotated inputs. And here's the kicker — the teacher isn't trained directly. It's updated by taking an exponential moving average of the student's weights.

Jane: So the teacher is always a slightly "older" and more stable version of the student. That's the self-distillation part. And because the teacher is always improving, the pseudo labels it generates are always getting better. It's a virtuous cycle.

Tom: Exactly. And they compare this against two previous methods — SGP and EYOC. SGP used a similar idea but had a lot of extra baggage: it needed hand-crafted features like FPFH to bootstrap the training, and it had a verifier module to filter out bad pseudo labels. The authors here show that all that extra stuff actually hurts performance.

Jane: And EYOC, which is a more recent method, used a momentum encoder like this paper, but they applied augmentation to both the teacher and the student. The authors found that this makes the bootstrap phase — the very beginning of training when the teacher is basically random — much harder. So they removed the augmentation from the teacher's input entirely.

Tom: That's such a subtle but important insight. You'd think giving the teacher the same augmented data would help it generalize, but in the unsupervised setting, it just makes the initial learning unstable. The numbers back this up — on the three deeMatch benchmark, their method hits a ninety-two point seven percent registration recall, compared to ninety point eight percent for SGP and seventy-six point eight percent for EYOC.

Jane: And on the radar dataset, it's even more dramatic. Their method gets ninety-five point eight percent registration recall, while SGP only gets ninety point two percent. That's a huge gap, and it really shows that their simplified approach generalizes much better across different sensors.

Tom: So the summary is: fewer moving parts, better results, and it works across modalities. That's a winning combination. But how do they actually make the training stable? Let's dig into that next.

Improvements: Jane: Tom, I want to pick up on that point about stability, because I think that's where the real contributions are. The paper makes two key improvements over prior work. First, they remove the need for hand-crafted bootstrap features. SGP used FPFH features to get the teacher started, but those features are designed for dense indoor point clouds and just fall apart on radar data.

Tom: Right, and that's a killer problem. If your bootstrap method doesn't work on a new sensor, your whole training pipeline is useless. The authors show that FPFH features on their radar dataset achieve a registration recall of just three point seven percent — basically random. So SGP is dead in the water before it even starts.

Jane: Exactly. And the second improvement is that they remove the pseudo label verifier. SGP had this extra module that would check the teacher's output and discard samples that looked too unreliable. But the authors found that this verifier actually hurt performance — it was throwing away challenging but learnable samples.

Tom: And they have the ablation to prove it. When they add the verifier back in, their registration recall drops from ninety-five point seven percent down to eighty-one point six percent on the radar subset. That's a massive drop. So the verifier was doing more harm than good.

Jane: But the biggest improvement, I think, is the augmentation strategy. In their setup, the teacher sees the original, unrotated point clouds, while the student sees the rotated versions. This is counterintuitive because you'd think the teacher should also see the augmented data to be robust. But the authors show that during the bootstrap phase, the teacher is too weak to handle the extra difficulty.

Tom: And they have a beautiful figure showing this. When the teacher gets augmented inputs, the student's feature-match recall crawls along near zero for the first ten epochs. But when the teacher gets the clean inputs, the student follows almost the same trajectory as a supervised model. It's night and day.

Jane: So the improvement here is really about understanding the bootstrapping dynamics. You have to give the teacher an easy job at the start, and only then can the student learn to handle the hard stuff. It's like teaching a kid to ride a bike — you don't start them on a mountain trail, you start them on a flat parking lot.

Tom: That's a great analogy, Jane. And it's these kinds of practical insights that make the paper so valuable. It's not just a new architecture — it's a set of design principles for unsupervised learning in geometric tasks. Now, let's talk about the first page of the paper, because there's some really interesting motivation there.

First Page: Jane: So Tom, looking at the first page of "Unsupervised Point Cloud Registration with Self-Distillation," the authors open with a really compelling motivation. They show this diagram comparing the cost and quality of data. On one side, you have point cloud pairs with ground truth — expensive and limited. On the other side, you have crowdsourced data from consumer cars — cheap and abundant.

Tom: And the key insight is that the quality of pseudo labels generated by their method is close to ground truth. That's the whole bet of this paper — that you can get ninety percent of the performance of supervised learning without paying for the labels. And their results on both three deeMatch and the radar dataset back that up.

Jane: Right, and they also position this within the broader context of autonomous driving. Collecting ground truth poses for a city is a massive undertaking. You need professional survey vehicles with expensive sensors, and you have to do it repeatedly as the city changes. But crowdsourced data from regular cars is being generated all the time.

Tom: And that's where I think this paper could have a real impact. If you can train registration models on this unlabeled crowdsourced data, you can keep your maps fresh without the cost of a full survey. That's a huge deal for companies like Bosch, which is why they're funding this research.

Jane: Exactly. And they also mention that their method simplifies the training procedure compared to related work. They don't need consecutive point cloud frames, they don't need progressive datasets, and they don't need hand-crafted features. That's a big deal for practitioners because it means the method is much easier to adapt to a new sensor or a new environment.

Tom: And the results speak for themselves. On three deeMatch, they're within zero point four percentage points of the supervised baseline on registration recall. On radar, they're within zero point eight percentage points. That's remarkable for an unsupervised method.

Jane: It really is. And I think the implications go beyond just registration. This self-distillation framework could be applied to other geometric tasks — like object detection or scene flow — where ground truth is also expensive. The principles they've established here — the augmentation strategy, the EMA teacher, the robust solver — are pretty general.

Tom: Absolutely. And that's what makes this paper so exciting. It's not just a better registration model; it's a template for how to do unsupervised learning in three dee perception. Let's bring in Lu and Meng to get their take on the broader impact.

Lu: I'm really excited about the generalizability here. The fact that they tested on radar — which has completely different noise characteristics than LiDAR or RGB-D — shows that the method isn't overfit to one modality. That's rare in this field.

Meng: And from an engineering standpoint, the simplicity is the killer feature. I've tried to implement SGP before, and the verifier and the FPFH bootstrapping are a nightmare to tune. This paper removes all of that. I could probably get this running on a new dataset in a day.

Conclusion: Tom: Alright, we've covered a lot of ground on "Unsupervised Point Cloud Registration with Self-Distillation." Let's wrap this up. Jane, what's the one-sentence takeaway?

Jane: The takeaway is that you can train a point cloud registration network without any ground truth labels by using a self-distillation setup where a stable teacher generates pseudo labels for an augmented student, and this simple recipe beats all prior unsupervised methods while nearly matching supervised performance.

Tom: And it does that on two very different datasets — indoor RGB-D scans and automotive radar. That's the kind of cross-modal robustness that makes a paper stand out.

Lu: I'd add that the theoretical insight about the bootstrap phase is the real gem. Knowing that you need to keep the teacher's input clean during early training is a principle that could apply to many other self-supervised learning tasks.

Meng: And practically, this means we can finally leverage the massive amounts of unlabeled data that's already being collected by fleets of vehicles. That's a game-changer for anyone building large-scale mapping systems.

Jane: And let's not forget the future work they mention — integrating a differentiable RANSAC could allow end-to-end training, which might push the performance even higher. There's a lot of room to build on this.

Tom: Absolutely. So to the authors — Christian, Thorben, André, and Alexandru — great work. This is a paper that I think will be cited heavily in the coming years. And to our listeners, thanks for tuning in. We'll be back soon with another exciting paper from arXiv.

Jane: Until next time, keep your point clouds aligned and your pseudo labels clean. Goodbye, everyone!

Christian Löwens, Thorben Funke, André Wagner, Alexandru Paul Condurache

Bosch Research · University of Lübeck · Robert Bosch GmbH · University of Lübeck

cs.CV, cs.LG, cs.RO

Submitted: 2026-08-10

Comments: Oral at BMVC 2024

Code: https://github.com/boschresearch/direg

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 61/100

Terminology

Summary

Summary

This paper introduces DiReg, a self-distillation approach for unsupervised rigid point cloud registration. The authors address the scalability problem of supervised deep learning methods, which require expensive ground truth poses for training. They propose a student-teacher architecture where the teacher generates pseudo labels on the fly to train the student, eliminating the need for ground truth labels.

The method is described as follows: "We adopt a student-teacher architecture, where the teacher generates pseudo labels on the fly to train the student. This continuously improves the pseudo labels, in contrast to SGP, which generates new labels only after several epochs of training. Our teacher network is updated using an EMA of the student’s parameters and therefore, shares the same architecture. Since we are in an unsupervised setting, this process is also termed self-distillation." The feature matcher consists of the FCGF feature extractor followed by a nearest neighbor search. The teacher receives the original point clouds, while the student receives augmented views. The teacher predicts correspondences via nearest neighbor search in feature space, then RANSAC (optionally followed by ICP) estimates a transformation optimizing the unsupervised inlier ratio. Refined correspondences are obtained by a nearest neighbor search in coordinate space, which serve as pseudo labels. The student is trained with a hardest-contrastive loss, and the teacher is updated via an exponential moving average (EMA) of the student's parameters with a cosine schedule.

Key contributions include simplifying unsupervised registration by removing the pseudo label verifier, hand-crafted bootstrap features, and progressive datasets used in prior work. The authors also demonstrate that applying data augmentation to the teacher's input can impede robust bootstrapping. They state: "we draw inspiration from Noisy Student Training and only apply the augmentation on the student’s input. Surprisingly, keeping the data’s original orientation substantially impacts the bootstrap phase as we demonstrate in our ablations. It makes the registration task much easier and accelerates the training. They further note that SGP utilizes classical features from FPFH for bootstrapping. These features are effective as initialization for RGB-D point clouds but not discriminative enough for radar... After removing the augmentation for the teacher, we found that FPFH features are counterproductive. So, we omitted them and instead trained them with a teacher network that is randomly initialized."

Experiments are conducted on two datasets: 3DMatch (RGB-D indoor scenes) and a proprietary automotive radar dataset. On 3DMatch, DiReg achieves a feature match recall (FMR) of 92.7%, registration recall (RR) of 91.6%, and inlier ratio (IR) of 24.1%, outperforming SGP (91.3% FMR, 90.8% RR, 22.4% IR) and EYOC (62.5% FMR, 76.8% RR, 10.6% IR), while the supervised FCGF baseline achieves 93.5% FMR, 92.0% RR, and 24.3% IR. On the radar dataset, DiReg achieves an RR of 95.8%, RTE of 0.413 m, and RRE of 0.156°, compared to supervised FCGF (96.6% RR, 0.355 m, 0.160°), SGP (90.2% RR, 0.596 m, 0.219°), and EYOC (91.7% RR, 0.487 m, 0.166°). The authors note that our approach’s performance is close to that of the supervised method on radar.

Ablation studies on a radar subset show that removing ICP does not degrade performance (RR 95.7% without ICP vs. 95.3% with ICP), the pseudo label verifier significantly hurts performance (RR drops to 81.6%), and FPFH features fail completely on radar (RR 3.7%). The authors also show that training without augmentation for the teacher follows the supervised trajectory, while training with augmentation requires more epochs to overcome the bootstrap phase.

The paper concludes: "This study presents a novel self-distillation framework for point cloud registration. We show how to bootstrap student-teacher networks unsupervised without the need for initial hand-crafted features, verifiers, or progressive datasets, while still reaching supervised performance on RGB-D and radar scans." Future work directions include non-contrastive losses to eliminate negative pairs and integrating differentiable RANSAC for end-to-end training.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved system can do:


  1. Remove the need for ground-truth poses during training
  • Implement the self-distillation framework (DiReg) where a teacher network (EMA-updated) generates pseudo labels on the fly, and a student network learns from augmented views.

  • Eliminate hand-crafted bootstrap features (e.g., FPFH), pseudo-label verifiers, and progressive datasets—all of which require modality-specific tuning.

  1. Apply augmentation only to the student, not the teacher
  • This stabilizes the bootstrap phase, preventing the randomly initialized teacher from failing on rotated inputs.

  • The student still learns rotation invariance, while the teacher sees the original orientation for reliable initial correspondences.

  1. Use a learning-free robust solver (RANSAC) in the teacher
  • RANSAC optimizes for the unsupervised inlier ratio, providing high-quality pseudo correspondences without any supervision.

  • Optionally, ICP can be added for fine-tuning, but the paper shows it is not always necessary (e.g., on radar).

  1. Replace the pseudo-label verifier with a simple distance threshold
  • The verifier in SGP removes challenging samples and hurts performance; instead, keep all samples and use a coordinate-space nearest neighbor search with a threshold (τ2) to refine correspondences.
  1. Use a contrastive loss with teacher-generated positive pairs
  • The hardest-contrastive loss (from FCGF) is adapted so that positive pairs come from the teacher’s refined correspondences, and negative pairs are mined accordingly—no ground truth needed.
  1. Update the teacher via exponential moving average (EMA) with a cosine schedule
  • This provides continuous, stable improvements to pseudo labels, unlike SGP’s epoch-wise label regeneration.

  • It also allows the option to share parameters (θt = θˢ) for memory efficiency with minimal performance loss.

  • Register point clouds from any sensor modality (RGB-D, automotive radar, LiDAR) without any labeled data, achieving performance close to supervised methods (e.g., 92.7% FMR on 3DMatch vs. 93.5% supervised; 95.8% RR on radar vs. 96.6% supervised).

  • Train on large-scale, crowdsourced, unlabeled data (e.g., from consumer vehicles) to learn robust features that generalize across environments, without the cost of ground-truth surveys.

  • Bootstrap from scratch—no need for classical features like FPFH, which fail on radar due to their reliance on 3D coordinates only.

  • Adapt to new modalities with minimal hyperparameter tuning—the method has fewer components (no verifier, no progressive dataset) and works out-of-the-box on both indoor RGB-D and sparse radar scans.

  • Provide a memory-efficient variant (shared student-teacher parameters) for deployment on resource-constrained devices, with only a 1% drop in registration recall.

  • Enable continuous learning—the EMA teacher updates pseudo labels in real-time, so the system improves during training without retraining from scratch.

  • Autonomous driving: Register radar or LiDAR scans from multiple drives to build high-definition maps, using only unlabeled fleet data.

  • Robotics: Align point clouds from different sensors (e.g., depth cameras) in unknown environments for SLAM or object manipulation.

  • AR/VR: Register 3D scans of indoor scenes for augmented reality without manual annotation.

  • Industrial inspection: Align point clouds from different scans of the same object (e.g., for quality control) without expensive metrology equipment.

This system is ready for deployment in scenarios where labeled data is scarce, expensive, or impossible to collect, while maintaining near-supervised accuracy.

Sources

Related papers