SonarVoxNet: Diver Detection in 3D Bounding Box using 3D Sonar

arXiv:2610.01644 · cs.RO · Submitted 2026-10-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Robotics Radio. Generated commentary on the latest robotics and control papers.

Rosa: Today's paper: "SonarVoxNet: Diver Detection in 3D Bounding Box using 3D Sonar".

Dev: Autonomous underwater vehicles require continuous tracking of a diver's 3D position and full-body orientation for safe human-robot interaction, but existing forward-looking sonar methods discard elevation information,

Rosa: First, who's behind it and why it matters.

Title and authors: Rosa: To kick things off, let's talk about SonarVoxNet: Diver Detection in three dee Bounding Box using three dee Sonar and who came up with this work. We need to set the context for what these researchers actually did.

Dev: I’m ready to hear the details on the team behind it so we can understand their background before diving into the technical specifics of SonarVoxNet.

Taro: Before we get too deep into the math, I want to know what kind of autonomy challenges this paper was designed to solve in terms of perception.

Rosa: This work comes from Eugene Park, Jiwon Lee, Seyoung Kan, Trung Dong, Xiaomin Lin, and Jane Shin; they are researchers who clearly have a strong background in both robotics and advanced computer vision.

Dev: I've looked at the authors’ previous work—they seem to be focused on areas like trajectory planning and resilient systems, which gives me confidence that this paper will have a rigorous engineering foundation.

Taro: That focus on resilience is exactly what we need when we talk about autonomy in unpredictable environments where standard perception methods fail, which is the main motivation here.

Rosa: The title itself tells us the main goal: to detect divers in three dee using sonar, specifically focusing on getting that three dee bounding box right.

Dev: So, if I boil it down for the listeners, this paper is about taking a limitation in current forward-looking sonar—the lack of elevation information—and fixing it by introducing a new detection method.

Taro: The implication here is that we can finally move past systems that are limited to upright objects and start dealing with the complexity of human body poses underwater.

Rosa: Precisely; this work suggests that full nine-DoF tracking is achievable with three dee sonar returns, which opens up a lot of possibilities for human-robot interaction.

Dev: I'm thinking the real impact is in creating tools where we can safely monitor divers or assist them in ways that require knowing their exact posture, not just their location.

Taro: It means that autonomous underwater vehicles won't just be bumping into things; they can actually understand the dynamic three dee state of a human subject.

Rosa: That's a big conceptual shift; it moves us from simple localization to understanding full body geometry in complex settings like caves.

Dev: I think the immediate practical application is in developing safer navigation algorithms for AUVs that need to operate near divers without causing interference due to unexpected body movement.

Taro: And for autonomy research, this sets a new benchmark for how perception systems should handle sparse, noisy data where ground truth orientation is scarce.

Rosa: So, we're looking at a system that uses a sophisticated voxel encoder and an anchor-free head to bridge that gap between what sonar gives us and what we need for safe tracking.

Dev: And I’m keen to see how well the authors handle the noise inherent in three dee sonar returns when they try to reconstruct those full nine-DoF boxes.

The paper's summary: Rosa: Now that we know who wrote it, let's look at what SonarVoxNet actually proposes in this paper and how it achieves its goal of full orientation detection.

Dev: I want to focus on the core methodology here—how they handle the transition from sparse three dee points to a detailed three dee bounding box prediction.

Taro: I'm interested in the specific mechanism they use to move away from old methods, like just using yaw-only output, and what that new representation actually does for tracking.

Rosa: SonarVoxNet adapts a voxel-based encoder and an anchor-free center-based detection head to process sparse sonar data, replacing the traditional yaw-only rotation with a continuous 6D rotation parameterization.

Dev: That continuous 6D parameterization is key; it’s not just guessing an angle anymore, it's predicting a full orientation using two vectors that are then mapped to a rotation matrix via Gram–Schmidt orthonormalization.

Taro: That sounds like the mechanism that makes the difference because it avoids singularities and allows for a continuous representation of any pose, which is exactly what we need for diverse poses.

Rosa: They also present Diverthree dee as a new dataset, comprising two hours, thirty-seven thousand ten frames with thirty-four thousand five hundred sixty-eight annotated instances collected at a natural cave-diving site near Gainesville, FL.

Dev: The authors emphasize that this dataset is significant because it's the first public sonar dataset with full three dee orientation labels for divers in poses rarely seen in pool settings.

Taro: Having data that captures realistic pitch and roll from natural environments is incredibly valuable because it trains the AI to generalize beyond simple, upright scenarios.

Rosa: To summarize, they use a voxelize-encode-detect structure that collapses the three dee feature map into a Bird-eye-view feature map, which then feeds into several branches in the detection head to predict center location and box extent.

Dev: The detection head has branches for center localization using a heatmap branch, an offset branch for displacement, and another for height regression to recover the full three dee center coordinates.

Taro: And critically, they also have a size branch that regresses the box dimensions in log space to predict the extent of the object accurately.

Rosa: So, in short, they tackle sparse sonar by using a dense feature map approach and then use specialized heads to recover not just location but also full three dee orientation and size information for the bounding box.

Dev: And I'm looking forward to seeing how those regression losses—the heatmap loss and the regression loss—actually guide the network toward making these complex, multi-dimensional predictions correctly.

The paper's improvements: Rosa: We’ve covered what SonarVoxNet does, but now let's focus on the specific enhancements and contributions they highlight in this paper.

Dev: I want to know about the key changes they made to the original detection pipeline that led to better performance metrics, like those scores on Diverthree dee.

Taro: What about the comparison they made between using full SO(three) rotation versus just yaw-only output? That seems like a major point of contention in their analysis.

Rosa: The authors explicitly show that adopting full SO(three) rotation is the dominant factor behind accurate three dee detection on sonar data, even more so than the choice of backbone used in their experiment.

Dev: That confirms our suspicion that representing the orientation correctly is more important than just having a complex feature extractor; it’s about the geometric representation itself.

Taro: The paper also points out that a sonar-specific, orientation-preserving augmentation strategy yielded additional gains in both detection accuracy and orientation fidelity compared to standard training methods.

Rosa: So, they're not just presenting a new model; they are showing that the combination of full rotation representation and smart data augmentation is what pushes the performance up.

Dev: I'm also interested in the causal geometry refinement stage, which uses an interacting multiple-model (IMM) filter to handle out-of-BEV-plane geometry during inference.

Taro: That refinement module is important because it addresses potential jitters in predicted geometry during inference by ensuring the operations only depend on frames one through t, which stabilizes the output.

Rosa: It sounds like they've built a robust system that doesn't just predict a static box but actively refines it using temporal context to make the final three dee bounding box more accurate.

Dev: That temporal refinement mechanism is what I think gives us the edge in terms of practical deployment, as it helps smooth out those inevitable small errors you get from noisy sensor data during live operation.

Conclusion: Rosa: So, to wrap up our discussion on SonarVoxNet: Diver Detection in three dee Bounding Box using three dee Sonar, we've seen that this paper successfully introduces a method for predicting full nine-DoF oriented bounding boxes from sparse sonar.

Dev: I think the main conclusion is that moving to a continuous 6D rotation parameterization and incorporating causal geometry refinement significantly improves detection accuracy and orientation fidelity over previous methods.

Taro: From an autonomy research viewpoint, the fact that they've provided a dataset like Diverthree dee means there’s now more real-world data available for training systems meant to understand complex human poses.

Rosa: It really opens up avenues for developing safer autonomous underwater vehicles capable of interacting with divers in ways that require knowing their precise three dee state.

Dev: The challenge remains in ensuring the causal refinement stage runs efficiently enough during live operation so we can get those real-time improvements we discussed earlier.

Taro: For future work, I think exploring how this system performs when faced with extreme environmental disturbances or completely novel object shapes would be a logical next step.

Rosa: It’s certainly a compelling paper that shows the potential of three dee sonar to move beyond simple localization toward true geometric understanding of dynamic subjects in the water.

Dev: I'm just hoping we can see this technology integrated into something where it can handle high-frequency updates without introducing significant lag or failure modes.

Taro: Well, SonarVoxNet is a solid piece of work that proves that three dee sonar has more capability than previously assumed for tracking dynamic objects in complex settings.

Eugene Park, Jiwon Lee, Seyoung Kan, Trung Dong, Xiaomin Lin, Jane Shin

cs.RO

Submitted: 2026-10-01

Updated: 2026-10-01

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 84/100

The gist: Autonomous underwater vehicles require continuous tracking of a diver's 3D position and full-body orientation for safe human-robot interaction, but existing forward-looking sonar methods discard

Key concepts

SonarVoxNet
A model that detects divers in 3D using sonar data by predicting full 9-DoF oriented bounding boxes. It uses a voxel encoder to process sparse sonar points and a detection head that predicts continuous 6D rotation parameters, allowing it to handle diverse diver poses accurately.
Voxelization and Encoding
This process takes raw, sparse 3D sonar points and organizes them into a regular grid of voxels. Points within the same voxel are then aggregated by a Voxel Feature Encoder (VFE) to create a fixed-length descriptor. This dense tensor is then processed through 3D convolutions to extract meaningful features.
Full SO(3) Rotation Representation
Instead of using just a single yaw angle, SonarVoxNet predicts the full 3D rotation matrix (R in SO(3)) by outputting two vectors. These vectors are then mapped to a continuous 6D representation using Gram–Schmidt orthonormalization, which is crucial for accurately modeling the wide range of diver orientations.
Causal Geometry Refinement
A lightweight stage applied during inference that uses an interacting multiple-model (IMM) filter in Euclidean space for position and a similar filter in the SO(3) tangent space for rotation. This ensures that geometry predictions are refined frame-by-frame, depending only on previous frames, to reduce geometric jitters.

Terminology

Summary

Autonomous underwater vehicles require continuous tracking of a diver's 3D position and full-body orientation for safe human-robot interaction, but existing forward-looking sonar methods discard elevation information, creating a fundamental geometric bottleneck. This paper introduces SonarVoxNet, the first 3D sonar diver detector to predict full 9-DoF oriented bounding boxes using a continuous 6D rotation parameterization, and presents Diver3D, the first public dataset with full 3D orientation labels for divers in diverse poses.

How it works

SonarVoxNet adapts a voxel-based encoder and an anchor-free center-based detection head to sparse, noisy 3D sonar returns. The pipeline follows a 'voxelize-encode-detect' structure: given a 3D sonar point cloud P = pi, points are quantized into a regular voxel grid of resolution (vx, vy, vz), and points within the same voxel are aggregated by a Voxel Feature Encoder (VFE) into a fixed-length descriptor. This dense voxel tensor is then processed by a 3D convolutional middle stage that gradually collapses the z-axis, producing a dense 2D feature map F ∈ R C×H×W. A 2D convolutional encoder then refines F into the Bird-eye-view (BEV) feature, which is shared by the detection head.

Detection Head and Rotation Representation

The detection head operates on the shared BEV feature map and consists of several sibling branches. For center localization, a heatmap branch predicts a single-diver class score map Hˆ ∈ [0, 1]H×W. An offset branch regresses the continuous displacement (δx, δy) from the discretized peak to the true center, and a separate branch regresses the object height z to recover the full 3D center c = (x, y, z). The size branch regresses box extent (l, w, h) in log space. Crucially for orientation, instead of a single yaw angle as in previous methods that assume upright targets, SonarVoxNet predicts a full 3D rotation R ∈ SO(3). This is achieved by outputting two vectors (a1, a2) ∈ R 6, which are mapped to a rotation matrix using Gram–Schmidt orthonormalization to yield the continuous 6D representation.

Training Objective and Causal Geometry Refinement

SonarVoxNet is trained end-to-end with a multi-task loss L = Lhm + λreg. The heatmap loss (Lhm) supervises the predicted heatmap Hˆ using a reduced Gaussian focal loss, where the term 1−Huvβ down-weights penalties near a true center. The regression loss (Lreg) minimizes the difference between predicted targets t⋆ and ground truth targets t⋆ = (δ⋆x, δ⋆y, z⋆, log l⋆, log w⋆, log h⋆,(ρ⋆)T)T. To address potential jitters in predicted geometry during inference, a lightweight causal refinement stage is applied. This stage uses an interacting multiple-model (IMM) filter in Euclidean space R cubed for the center and an analogous filter in the tangent space of SO(3) for rotation, ensuring that the refinement operations depend only on frames 1:t.

Dataset and Performance Analysis

The model is evaluated on Diver3D, the first public 3D sonar dataset with full 3D orientation labels for human divers. The dataset comprises 2 hours, 37,010 frames and 34,568 annotated instances collected at a natural cave-diving site. Annotators provide a dedicated full SO(3) bounding box defined relative to a diver-centric right-handed body frame. Performance comparisons on the Diver3D test set show that SonarVoxNet (no augmentation) achieves an AP3D@0.50 of 0.233 and an SDS of 0.574, whereas SonarVoxNet (Ours) achieves a score of 0.359 and a SDS of 0.673, demonstrating an improvement in detection accuracy and orientation fidelity that is the dominant factor behind accurate detection on sonar data. Furthermore, the causal geometry refinement module improves every geometry metric—translation, scale, and full-3D orientation—and every AP3D threshold.

Ablation Study and Conclusion

The ablation study confirms that adopting full-SO(3) rotation rather than yaw-only is the dominant factor behind accurate 3D detection on sonar data. The continuous 6D representation was found to be superior to other parameterizations, as it avoids singularities and is essential for regressing the wide range of diver poses. While the sparse backbone (sparse voxel) performs level or better on strict AP3D@0.

Improvements for AI systems

Here are the specific improvements an AI system, based on SonarVoxNet and Diver3D, can achieve:

  1. Enhanced Underwater Human-Robot Interaction (HRI) Safety: The system will be able to continuously track a diver's full 9-DoF orientation (position and complete body pose—pitch, roll, yaw). This allows for significantly safer autonomous underwater vehicle (AUV) navigation and interaction by enabling the AUV to predict where a moving diver will go in real-time, rather than relying on simplified or unreliable position estimates.

  2. Robust Detection in Turbid/Noisy Underwater Environments: By leveraging 3D sonar data with a full SO(3) rotation head, the system overcomes the fundamental limitation of yaw-only detectors. This makes object detection reliable even when targets are pitching and rolling freely (as divers do), improving performance in real-world cave or natural underwater settings where optical vision fails due to scattering and turbidity.

  3. Accurate Pose Estimation for Diver Tracking: The system can accurately estimate the full 3D orientation of a diver from sparse, noisy 3D sonar returns alone. This is a significant leap, as it provides the necessary input for subsequent, more complex tasks like per-joint skeletal pose estimation and diver tracking algorithms that are currently impossible with standard forward-looking sonar.

  4. Improved Model Generalization: The system benefits from a dataset (Diver3D) explicitly annotated with diverse, nonupright poses collected in natural environments. This training allows the model to generalize its understanding of human body geometry across a much wider range of orientations than models trained on upright terrestrial data, making it more robust for deployment in varied aquatic scenarios.

  5. Real-Time Geometric Refinement: The inclusion of the causal geometry refinement module allows the system to produce spatially coherent 3D bounding boxes at inference time. This corrects jitters and improves the quality of localization, scale, and orientation estimates by iteratively refining predictions using temporal information (short tracks) without requiring further training or complex real-time computation beyond what is necessary for online processing.

In essence, this improved AI system transforms underwater perception from simple object localization into a sophisticated tool capable of understanding the dynamic 3D posture of human subjects in complex aquatic environments.

Sources

Related papers