The Common Objects Underwater (COU) Dataset for Robust Underwater Object Detection
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "The Common Objects Underwater (COU) Dataset for Robust Underwater Object Detection".
Jane: The Common Objects Underwater (COU) Dataset addresses a critical gap in underwater object detection by providing an instance-segmented image collection of commonly found man-made objects across diverse aquatic environments,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: The title itself, "The Common Objects Underwater Dataset," really tells you what they did, but it’s more than just a collection of pictures of things under water.
Jane: It’s about moving past just marine life and giving computer vision models data on the man-made stuff we actually see in underwater environments.
Lu: They built this dataset to bridge that gap between what we train robots on land and what happens when those robots go into a real, messy aquatic setting.
Meng: So, the main implication is that if you want an autonomous underwater vehicle to reliably find tools or debris, they need training data that reflects those specific conditions.
Tom: Exactly. It changes the training ground entirely for any AI trying to navigate or inspect underwater without getting confused by unexpected objects.
Jane: The authors put together this collection across different water types—pools, lakes, and the ocean—which gives their models a much tougher challenge than just one type of water.
Lu: They didn't just collect data; they specifically focused on lighting conditions and visibility changes because that’s a huge variable in real underwater scenes.
Meng: From an engineering standpoint, this means we can build detectors that aren't just good in a lab setting but actually work when the robot is operating in murky lake water or deep ocean.
Tom: It sets a new standard for what makes an object detection dataset useful: it has to be diverse enough to handle the real-world mess.
Jane: The authors are showing us how instance segmentation, not just simple bounding boxes, helps capture those specific objects with more accuracy underwater.
Lu: It’s about making sure the AI understands the shape and context of a plastic bottle or a dive tool in three dimensions, not just as a blob on a 2D image <ref:2502.20651#pg1>.
Tom: This work lays some really solid groundwork for deploying these systems on actual AUVs that need to perform inspection tasks.
Jane: It’s important because it shows that creating domain-specific data is one of the most effective ways to make AI useful in complex, real-world physical spaces.
Lu: The next thing we should look at is how researchers are combining this with other existing datasets or exploring how these models handle the specific visual degradation water causes.
Conclusion: Tom: So we're wrapping up our look at the Common Objects Underwater Dataset for Robust Underwater Object Detection.
Jane: It boils down to this collection of segmented images that gives AI a much better picture of what’s actually floating around underwater, not just what’s swimming.
Lu: The authors really focused on making sure this data covers things people actually need to find when they are working in the deep ocean or on a lake.
Meng: So for someone listening who's just driving or walking, it means the AI being used by a robot won't just look for fish; it can spot trash, tools, and even other robots in unexpected places.
Tom: Right. It’s about training those systems to handle the real-world mess of underwater exploration without getting totally confused by everything around them.
Jane: They did this by collecting data across different water settings—pools, lakes, and the open ocean—which makes sure the AI learns to adapt to changing light and visibility conditions.
Lu: This approach is key because it shows that just having a lot of general underwater pictures isn't enough; you need specific examples of objects in those varied environments.
Tom: And the authors made their point about using instance segmentation, which helps the AI understand exactly what an object is, not just a rough outline.
Jane: That precision matters when you’re trying to identify small pieces of debris or specialized dive equipment in a murky setting.
Lu: The real value here is how this dataset sets the stage for building detectors that are fast enough to run right on the robot itself, which is what they aimed for originally.
Meng: So if you're thinking about deploying new underwater technology, this paper shows you what kind of training data actually makes a system reliable in a practical field mission.
Tom: It’s establishing a foundation where AI training moves away from just looking at land and starts focusing on the specific visual problems underwater presents.
Jane: And they're laying out how to build even more specialized collections for different kinds of underwater challenges moving forward.
Rishi Mukherjee, Sakshi Singh, Jack McWilliams, *Junaed Sattar
Department of Computer Science & Engineering · Minnesota Robotics Institute
cs.CV
Submitted: 2025-02-28
Updated: 2025-02-28
Code: https://github.com/facebookresearch/detectron2
Importance score: 77/100
The gist: The Common Objects Underwater (COU) Dataset addresses a critical gap in underwater object detection by providing an instance-segmented image collection of commonly found man-made objects across
Key concepts
- Instance Segmentation
- This is a computer vision technique that goes beyond simple bounding boxes by outlining every individual object within an image. For underwater objects, this means precisely defining the exact shape and boundaries of items like plastic bottles or dive tools, which is crucial for accurate detection on robots.
- Domain Shift Mitigation
- Domain shift occurs when a model trained on one type of data (like land photos) performs poorly on a different environment (like underwater). COU tackles this by training models specifically on varied underwater conditions—different lighting, turbidity, and locations—ensuring the resulting detectors are robust across these challenging real-world settings.
- Instance-Segmented Collection
- The dataset collects images where every object is not just boxed but fully outlined. This high level of detail helps AI understand the context and shape of underwater items, which is vital for Autonomous Underwater Vehicles (AUVs) to correctly identify and interact with objects in complex aquatic settings.
Terminology
Summary
The Common Objects Underwater (COU) Dataset addresses a critical gap in underwater object detection by providing an instance-segmented image collection of commonly found man-made objects across diverse aquatic environments, making it invaluable for training robust, real-time detectors for Autonomous Underwater Vehicles (AUVs). This dataset is significant because it moves beyond terrestrial datasets to provide domain-specific data that mitigates the severe domain shift encountered when applying models trained on land to underwater scenarios.
The gist: COU contains approximately 10K segmented images, annotated from images collected during a number of underwater robot field trials in diverse locations.
Dataset Scope and Diversity
The COU dataset was created to address the lack of datasets with robust class coverage curated for underwater instance segmentation and to address the lack of diversity in object classes since commonly available underwater image datasets focus only on marine life <ref:2502.20651#pg2>. Currently, COU contains images from both closed-water (pool) and open-water (lakes and oceans) environments, of 24 different classes of objects including marine debris, dive tools, and AUVs. The dataset expands on terrestrial object categories by annotating a wide range of underwater objects such as pollution: plastic bags, bottles,
and dive equipment: snorkels, flippers, goggles
<ref:2502.20651#pg4>.
Data Collection and Environment Selection
The key motif for creating COU was to maximize the diversity of lighting conditions and environments to enable robust use onboard underwater robots <ref:2502.20651#pg5>. Data was collected from three different environments: (i) in a well-lit pool setting, (ii) two lake locations with low visibility underwater, and (iii) in the ocean with mixed-lighting conditions <ref:2502.20651#pg8>. Specific environments included a well-lit pool setting at maximum depth of 4.57 m, two Minnesota lakes (Lake Superior and Green Lake) with high turbidity, and the ocean off the coast of Barbados with better lighting conditions but little turbidity <ref:2502.20651#pg9>. The footage was captured using a GoProTM camera at a resolution of 1920 × 1080, and 30 frames per second (FPS), and with a linear lens setting <ref:2502.20651#pg9>.
Annotation Process
The labeling process for COU was completed in two phases <ref:2502.20651#pg10>. During the first phase, images were manually annotated with bounding box annotations using the CVAT labeling tool <ref:2502.20651#pg10>. Next, the Segment Anything Model (SAM) was used to automate the segmentation process by passing bounding box annotations along with the images <ref:2502.20651#pg10>. This automation allows for speed, as it takes roughly 12 − 15 seconds to annotate a single bounding box, while a segmentation mask can take over 2 minutes on average <ref:2502.20651#pg10>. The main hardware used was an Nvidia 4080M GPU, and the processing time for the SAM model was 0.56 seconds per image <ref:2502.20651#pg10>.
Experimental Evaluation
To assess efficacy, COU was evaluated using three state-of-the-art models: YOLOv9 [8], Mask R-CNN [10], and Mask2Former [31]. Benchmarking experiments compared performance on the validation and test sets of COU, showing that the improved performance of COU-trained detectors over those solely trained on terrestrial data demonstrates the clear advantage of training with annotated underwater images. Efficiency experiments measured inference speed and hardware used, testing models on a Nvidia Jetson Orin ‘NX’ and an Nvidia RTX 2080 Ti <ref:2502.20651#pg5>. Field robot experiments conducted qualitative evaluations using unseen data from ocean environments with the YOLOv9-C model deployed on the MeCO AUV <ref:2502.20651#pg8>.
Performance Metrics
The evaluation utilized standard accuracy and efficiency metrics, including Average Precision (AP) for pixel-wise segmentation <ref:2502.20651#pg9>. The Mask R-CNN model showed the highest mAP@.5-.95 in the models trained from scratch on COU <ref:2502.20651#pg9>. However, YOLO models performed worse on average than the other two architectures, trading off performance for accuracy, compared to Mask R-CNN and Mask2Former which are known for precise segmentations <ref:2502.20651#pg9>. The compact YOLOv9-C model showed the fastest response time on both hardware platforms with an inference time of 150 milliseconds on the Jetson Orin, equating to roughly 6−7 FPS <ref:2502.20651#pg9>. This performance satisfies the semi-real time criterion for deployment on edge computing devices/onboard compute units <ref:2502.20651#pg9>.
REFERENCES
[1] IndustryARC, “Underwater Unmanned Vehicles Market Report,” 2025. [Online]. Available: https://www.industryarc.com/Report/7397/underwater-unmanned-vehicles-market-report.html <ref:2502.20651#pg9>
[2] S. Fattah, F. Abedin, M. Ansary, M. Rokib, N. Saha, and C. Shahnaz, “R3Diver: Remote Robotic Rescue Diver for Rapid Underwater Search and Rescue Operation,” in 2016 IEEE Region 10 Conference (TENCON), 2016, pp. 3280–3283 <ref:2502.20651#pg10>
[3] S. Venkatesan, “AUV for Search & Rescue at Sea - An Innovative Approach,” in 2016 IEEE/OES Autonomous Underwater Vehicles (AUV), 2016, pp. 1–9 <ref:2502.20651#pg10>
[4] N. M. Benoist, K. J. Morris, B. J. Bett, J. M. Durden, V. A. Huvenne, T. P. Le Bas, R. B., Wynn, S., Ware, and H., Ruhl, “Monitoring Mosaic Biotopes in a Marine Conservation Zone by Autonomous Underwater Vehicle,” Conservation Biology, vol 33 no 5 pp 1174–1186 <ref:2502.20651#pg10>
[5] T.-Y. Lin, M. Maire, S. J. Belongie, L. D., Bourdev, R., Girshick, J., Hays, P., Perona, D., Ramanan, P., Dollar, C.-L., Zitnick, L.-C.: “Microsoft COCO: Common Objects in Context,” CoRR vol abs/1405.0312 2014 <ref:2502.20651#pg10>
[6] J. Deng, W. Dong, R. Socher, L.-J., Li, K., Li, and L.-F.: “ImageNet: A Large-scale Hierarchical Image Database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition 2009 pp 248–255 <ref:2502.20651#pg10>
[7] H. Lu, Y. Li, Y. Zhang, M. Chen, S., Serikawa, and H., Kim: “Underwater Optical Image Processing: A Comprehensive Review,” arXiv preprint arXiv:1702.03600 2017 <ref:2502.20651#pg10>
[8] C.-Y. Wang, I.-H., Yeh, and H.-Y. M.: “YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information,” 2024 <ref:2502.20651#pg10>
[9] N. Carion, F., Massa, G., Synnaeve, N., Usunier, A., Kirillov, A., and Zagoruyko, S.: “End-to-End Object Detection with Transformers,” 2020 <ref:2502.20651#pg10>
[10] K. He, G., Gkioxari, P., Dollar, and R. Girshick: “Mask R-CNN,” 2018 <ref:2502.20651#pg10>
[11] M. Everingham, L., Van Gool, C. K. I., Williams, J., Winn, and A.
Improvements for AI systems
-
The improved AI system can be trained to perform robust object detection and segmentation on Autonomous Underwater Vehicles (AUVs) across diverse aquatic environments by utilizing COU, which addresses
the lack of diversity in object classes since the commonly available underwater image datasets focus only on marine life.
-
The system will exhibit improved performance over models trained solely on terrestrial data because it mitigates domain shift by focusing
solely on underwater scenes,
leading toimproved generalization to underwater conditions.
-
The AI can achieve high-quality, pixel-level object boundaries by leveraging the dataset's focus on precise segmentation, as COU provides annotations that account for
the unique visual characteristics of underwater objects.
-
The system will be capable of recognizing a diverse set of man-made objects relevant to human-robot collaboration, including
marine debris, dive tools, and AUVs,
which expands beyond the limitations of existing datasets. -
The system can be optimized for real-time deployment on edge computing devices by utilizing models like YOLOv9-C, achieving inference times such as
150 ms/img
on the Jetson Orin NX, whichequates to roughly 6−7 FPS.
Sources
- Underwater Optical Image Processing: A Comprehensive Review
- YOLOv9: Learning What You Want to Learn Using Programmable Gradient Information
- End-to-End Object Detection with Transformers
- Mask R-CNN
- TrashCan: A Semantically-Segmented Dataset towards Visual Detection of Marine Debris
- Deep Residual Learning for Image Recognition
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models