SomBench: Benchmark Dataset for Advancing Machine Learning in Lunar Science

arXiv:2609.13277 · cs.CV, cs.LG · Submitted 2026-09-08 · Read on arXiv

cs.CV, cs.LG

Submitted: 2026-09-08

Updated: 2026-09-08

License: http://creativecommons.org/licenses/by/4.0/

The gist: Lunar orbital missions, such as Lunar Reconnaissance Orbiter, Kaguya/SELENE, Gravity Recovery and Interior Laboratory, and Lunar Prospector, among others, provide rich multi-instrument observations,

Terminology

Abstract

Lunar orbital missions, such as Lunar Reconnaissance Orbiter, Kaguya/SELENE, Gravity Recovery and Interior Laboratory, and Lunar Prospector, among others, provide rich multi-instrument observations, but their heterogeneity in sampling, projection, and conventions limits reproducible machine learning (ML). We introduce SomBench, a unified, spatially-aligned, ML-ready lunar dataset aggregating 30+ co-registered layers from ten instruments across four missions, spanning 1 meter to 20 kilometer/pixel and covering 82 degree latitude in 90 Lunar Transverse Mercator zones with two polar stereographic caps. An image-anchored tiling pipeline yields pretraining-ready multimodal tile views with leakage-safe splits, distributed as netCDF with Parquet catalogs. An application benchmark suite spans impact processes, volcanic history, and polar volatiles. Baseline experiments with ResNet-50 and SwinV2-B models confirm that each benchmark task is learnable from the released inputs, establishing reference points for future model development.

Related papers