A General-Purpose Diversified 2D Seismic Image Dataset from NAMSS

arXiv:2602.04890 · physics.geo-ph, cs.AI, cs.CV, cs.LG · Submitted 2026-01-27 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "A General-Purpose Diversified 2D Seismic Image Dataset from NAMSS".

Jane: We introduce Unicamp-NAMSS, a large, diverse, and geographically distributed collection of migrated 2D seismic sections from the National Archive of Marine Seismic Surveys (NAMSS),

Tom: First, who's behind it and why it matters.

Paper summary: Tom: So, wrapping up our discussion on "A General-Purpose Diversified 2D Seismic Image Dataset from NAMSS," the main point is that this collection offers a large, varied foundation for modern machine learning research in geophysics <ref:2602.04890#pg1>.

Jane: The authors of this paper have clearly aimed to create something that covers many acquisition conditions and geological settings, which is exactly what makes it valuable for training models <ref:2602.04890#pg1>.

Lu: They successfully demonstrated that the data has substantial variability both within and across regions, suggesting that geological setting and survey geometry are the primary sources of structural differences <ref:2602.04890#pg1>.

Meng: From an engineering view, this means we have a massive, pre-processed dataset that is ready to be used for training foundation models without having to spend years curating a comparable resource from scratch <ref:2602.04890#pg1>.

Lalam: I think the implication is that this kind of standardized, diverse data could really encourage more cross-disciplinary collaboration between geophysicists and AI engineers to tackle complex subsurface problems <ref:2602.04890#pg1>.

Tom: It’s about providing a resource that covers a broader portion of the seismic appearance space than many existing benchmarks, making it a strong tool for model pretraining <ref:2602.04890#pg1>.

Jane: Overall, the Unicamp-NAMSS dataset represents a significant step toward enabling more robust and generalizable AI applications in the geophysical domain <ref:2602.04890#pg1>.

Conclusion: Tom: So, we've been diving deep into how they put together this massive seismic collection from NAMSS, and now it's time to talk about what that title actually means for us in real life.

Jane: I think the authors really nailed it with that title because they aren't just handing over a bunch of old files; they are giving us a set of tools for everyone who wants to do AI work in this field.

Lu: Exactly, and when you break down "general-purpose diversified," it tells us this isn't some narrow dataset for one specific type of geology; it’s broad enough that models trained on it should actually perform well on totally new problems.

Meng: From an engineering standpoint, what I see is the sheer scale of this diversity, which means we can test robustness in ways we just couldn't before without collecting entirely new data.

Lalam: And from my perspective as a large language model, this collection provides a really rich semantic space for understanding subsurface structures across different scales and conditions, which could fundamentally improve how we learn to interpret geophysical signals.

Tom: It really comes down to the fact that this dataset covers so many different acquisition methods and geological settings that it sets a new baseline for what a strong foundation model needs.

Jane: And I think the authors’ focus on balancing the distribution across training, validation, and test sets makes it incredibly useful for rigorously evaluating how well these models actually generalize in practice.

Lu: The implication here is that we can move past training models that only work in one specific scenario and start building systems that handle real-world complexity much better.

Meng: If we can get a model to learn from this breadth, the practical impact is huge for things like autonomous subsurface exploration where you don't know what to expect.

Lalam: It could really help shift the culture around geophysical modeling by making sophisticated AI techniques accessible and applicable across a wider range of scientific questions.

Tom: So, this collection isn't just data; it’s a new kind of standardized laboratory for testing the limits of our current geophysical AI methods.

Jane: And we'll be looking at how researchers start using this to build more reliable tools for understanding the Earth beneath our feet next.

Instituto de Computação, Universidade Estadual de Campinas (Unicamp)

physics.geo-ph, cs.AI, cs.CV, cs.LG

Submitted: 2026-01-27

Updated: 2026-01-27

Journal ref: IEEE Data Descriptions, vol. 3, pp. 595-604, 2026

DOI: 10.1109/IEEEDATA.2026.3713847

Code: https://github.com/discovery-unicamp/namss-dataset

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 79/100

The gist: We introduce Unicamp-NAMSS, a large, diverse, and geographically distributed collection of migrated 2D seismic sections from the National Archive of Marine Seismic Surveys (NAMSS), designed to

Key concepts

Unicamp-NAMSS
A large, diverse collection of 2D seismic sections migrated from the National Archive of Marine Seismic Surveys (NAMSS). It was constructed by selecting and balancing surveys to ensure each contributed similar data, providing broad variability for machine learning.
Data Preprocessing
Rigorous steps taken on raw SEG-Y files to create a clean, uniform image dataset. This included reading samples into matrices, amplitude normalization to the [-1, 1] interval by dividing by the global maximum, and saving the final data in TIFF format for processing.
Macro-regions
The 122 survey areas were manually grouped into nine non-overlapping macro-regions. These regions were then used to create training, validation, and test splits to robustly evaluate how well models generalize across different geological and acquisition conditions.
Embedding-space Analysis
Using deep learning models like ResNet-50 and ViT-B/14 to project the seismic data into a learned embedding space using UMAP. This analysis confirmed that structural variability is primarily driven by geological setting, survey geometry, and processing workflows.

Terminology

Summary

We introduce Unicamp-NAMSS, a large, diverse, and geographically distributed collection of migrated 2D seismic sections from the National Archive of Marine Seismic Surveys (NAMSS), designed to support modern machine learning research in geophysics by providing substantial variability across acquisition conditions and geological settings.

The gist

Unicamp-NAMSS is a large, diverse, and geographically distributed collection of migrated 2D seismic sections designed to support modern machine learning research in geophysics.

Dataset Construction and Collection Process

The dataset was constructed from the National Archive of Marine Seismic Surveys (NAMSS), which contains decades of publicly available marine seismic data acquired across multiple regions, acquisition conditions, and geological settings. The initial search utilized the NAMSS map-based tool to locate surveys with 2D multichannel seismic data, leading to a list of 547 distinct 2D multichannel surveys. From these, 143 contained at least one migrated dataset. By matching filename patterns with regular expressions, the researchers identified 9,350 migrated files, totaling 157 GB. To ensure reliability and prevent bias from large volumes, the researchers balanced the dataset so that each survey contributed a similar amount of data. This was achieved by selecting a maximum threshold of 300 MB per survey; surveys exceeding this limit were randomly undersampled following the strategy by He and Garcia [15]. The final collection resulted in 2,588 seismic images time-migrated seismic images with 4 ms sampling interval.

Data Preprocessing and Standardization

The raw SEG-Y files underwent several rigorous preprocessing steps to create a clean, uniform image-based dataset. These steps included:

  1. Reading: Seismic samples were extracted from each trace and converted from IBM hexadecimal floating-point format to IEEE 754 double-precision when necessary. The traces were assembled into matrices of size (number of samples per trace) × (number of traces), keeping the original order without spatial reordering.

  2. Normalization: Each matrix was amplitude-normalized to the interval [−1, 1] by dividing all samples by the global absolute maximum.

  3. Writing: The normalized data were saved in TIFF format, which preserves 4-byte floating-point precision and is convenient for pretraining and downstream processing tasks.

During cleaning, specific quality control measures were implemented. Duplicate exploration areas were removed, such as when one survey was fully contained within another. To maintain uniform vertical sampling across the dataset, only files with a sampling interval of 4 ms—which represented 91 % of all collected files—were retained and others were removed. Furthermore, 59 corrupted or unreadable files (originating from 7 surveys) were removed, and identical files across multiple surveys were discarded.

Dataset Partitioning and Distribution

To ensure robust evaluation of generalization to unseen geological and acquisition conditions, the dataset was partitioned into non-overlapping macro-regions. The researchers manually grouped the 122 survey areas into nine macro-regions, avoiding spatial redundancy between them. These macro-regions were then assigned to training, validation, and test splits as follows:

Training:

Gulf of Alaska, North Atlantic Ocean, Arctic Ocean, Southern California.

Validation:

Caribbean Sea, Northern California.

Test:

Argentine Basin, Bering Sea, Washington-Oregon.

The split was manually adjusted to achieve approximately 80 % training, 10 % validation, and 10 % testing of the total dataset size, ensuring that no survey dominates the distribution. The resulting structure allows for a robust evaluation of generalization to unseen geological and acquisition conditions.

Technical Validation and Variability Analysis

The quality, diversity, and representativeness of Unicamp-NAMSS were validated through embedding-space analysis using both convolutional (ResNet-50) and transformer-based (ViT-B/14) models. The data was projected into a learned embedding space using UMAP to visualize structural and textural variability. These analyses showed that Unicamp-NAMSS exhibits substantial variability within and across regions, while maintaining coherent structure across acquisition macro-region and survey types. Specifically, the DINOv2 embeddings demonstrated a clear separation for certain macro-regions like the Bering Sea and Arctic Ocean, suggesting that geological setting, survey geometry, and processing workflows appear to be the primary sources of structural variability.

Comparison with Existing Benchmarks

Comparisons with widely used interpretation datasets (Parihaka and F3 Block) demonstrated that Unicamp-NAMSS covers a broader portion of the seismic appearance space, making it a strong candidate for machine learning model pretraining.

Improvements for AI systems

As a fastidious and diligent researcher, I have analyzed the Unicamp-NAMSS dataset paper, focusing on its strengths: large scale (2,588 images), high diversity (geographic regions, acquisition conditions), and balanced partitioning.

Here are the specific improvements to AI systems that can be made using this dataset:


  1. Enhance Generalization in Seismic Foundation Models (Self-Supervised Learning/Pretraining)

  2. Improve Domain Adaptation for Super-Resolution and Denoising Tasks

  3. Develop Robust Semantic Segmentation and Attribute Prediction Models

  4. The improved system can function as a Seismic Vision Foundation Model. By pretraining models (like ResNet-50 or ViT-B/14) on the diverse Unicamp-NAMSS embeddings, the model will learn generalized, low-level structural and textural features across vastly different acquisition geometries (trace spacing from 4m to 305m) and geological settings.

  5. This foundation model can then be efficiently fine-tuned for specific tasks. For instance, a self-supervised pretraining approach allows the model to learn robust representations without requiring expensive human labels, addressing the scarcity of labeled seismic data in specialized domains (e.g., deep subsurface geology).

  6. The improved system can perform state-of-the-art image restoration and enhancement tasks.

  7. For Super-Resolution: The model can be trained to reconstruct higher-resolution seismic sections from the 4ms sampling intervals by leveraging its learned knowledge of clean structures, effectively upscaling the data while preserving geological features rather than introducing artifacts common in traditional interpolation methods.

  8. For Denoising: The model can be fine-tuned on noisy versions of Unicamp-NAMSS data to learn to distinguish between acquisition noise/processing artifacts and true subsurface signal, leading to significantly cleaner seismic images suitable for interpretation.

  9. The improved system can perform high-accuracy geological feature extraction and classification.

  10. Attribute Prediction: The model can be trained (or fine-tuned) on the dataset to predict specific geological attributes (e.g., lithology, fault presence, fluid saturation proxies) directly from the seismic image without needing extensive manual labeling for every new target attribute.

  11. Semantic Segmentation: By leveraging the diversity of macro-regions and acquisition parameters, the system can be trained to perform pixel-level segmentation of subsurface structures (e.g., sedimentary layers or geological boundaries), providing a more flexible tool than models trained on single-domain data.

  12. Model Selection and Robustness Benchmarking

  13. The dataset's clear separation between its embeddings and those from localized datasets (Parihaka, F3) allows the system to be used for rigorous benchmarking. It can evaluate whether a model is truly learning general geophysical principles or merely memorizing patterns specific to a narrow acquisition domain, ensuring higher confidence in its deployment readiness.

Abstract

We introduce the Unicamp-NAMSS dataset, a large, diverse, and geographically distributed collection of migrated 2D seismic sections designed to support modern machine learning research in geophysics. We constructed the dataset from the National Archive of Marine Seismic Surveys (NAMSS), which contains decades of publicly available marine seismic data acquired across multiple regions, acquisition conditions, and geological settings. After a comprehensive collection and filtering process, we obtained 2588 cleaned and standardized seismic sections from 122 survey areas, covering a wide range of vertical and horizontal sampling characteristics. To ensure reliable experimentation, we balanced the dataset so that no survey dominates the distribution, and partitioned it into non-overlapping macro-regions for training, validation, and testing. This region-disjoint split allows robust evaluation of generalization to unseen geological and acquisition conditions. We validated the dataset through quantitative and embedding-space analyses using both convolutional and transformer-based models. These analyses showed that Unicamp-NAMSS exhibits substantial variability within and across regions, while maintaining coherent structure across acquisition macro-region and survey types. Comparisons with widely used interpretation datasets (Parihaka and F3 Block) further demonstrated that Unicamp-NAMSS covers a broader portion of the seismic appearance space, making it a strong candidate for machine learning model pretraining. The dataset, therefore, provides a valuable resource for machine learning tasks, including self-supervised representation learning, transfer learning, benchmarking supervised tasks such as super-resolution or attribute prediction, and studying domain adaptation in seismic interpretation.

Sources

Related papers