JW-FD: A 15-Year Multimodal Dataset for Solar Flare Forecasting

summary

Video file (mp4)

The gist

Solar flares pose severe space weather hazards, and reliable forecasting remains a central challenge for heliophysics and operational space weather services.

In short

JW-FD is a comprehensive, 15-year dataset (2011–2025) for solar flare forecasting, covering Solar Cycles 24 through 25. It links images, magnetic features, and flare labels across multiple modalities like FITS and PNG. This resource allows researchers to train advanced machine learning models for predicting space weather events using long-term data.

Key concepts

Multimodal Coregistration
This means combining different types of data—like physical images (FITS), visual pictures (PNG), numerical features (CSV), and video sequences (MP4)—so they all refer to the exact same solar active region at the same time. This helps models learn richer patterns from various data sources simultaneously.
Flare Labeling Scheme
The dataset uses a specific binary label called 'flare_label_τ_hhr' to mark when a flare occurs. A record is labeled '1' if any flare with a GOES intensity higher than a set threshold occurs within the specified time window relative to the AR's start time, and '0' otherwise.
AR Level Partitioning
Instead of splitting data by calendar dates, JW-FD splits it based on individual Active Regions (ARs). This is done to prevent temporal leakage, ensuring that all data points belonging to one specific solar region are kept together in either the training or testing set.
Multimodal Fusion
This refers to the process of merging information from different data types. For example, researchers can combine visual inputs (PNG/MP4) with magnetic measurements and flare labels from the same record to build more accurate forecasting systems.

Terminology used across episodes

This episode discusses

The paper

JW-FD: A 15-Year Multimodal Dataset for Solar Flare Forecasting · Read on arXiv

Shao Mingfu, Lin Jiaben, *Wang Hui*, Tong Liyue, *Yang Chen*, *Zhang Yin*, *Li Yuyang*

State Key Laboratory of Solar Activity and Space Weather, NAOC, Beijing 100101, P. R. China · University of Chinese Academy of Sciences

Solar flares drive severe space weather hazards, and forecasting their occurrence remains a central challenge for both heliophysics and operational space weather services. Data driven methods require long horizon datasets in which images, magnetic features, and flare labels are coregistered in space and time. We present JW-FD (JW-Flare Dataset), a 15 year multimodal release spanning 1 January 2011 through 31 December 2025, constructed from SDO/HMI line of sight magnetograms, NOAA Solar Region Summary reports, and NOAA X-ray flare event lists. The dataset comprises 3, 064 independent active regions and 1, 991, 247 coregistered magnetogram crops, together with FITS, PNG, CSV, and MP4 modalities. Each sample provides 29 magnetic features linked to configurable flare labels under a strict pre-eruption window spanning seven forecast horizons and four GOES intensity thresholds. An 8:1:1 split at the active region level is adopted to prevent temporal leakage between partitions. PNG branches are released at six magnetic saturation thresholds, and internal Transformer experiments on at least C1.0 forecasting suggest B th=1000 G as a preliminary default, although the optimal saturation is model and task dependent. The open source construction pipeline is available at https://github.com/Xiaoxuan-1/JW-FD.

DOI: 10.3390/universe12090281

Transcript

Introduction to the show: ident: Astrophysics Radio. Generated commentary on the latest astrophysics papers.

Vera: I'm Vera, and with me are Jocelyn and Subrahmanyan, guest researcher.

Jocelyn: Today's paper: "JW-FD: A 15-Year Multimodal Dataset for Solar Flare Forecasting".

Vera: Solar flares pose severe space weather hazards, and reliable forecasting remains a central challenge for heliophysics and operational space weather services.

Jocelyn: First, who's behind it and why it matters.

Title and authors: Vera: So, we started by looking at the title and authors of the paper JW-FD: A fifteen-Year Multimodal Dataset for Solar Flare Forecasting. It's important to understand that this isn't just another data dump; it's a structured collection built by Shao Mingfu and his team to tackle a major challenge in space weather forecasting.

Jocelyn: I agree, Vera; the title itself tells us immediately that the authors are focused on creating something multimodal and long-horizon, which is exactly what we need when dealing with solar flares that happen over extended periods.

Subrahmanyan: Theoretically, the authors’ focus on a fifteen-year span suggests they are trying to capture phenomena across multiple solar cycles, which is vital for understanding the longer-term magnetic storage and release mechanisms in the Sun.

Vera: That’s right; they aren't just looking at one short event or a single solar cycle; they want to provide enough data depth so that models can see patterns that span those long timescales.

Jocelyn: The authors are clearly aiming for a shared infrastructure, meaning this isn't just their private collection; it’s intended to be something the broader heliophysics and space weather community can actually use for their own research.

Subrahmanyan: If they succeed in creating a reusable corpus, it could significantly lower the barrier for researchers who want to apply advanced machine learning techniques to solar phenomena without needing proprietary access to massive datasets.

Vera: That’s the main goal of making it accessible; by providing this coregistered data set, they are setting up a standard for how multimodal flare forecasting research should be conducted in this field.

Jocelyn: And looking at the authors, they seem to have a solid background spanning observational astronomy and data science, which is what you need when you're building something so complex as JW-FD: A fifteen-Year Multimodal Dataset for Solar Flare Forecasting.

Subrahmanyan: Their combined expertise in connecting the physics of solar activity with modern data analysis methods is exactly what makes this paper important to the community.

Vera: So, in short, this paper introduces a new way to structure long-term solar flare forecasting data that’s accessible to everyone interested in applying modern computational methods.

Jocelyn: It's about creating a foundation for future work rather than just reporting on one specific finding; it’s about building something lasting.

Subrahmanyan: And from a theoretical standpoint, it provides the necessary empirical backbone to test hypotheses about solar magnetic evolution over extended timescales.

The paper's summary: Vera: Now we get into the paper's summary, and it really lays out exactly what JW-FD: A fifteen-Year Multimodal Dataset for Solar Flare Forecasting is all about—it’s a comprehensive, fifteen-year release spanning from January two thousand eleven through December two thousand twenty-five.

Jocelyn: The summary emphasizes that the key challenge they are addressing is the need for data where images, magnetic features, and flare labels are all coregistered in space and time under a leakage-aware partitioning scheme.

Subrahmanyan: This coregistration requirement is what separates it from previous datasets; we're not just getting one type of data; we're getting the physical context alongside the observational snapshots.

Vera: Exactly, and they detail the specific modalities included, including FITS for physics data, PNG for images, CSV for structured feature vectors, and MP4 for evolution videos.

Jocelyn: And they stress that this multimodal release covers solar cycles twenty-four through twenty-five and includes a total of three thousand sixty-four independent active regions with a total of one million coregistered magnetogram crops.

Subrahmanyan: Having that specific coverage across those cycles means the dataset is designed to capture the full range of magnetic activity we expect over a substantial period.

Vera: It’s also important to note that each record links twenty-nine magnetic features—categorized into gradient, neutral line, wavelet, and flux descriptors—to flare labels defined based on seven forecast horizons and four GOES intensity thresholds.

Jocelyn: That detailed labeling structure is key because it allows for flexible evaluation based on various risk metrics, which is a major strength over simpler datasets.

Subrahmanyan: The methodology described in the summary indicates that they used a six-step pipeline to build this dataset from NOAA Events, SRS reports, and HMI data to ensure the inputs are sourced from reliable scientific instruments.

Vera: It's about taking raw observational data and applying specific processing steps—like cropping active regions at six hundred times six hundred pixels and converting FITS to PNG using a magnetic saturation threshold Bth.

Jocelyn: And they also created time-compressed MP4 evolution videos for each active region, which is a big addition for any research involving temporal dynamics.

Subrahmanyan: So, the summary paints a picture of a dataset that is meticulously constructed to serve as shared infrastructure for classical ML, deep learning, and multimodal flare forecasting research.

Vera: It’s designed to be the central resource where different types of data can be combined and analyzed together in a way that was previously difficult to achieve.

Jocelyn: And the goal is clearly to provide a rich resource covering solar cycles twenty-four through twenty-five so researchers can really test their predictive capabilities against realistic, long-term solar activity.

Subrahmanyan: It’s about providing the necessary empirical backbone for testing models that try to predict complex, time-dependent events in space weather.

The paper's improvements: Vera: Now let's discuss the improvements they suggest for this dataset, because it’s not just about what they have, but how they plan to make it even more effective for researchers. It seems the authors are pushing us to think about advanced applications.

Jocelyn: I see a few key areas where they suggest enhancing the utility of this data: specifically improving solar flare forecasting models using vision and video models.

Subrahmanyan: They suggest training Convolutional Neural Networks or Vision Transformers directly on those six hundred times six hundred pixel PNG crops to identify spatial precursors of flares, which lets models learn visual signatures in magnetic field morphology that precede eruptions.

Vera: That moves the focus from just using features to actually analyzing the visual structure, which is a significant step forward for identifying subtle morphological changes before an event occurs.

Jocelyn: They also propose utilizing those sixty-six-dimensional CSV records for parameter-based classification, suggesting classical classifiers or Gradient Boosted Trees like XGBoost to compare their predictive power against the visual inputs.

Subrahmanyan: That comparison is important; it allows researchers to see if the physical descriptors alone can predict flares as effectively as a model that looks at the actual magnetic structure in an image.

Vera: Furthermore, they suggest training Long Short-Term Memory networks on the time-series data derived from those sixty-six magnetic features to capture temporal dependencies and forecast probability or intensity over various horizons like one hour to seventy-two hours.

Jocelyn: That’s interesting; integrating temporal modeling with both visual and feature data seems like a powerful way to capture the full picture of what's happening in an active region over time.

Subrahmanyan: When we combine that temporal aspect with the multimodal fusion idea, it enables AI systems to learn spatio-temporal patterns—how magnetic features evolve over time leading up to an eruption, which is crucial for understanding flare dynamics missed by static snapshots alone.

Vera: And they also point out that the system can be designed to handle those configurable labels across different risk levels, allowing operational services to select the most relevant forecast based on their specific hazard tolerance.

Jocelyn: That threshold-aware prediction capability seems very practical for real-world applications because it lets you tailor the output to your specific needs rather than just one fixed output.

Subrahmanyan: The rigorous AR level split, which they employ, helps ensure model generalization across different solar cycle environments by allowing us to test how well a model trained on one set of conditions performs on another.

Vera: And that strict partitioning is also key for ensuring leakage-aware training, meaning we aren't accidentally contaminating the training data with information from an active region that isn't actually in the test set.

Jocelyn: So, to summarize these improvements are centered around building models that can handle visual inputs, physical parameters, and temporal dynamics simultaneously for better flare forecasting.

Subrahmanyan: And this holistic approach is what really allows us to connect the observational data to the underlying physics of solar flares in a more comprehensive manner than any single modality could achieve alone.

Conclusion: Vera: So, as we wrap up our discussion on JW-FD: A fifteen-Year Multimodal Dataset for Solar Flare Forecasting, it seems the paper concludes that this dataset is a significant step forward in providing the necessary long horizon data for space weather forecasting.

Jocelyn: It really does establish a new benchmark by successfully coregistering different types of data—images, features, and labels—for research purposes.

Subrahmanyan: The ultimate implication is that it gives researchers a standardized way to test their theories about the physics of solar flares using real-world, long-term observational evidence.

Vera: For me, I think it’s exciting because we now have a foundation to train more advanced AI systems that can actually look into the physical precursors of flares with high fidelity.

Jocelyn: I'm looking forward to seeing how these models are applied, especially with the multimodal fusion capabilities they’ve put into this dataset: JW-FD: A fifteen-Year Multimodal Dataset for Solar Flare Forecasting.

Subrahmanyan: From my side, it confirms that the long-horizon data is necessary to connect our theoretical predictions about solar magnetic structures to what we actually observe.

Vera: We’re definitely ready to look at the next paper when it comes, but for now, this JW-FD: A fifteen-Year Multimodal Dataset for Solar Flare Forecasting has given us a lot of concrete material to work with.

More episodes

← Home