Comparing Gaia, NED and SIMBAD source classifications in nearby galaxies

arXiv:2408.12717 · astro-ph.GA · Submitted 2024-08-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Astrophysics Radio. Generated commentary on the latest astrophysics papers.

Vera: Today's paper: "Comparing Gaia, NED and SIMBAD source classifications in nearby galaxies".

Jocelyn: Gaia Data Release 3 (DR3) provides a new standard for source classification,

Vera: First, who's behind it and why it matters.

Paper summary: Vera: So to wrap up what we've discussed, this paper titled "Comparing Gaia, NED and SIMBAD source classifications in nearby galaxies" is essentially testing how well the new Gaia DR3 classifications compare to those found in literature databases like NED and SIMBAD when we focus on sources within twice the Holmberg radius of nearby galaxies. The core thesis they are exploring is whether Gaia's classification system aligns well with these existing, more detailed, but heterogeneous literature classifications for sources in this specific local volume.

Jocelyn: They claim that by crossmatching these catalogues, which involves matching approximately three point two times one hundred five unique Gaia matches for four times one hundred five sources across one thousand forty galaxies in the Local Volume Galaxy catalogue, they are providing a new standard for source classifications based on the completeness and uniformity of the Gaia data <ref:2408.12717#pg0,galaxies in the Local Volume Galaxy catalogue>.

Subrahmanyan: The significance lies in using these matched catalogues to evaluate Gaia's performance, which is a key step in assessing the classification accuracy of Gaia itself across different types of sources present in these nearby galaxies. It sets up a direct comparison between an observational standard and established literature classifications for contextualizing the data.

Vera: And what matters from their findings is that they found that while Gaia's balanced accuracy isn't high when compared to those truth values, it still manages to perform well on classifying single stars as identified by both NED and SIMBAD, even though its performance varies significantly depending on the specific type of source being classified.

Jocelyn: That variation is what makes it interesting; they show that for certain sources, Gaia's accuracy holds up well against literature truth values, but then it drops when we look at background galaxies or quasars compared to the NED and SIMBAD classifications.

Subrahmanyan: This suggests that the utility of Gaia’s classification depends heavily on what we are trying to classify; it’s not a uniform performance across all source types, which is an important nuance for anyone planning future astrophysical analyses using these datasets.

Vera: And they also looked at sources with ambiguous classifications, like those labeled as star clusters or molecular clouds in literature, and the study found that Gaia's classifications for those things were primarily star clusters, H ii regions, and molecular clouds.

Jocelyn: That comparison of Gaia’s results against these literature labels helps us understand how the machine learning algorithms are interpreting sources that might be poorly defined in traditional catalogs. It gives us insight into the inherent biases within the classification modules themselves.

Subrahmanyan: So, this paper provides a comparative framework, using NED and SIMBAD as benchmarks to test Gaia's output, which is fundamentally important for establishing a reliable baseline for how we interpret source catalogs in our local cosmic neighborhood.

Conclusion: Vera: So to conclude this discussion on "Comparing Gaia, NED and SIMBAD source classifications in nearby galaxies," the authors Hales and Barmby have done a solid piece of work by comparing their new Gaia DR3 classifications against established literature from NED and SIMBAD for sources near us. The main implication is that while Gaia offers a new classification standard, we need to be realistic about where its accuracy is highest—it seems strongest when classifying single stars against literature truth values.

Jocelyn: I think the real significance is understanding the trade-offs; you get high accuracy on some source types but lower performance on others, like quasars or background galaxies. This tells us that no single classification system can perfectly describe every object in our local neighborhood.

Subrahmanyan: From a cosmic perspective, this work helps us calibrate our expectations when we use these classifications for larger structure studies; it’s about understanding the limitations inherent in the observational data and how those limitations affect our models of nearby galaxies.

Vera: Exactly, so we see that Gaia is a useful tool, but it requires careful application based on what kind of source we are interested in when working with this data set.

Jocelyn: It’s a good reminder that the accuracy isn't static across all source types and that literature databases still have their specific strengths in certain areas, which is important for using them together effectively.

Subrahmanyan: This comparative approach is valuable because it highlights how different observational methods contribute to the overall picture, helping us refine our understanding of what we observe in the local universe.

J. Hales, P. Barmby

Department of Physics & Astronomy, Western University

astro-ph.GA

Submitted: 2024-08-22

Updated: 2024-08-22

Comments: MNRAS in press; 12 pages, 7 figures

Journal ref: 2024, MNRAS, vol 533 p3415

DOI: 10.1093/mnras/stae2026

Code: https://github.com/mshubat/galaxy_data_mines

Project page: https://archives.esac.esa.int/gaia

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 74/100

The gist: Gaia Data Release 3 (DR3) provides a new standard for source classification, and this study compares its classifications against those from literature databases like NED and SIMBAD to understand how

Key concepts

Gaia Data Release 3 (DR3)
This is the new standard system used by Gaia to categorize astronomical objects. The study tests how well this new system aligns with existing classifications found in other major databases like NED and SIMBAD, which are used for cross-referencing astronomical sources.
Discrete Source Classifier (CU8-DSC)
This is the machine learning module within Gaia that assigns a class to each source. It uses algorithms like Allosmod and Specmod to determine the most probable category for a source based on its characteristics, assigning a class with over 50% probability.
Holmberg Radius
This defines the specific area around nearby galaxies that the researchers focused on studying. Sources located within this radius are retrieved from literature databases like NED and SIMBAD to compare them directly with Gaia's classifications.

Terminology

Summary

Gaia Data Release 3 (DR3) provides a new standard for source classification, and this study compares its classifications against those from literature databases like NED and SIMBAD to understand how these different classification schemes align for sources in nearby galaxies. The core finding is that while Gaia's balanced accuracy is relatively low when compared to literature truth values, it performs well on classifying single stars as identified by both NED and SIMBAD, and its performance metrics vary significantly depending on the source type being classified.

How it works

The researchers compare the Gaia classifications from the Discrete Source Classifier (CU8-DSC) module against the more detailed and heterogeneous classifications found in NED and/or SIMBAD for sources located within twice the Holmberg radius of nearby galaxies. The comparison involves matching approximately 3.2×105 unique Gaia matches for 4×105 sources across 1040 galaxies in the Local Volume Galaxy catalogue. The analysis treats NED and SIMBAD classifications as truth values to evaluate Gaia's performance, which is a key step in assessing the classification accuracy of Gaia.

Galaxy Sample and Source Retrieval

The study utilizes the June 2022 version of the Local Volume Galaxy catalogue (LVG) as its parent sample, which contains entries for 1421 galaxies within 11 Mpc of the Milky Way or with radial velocities less than 600 km s−1. For each galaxy, sources within the Holmberg radius are retrieved from NED and SIMBAD using the open-source command line tool galaxy data mines2. A tolerance of 5 arcseconds is adopted for crossmatching between NED and SIMBAD sources, a choice made to account for the likely worse astrometric precision and heterogeneity of literature data compared to Gaia's precision.

Classifying Gaia Sources

The classification of Gaia sources is performed using the Discrete Source Classifier (CU8-DSC) module, which employs machine learning algorithms including Allosmod, Specmod, and Combmod. Each source is assigned a class with the highest combined probability above 50 per cent. The results show that the sample classified as stars by classprob dsc combmod has "> 99 per cent completeness and purity, although this metric is noted as not particularly meaningful due to the dominance of stars in the sample. Conversely, classified galaxy and quasar samples have high completeness (0.94 and 0.92 respectively), they have rather low purity (0.22 and 0.24)."

Crossmatching NED and SIMBAD Sources to Gaia Sources

The researchers matched the combined list of NED-only, SIMBAD-only, and NED+SIMBAD sources to Gaia sources using the match to catalog sky method based solely on sky position. A conservative approach was taken by removing all Gaia sources that are matched to more than one NED and/or SIMBAD source from the comparison, reducing the number of involved sources to 272422. The coordinate offsets between NED and SIMBAD sources and their corresponding Gaia matches were plotted, showing mean offsets in both right ascension and declination of "< 0′′02."

Comparing Source Classifications between Databases

The performance metrics for Gaia classification against NED or SIMBAD classifications are summarized in Table 3. When NED classes are considered as truth values, the overall accuracy is 0.80, and the balanced accuracy is 0.45. Similarly, when SIMBAD classes are considered as truth values, the overall accuracy is 0.83 and the balanced accuracy is 0.47. The study highlights that "Gaia performs well (metrics > 0.8) on classifying single stars as identified by both NED and SIMBAD," while performance decreases for literature-identified background galaxies (purity 0.7–0.8, completeness 0.3–0.6).

Ambiguous and Unclassified Sources

The analysis also examined sources with ambiguous classifications or those unclassified by Gaia, such as NED/SIMBAD ‘unclassified’ categories like star clusters or molecular clouds. For the NED-unclassified sources, their Gaia classifications were primarily star clusters, H ii regions, and molecular clouds, with a distribution of 66 per cent star and 16 per cent galaxy. The study concludes that Gaia sources in the vicinity of nearby galaxies differ in their classification distribution from Gaia sources in the (Galactic) field.

Summary and Conclusions

The analysis demonstrates that Agreement between Gaia classification and literature classification found in NED and SIMBAD is best for stars, and decreases for quasars, (background) galaxies and white dwarfs, being lowest for binary stars. The work concludes that while broad Gaia classifications do not map particularly well onto the nature of sources near nearby galaxies, the small sample of galaxies and quasars in the purer Gaia sample does have higher classification metrics.

Improvements for AI systems

Based on a thorough review of this scientific paper, here are specific improvements that can be made to existing and future AI systems, along with what those improved systems could achieve:


) 1. Enhancing Source Classification Robustness via Heterogeneous Truth Values:

The paper demonstrates that Gaia classifications are highly dependent on the truth used for comparison (NED vs. SIMBAD). It shows performance varies significantly depending on whether NED or SIMBAD is treated as the ground truth, and that agreement is best for stars but decreases for quasars and galaxies.

  • The improved AI system could be a Hybrid Classification Engine trained not just on Gaia outputs, but explicitly on the discrepancies observed between literature databases (NED/SIMBAD) to improve its confidence scores.

  • This system could perform source classification by dynamically weighting the input based on known data quality or source type. For example, when classifying a potential galaxy candidate, the system would prioritize classifications where both NED and SIMBAD agree over those where they disagree, effectively mitigating the no ground truth problem by using consensus as a meta-truth.

  • This allows for a more nuanced output that can distinguish between sources with high confidence (where literature and Gaia align) and those with significant classification ambiguity.

) 2. Developing Adaptive Classification Thresholds Based on Source Context:

The paper reveals that the performance metrics (Purity/Completeness/Accuracy) change drastically depending on the class being predicted (e.g., excellent for single stars, poor for binaries).

  • An improved AI system could implement Context-Aware Thresholding. Instead of a single classification threshold (like the 50% probability used in Section 2.2), the system would dynamically adjust this threshold based on metadata such as:

  • Source distance (to account for background vs. foreground contamination).

  • The source's literature classification confidence score (if available).

  • The presence of specific features like ambiguous classifications mentioned in Section 3.4.

  • This would allow the system to be highly sensitive when dealing with rare, high-value classes (like white dwarfs or binary stars) and more lenient when classifying complex, unclassified components like molecular clouds or H II regions where literature is inherently ambiguous.

) 3. Creating Specialized Models for Literature-Ambiguous Classes:

The analysis highlights that sources classified as unclassified by Gaia are often star clusters, H II regions, or molecular clouds—categories poorly captured by the standard five classes.

  • The improved system should incorporate a Literature Gap Filler Module. This module would be specifically trained to map ambiguous literature classifications (like 'infrared source' or 'HII') to the broader physical categories (star cluster, nebula) suggested by the Gaia confusion matrices (Table 4).

  • This allows the AI to move beyond rigid categorical output and provide a probabilistic range of physical interpretations, explicitly flagging sources where its classification relies on inference from literature rather than direct Gaia feature mapping.

) 4. Optimizing Cross-Matching for Literature Uncertainty:

The paper notes that astrometric uncertainties in NED/SIMBAD data dominate the matching success, leading to multiple matches for some Gaia sources.

  • An improved AI system should utilize a Probabilistic Match Prior during the cross-matching phase. Instead of relying solely on hard positional constraints (like the 5 arcsec tolerance used here), it should incorporate Bayesian priors based on the known uncertainty distributions of NED and SIMBAD coordinates.

  • This would allow the system to intelligently resolve multiple matches by calculating a weighted likelihood that accounts for each potential match's associated astrometric error, leading to a more accurate one-to-one source assignment, especially in crowded fields where literature data is inherently noisy.

) What the Improved AI System Can Do:

The resulting improved system would be a sophisticated astronomical classification pipeline capable of:

  1. Accurately classifying astronomical sources (stars, galaxies, quasars) in nearby galaxies with a validated accuracy of 0.80–0.83 when using NED/SIMBAD as truth values.

  2. Provide confidence scores for every classification, distinguishing between high-certainty classifications and those derived from ambiguous literature data.

  3. Perform source identification that is robust against the inherent noise and heterogeneity of legacy astronomical databases (NED/SIMBAD).

  4. Identify sources where the literature classification itself is ambiguous, suggesting a need for further multi-wavelength observational follow-up or specialized physical modeling, thereby guiding future scientific research efforts precisely where they are most needed.

Related papers