When X-ray Features Fail to Identify Intrinsic Emitters: Label Noise and Luminosity Overlap in Machine Learning Classification of AGN and Star-forming Galaxies
Jaymin Ding
astro-ph.GA, astro-ph.IM
Submitted: 2026-07-14
Comments: 7 pages, 4 figures. Submitted to RASTI
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: In a previous paper, we found that adding an X-ray flux feature to a Random Forest classifier of active galactic nuclei (AGN) and star-forming galaxies (SFGs) coincided with a decrease in
Terminology
Abstract
In a previous paper, we found that adding an X-ray flux feature to a Random Forest classifier of active galactic nuclei (AGN) and star-forming galaxies (SFGs) coincided with a decrease in classification accuracy from 97.51% to 89.26%, a counterintuitive result given the prevailing theory that X-rays are a reliable AGN diagnostic. This paper investigates the source of that discrepancy through both astrophysical and machine learning lenses. On the astrophysical side, we show that the X-ray luminosities of AGN and SFGs in the sample substantially overlap across the canonical 10 42 erg/s threshold, reflecting the moderate luminosity characteristic of an optically-selected SDSS sample and the contribution of high-mass X-ray binaries (HMXBs) to SFG emission. On the machine learning side, we argue that the BPT-derived training labels constitute instance-dependent label noise: label uncertainty is concentrated near the BPT demarcation line, where X-ray data would be most useful as a discriminator. We show that misclassified objects in 5-fold cross-validation cluster primarily near the Kewley demarcation curve, with a median distance of 0.123 dex compared to 0.743 dex for correctly classified objects (Kolmogorov-Smirnov D = 0.745,; p = 1.67 times 10-9). A controlled cross-validation decomposition shows that the previously reported decrease reflects predominantly sample selection rather than the X-ray feature, which exhibits a small but significantly negative permutation importance: the model makes limited, net-detrimental use of it, though its effect on overall accuracy is negligible. We conclude that future classifiers operating on optically-selected samples should employ independent label sources or noise-robust training methods.
Sources
Related papers
- Apparent Stability in Self-Gravitating Turbulence and the Evolution of Molecular Clouds
- Two sets of potential-density basis pairs for the study of radial perturbations in collisionless spherical stellar systems
- Constraining reionization-era Ly alpha escape with JELS-MUSE: a highly complete H alpha-selected sample at z about6.1
- Deriving volume density profiles of filaments from observed surface densities
- Little Red Dots and Supermassive Black Hole Seed Formation in Ultralight Dark Matter Halos
- MEGATRON: how the first stars can create an iron metallicity plateau in the smallest dwarf galaxies