Estimating the local star formation rate density from ASKAP RACS

arXiv:2608.10347 · astro-ph.GA · Submitted 2026-08-11 · Read on arXiv

Sonja Panjkov, O. Ivy Wong., Rachel L. Webster

University of Melbourne · Commonwealth Scientific and Industrial Research Organisation · University of Western Australia

astro-ph.GA

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 17 pages, 11 figures, published in PASA

Journal ref: Publications fo the Astronomical Society of Australia 43, e094, 1-17 (2026)

DOI: 10.1017/pasa.2026.10228

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper presents a novel method to estimate the local star formation rate density (SFRD) using supervised machine learning to identify star-forming galaxies (SFGs) from the Australian Square

Terminology

Summary

This paper presents a novel method to estimate the local star formation rate density (SFRD) using supervised machine learning to identify star-forming galaxies (SFGs) from the Australian Square Kilometre Array Pathfinder (ASKAP) Rapid ASKAP Continuum Survey (RACS). The authors implemented a gradient-boosted decision tree model (XGBoost) to classify extragalactic sources from the Beck et al. (2022) catalogue as either galaxies or quasars, using 105 features from the RACS-mid and WISE All-Sky catalogues. The final cross-matched sample consisted of 389,392 sources, split into 70%-15%-15% for training, validation, and testing. After trialling various resampling techniques, random oversampling of the quasar minority class was found to be optimal. The optimised model achieved a weighted F1 score of 0.93 and an accuracy of 0.94 on the test dataset, ultimately classifying 336,674 sources as galaxies and 52,718 sources as quasars.

Using the resulting galaxy sample, the authors determined star-formation rates (SFRs) from 1.4-GHz radio-continuum emission. They depth-matched the radio and infrared catalogues to correct for different sensitivities, following the methodology of Molnár et al. (2021). This process reduced the model-selected galaxy sample from 213,490 sources (after a signal-to-noise cut) to 13,325 sources, with 11,293 sources in the z < 0.1 redshift range. A modified 1.4-GHz SFR calibration was determined from the best-fit parameters of the infrared-radio correlation (IRRC) for the z < 0.1 depth-matched sample, yielding the prescription: log(SFR) = (0.745 ± 0.088) log(L1.4 GHz) + (−15.8 ± 1.9).

The local SFRD was calculated using the 1/Vmax method to correct for observational bias, and corrected for completeness using a luminosity function comparison with Matthews et al. (2021). The completeness-corrected, z < 0.1 SFRD was determined to be (1.4 ± 0.5) × 10−2 M⊙ yr−1 Mpc−3 using 11,293 sources. This value is consistent with previous results from Mauch & Sadler (2007), Upjohn et al. (2019), and the analytical fits of Hopkins & Beacom (2006) and Behroozi et al. (2013). The uncorrected value was (1.6 ± 0.6) × 10−3 M⊙ yr−1 Mpc−3. The authors also calculated SFRDs for higher redshift bins using the Molnár et al. (2021) prescription, but these were significantly underestimated due to the shallowness of RACS-mid beyond z ∼ 0.1.

The study demonstrates the feasibility of using supervised learning to identify large populations of SFGs to investigate SFRD evolution. The authors note that the model struggles with high-redshift sources (z > 0.5), and that the AGN contamination fraction of the galaxy sample was estimated at 0.05 across all redshifts (0.03 for z < 0.1) via cross-matching with the Milliquas catalogue. The paper concludes that future deeper surveys such as EMU will enable the cosmic SFRD to be probed out to higher redshifts with improved precision.

Improvements for AI systems

Improvements to AI Systems:

  1. Enhanced Class Imbalance Handling for Rare-Class Detection: The finding that random oversampling of the minority class (quasars) outperformed other resampling techniques (e.g., SMOTE, undersampling) provides a concrete, validated strategy for AI systems dealing with highly imbalanced astronomical or scientific datasets. An improved AI system can automatically implement random oversampling as a default baseline for binary classification tasks where the minority class represents <15% of the data, reducing the need for exhaustive trial-and-error.

  2. Feature Engineering from Multi-Wavelength Catalogues: The use of 105 features combining radio (RACS-mid) and infrared (WISE) data demonstrates that AI systems can be improved by integrating heterogeneous, multi-wavelength photometric features rather than relying on single-band data. An improved AI system can be designed to automatically generate and select cross-matched features from multiple survey catalogues, improving classification accuracy for extragalactic sources.

  3. Depth-Matching Correction for Sensitivity Bias: The paper’s methodology of depth-matching radio and infrared catalogues before applying AI-based SFR estimation corrects for selection effects. An improved AI system can incorporate a pre-processing module that aligns catalogue sensitivities (e.g., via flux-density limits) before training or inference, preventing systematic underestimation of star formation rates in shallower surveys.

  4. Calibration of Physical Prescriptions via IRRC Fitting: The derivation of a modified 1.4-GHz SFR calibration (log(SFR) = 0.745 log(L1.4) − 15.8) from the infrared-radio correlation provides a new, AI-validated empirical relation. An improved AI system can use this calibration as a prior or transfer-learning target for estimating SFRs in future radio surveys, reducing reliance on outdated or less accurate prescriptions.

  5. Redshift-Aware Uncertainty Quantification: The model’s known degradation at z > 0.5 and the significant underestimation of SFRD at higher redshifts highlight the need for AI systems to output confidence intervals or reliability flags per source. An improved AI system can be trained to predict not just class labels but also a trustworthiness score based on feature completeness and redshift, allowing downstream cosmological analyses to exclude or down-weight unreliable predictions.

  6. Completeness Correction via Luminosity Function Comparison: The use of a luminosity function comparison (with Matthews et al. 2021) to correct the AI-selected sample for completeness is a transferable technique. An improved AI system can include a post-processing step that compares the predicted source population’s luminosity function against a known reference to estimate and correct for selection biases, improving the accuracy of derived astrophysical quantities like SFRD.

What the Improved AI System Can Do:

  • Automatically classify millions of extragalactic radio sources into galaxies vs. quasars with >93% weighted F1 score, even when the quasar class is rare, using random oversampling and multi-wavelength features.

  • Estimate local star formation rates and the cosmic star formation rate density (SFRD) at z < 0.1 with an accuracy of (1.4 ± 0.5) × 10−2 M⊙ yr−1 Mpc−3, consistent with established literature, while flagging unreliable high-redshift predictions.

  • Provide per-source uncertainty estimates and completeness-corrected outputs, enabling robust statistical analyses of galaxy evolution without manual catalogue curation.

  • Transfer its trained classification and SFR calibration to deeper future surveys (e.g., EMU), automatically adapting to new sensitivity limits via depth-matching and luminosity-function corrections, thereby probing SFRD evolution to higher redshifts with improved precision.

Abstract

Understanding the evolution of the cosmic star formation rate density (SFRD) is key to uncovering how the Universe arrived at its present state. This paper presents a novel and efficient method to estimate the local SFRD, which uses supervised machine learning to first identify a population of star-forming galaxies (SFGs). Next, star-formation rates (SFRs) are determined using the 1.4-GHz radio-continuum emission detected by the Australian Square Kilometre Array Pathfinder (ASKAP). Specifically, a gradient-boosted decision tree model was implemented to classify extragalactic sources from the Beck et al. (2022, MNRAS, 515, 4711) catalogue as either galaxies or quasars using RACS-mid and WISE photometry. The full sample, consisting of 389,392 sources, was partitioned into a 70%-15%-15% split for training, validating, and testing. The optimised model achieved a weighted F1 score of 0.93 and an accuracy of 0.94 on the test dataset, ultimately classifying 336,674 sources as galaxies and 52,718 sources as quasars. Using the resulting z<0.1 depth-matched galaxy sample and the photometric redshift predictions from Beck et al. (2022, MNRAS, 515, 4711), a modified 1.4-GHz SFR calibration was determined, yielding a local, completeness-corrected, z<0.1 SFRD of (1.4 plus or minus 0.5) times 10-2; M, yr-1,Mpc-3 using 11,293 sources. This value is consistent with previous results. Thus, this study demonstrates the feasibility of using supervised learning to identify large populations of SFGs in order to investigate the SFRD evolution. This presents an exciting prospect for future, deeper surveys such as EMU, which will enable the cosmic SFRD to be probed out to higher redshifts.

Sources

Related papers