FLAGS II: Constraining Galaxy Formation Models with Dimensionality Reduction of Direct Observables
Jack C. Turner, Stephen M. Wilkins, William J. Roper, Aswin P. Vijayan
University of Sussex · University of Malta
astro-ph.GA
Submitted: 2026-08-12
Updated: 2026-08-14
Comments: 18 pages, 10 figures. Submitted to the Open Journal of Astrophysics
Code: https://github.com/jackcturner/crest
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: FLAGS II: CONSTRAINING GALAXY FORMATION MODELS WITH DIMENSIONALITY REDUCTION OF DIRECT OBSERVABLES This paper presents the second installment of the FLAGS series, which aims to derive robust
Terminology
Summary
FLAGS II: CONSTRAINING GALAXY FORMATION MODELS WITH DIMENSIONALITY REDUCTION OF DIRECT OBSERVABLES
This paper presents the second installment of the FLAGS series, which aims to derive robust constraints on galaxy formation models using direct observables. The authors investigate whether 2D embeddings of JWST and HST photometric fluxes, constructed using the non-linear dimensionality reduction algorithm UMAP, preserve sufficient information to statistically differentiate between five galaxy formation models.
Data and Methods:
-
Observations: The authors use JWST and HST imaging of the GOODS-S field, sourced from the DAWN JWST Archive (DJA). The JWST data was collected primarily as part of the JADES survey, spanning 100 arcmin2. They selected eight JWST/NIRCam filters (F090W, F115W, F150W, F200W, F277W, F356W, F410M, F444W), supplemented by HST/WFC3 F125W and F160W, and HST/ACS F435W, F606W, and F814W, for a total of thirteen filters.
-
Source extraction: Galaxies were detected using Source Extractor on an inverse-variance weighted stack of F277W, F356W, and F444W images. Photometry was measured in Kron apertures with aperture corrections applied. After masking, the final unmasked sky area is 43 arcmin2.
-
Sample selection: The authors apply a pseudo mass cut using the F444W apparent magnitude, selecting all sources with mF444W 99% having M* > 107 M⊙. They require each galaxy to be detected with S/N > 3 in at least five filters.
-
Models: Five models are compared: two semi-analytic models (sc-sam and sage), and three semi-empirical models (jaguar, spritz1, and spritz4). Synthetic noise is added to model fluxes to account for observational uncertainties.
-
Dimensionality reduction: The authors demonstrate that comparing distributions in the native thirteen-dimensional space is unfeasible due to sparse sampling and memory demands. They show that even with decile binning, the space becomes sparsely sampled beyond two dimensions, and the memory footprint for thirteen dimensions exceeds 20 TB. UMAP is used to generate 2D embeddings, fitting each model 500 times with different random seeds to account for stochasticity.
Key Results:
-
Euclidean metric (q ≡ ρ): Using a euclidean distance metric and comparing sky area densities, the authors find that jaguar performs best overall with S e = 408+33/−25. sc-sam is the best causal model with S e = 2532+162/−75, performing almost twice as well as sage (S e = 4731+89/−99). The spritz models are statistically consistent with each other. The authors note that "jaguar reproduces the population of bright galaxies (mAB < 26) in GOODS-S six times as well as sc-sam and twelve times as well as sage."
-
Cosine metric (q ≡ f): Using a cosine distance metric and comparing fractions of datasets, the models perform more poorly overall. jaguar remains the standout performer with S c = 880+88/−66. sc-sam is second best with S c = 3538+126/−147, sage worsens to S c = 5377+181/−232, and both spritz models have S c > 8000. The spritz models fail to occupy large regions of the space due to their template-based approach, which limits them to 39 rest-frame SED shapes.
-
Comparison with PCA: The authors find that PCA produces systematically lower scores than UMAP, showing that ignoring non-linear components oversimplifies distributions. The difference ranges from 25% for sage to >90% for jaguar. Importantly, the relative performance of models changes between PCA and UMAP spaces, with sc-sam performing better than spritz1 in UMAP but worse in PCA.
-
Physical interpretation: The most discrepant bin in the cosine space contains galaxies identified as dust-poor starbursts at 2.5 750 Å). The authors suggest that the failure of sc-sam and sage to produce comparable galaxies may be due to limited snapshot cadence in the underlying DMO simulation, which washes out short-lived starburst features. The lack of nebular emission modelling in sage also contributes to discrepancies.
-
Degradation layer: By removing each filter in turn, the authors identify that longer wavelengths appear to be more poorly modelled. Removing F200W reduces the sage score by 10%, likely due to [OIII] emission at z 3. The spritz models show a steady decline between F200W, F277W, and F356W, suggesting issues with rest-frame emission in the range 1 < λ rest < 2 µm.
-
Systematics: The results are largely insensitive to systematic uncertainties from source extraction parameters, source extraction codes (SEP, Photutils, ProFound), noise scaling, and UMAP hyperparameters. No justifiable variation changes the ordering of model performance. The absolute agreement varies at the < 0.1 dex level.
-
Computational efficiency: Each 500-iteration UMAP fit takes 0.007 seconds of CPU time per object, which is >100 times faster than Bayesian SED fitting with Bagpipes (>400 seconds per object). The authors suggest this makes the approach ideal for large surveys such as LSST and Euclid, and for simulation-based inference with CAMELS.
Conclusions: The authors demonstrate that UMAP dimensionality reduction of direct observables provides a viable alternative to SED fitting for statistically constraining galaxy formation models. The approach preserves sufficient non-linear information to differentiate between models, is largely insensitive to systematic uncertainties, and is computationally efficient enough for large datasets. The template SED approach of spritz and the lack of photoionisation in sage cause significant discrepancies, highlighting the importance of comprehensive forward modelling.
Improvements for AI systems
Improvements to AI systems:
-
Non-linear dimensionality reduction for model comparison: Implement UMAP-based 2D embeddings of raw multi-band photometry as a drop-in replacement for traditional SED fitting in simulation-based inference. The AI system can now compare galaxy formation models against observations in a fraction of the computational cost (0.007s/object vs >400s/object), enabling real-time model selection for large surveys like LSST and Euclid.
-
Stochastic ensemble fitting with uncertainty quantification: Use the 500-random-seed UMAP fitting approach to produce robust statistical scores with asymmetric error bars (e.g., S e = 408+33/−25). An improved AI system can automatically propagate stochasticity from dimensionality reduction into model rankings, providing confidence intervals on model performance rather than point estimates.
-
Metric-aware model evaluation: Adopt dual metrics (Euclidean for density comparison, cosine for fraction comparison) to capture different aspects of model fidelity. The AI system can now flag models that match overall galaxy densities but fail in distribution shape (e.g., spritz models), enabling more nuanced diagnostics of where models break down.
-
Filter-specific degradation analysis: Integrate the
leave-one-filter-out
approach into model validation pipelines. The AI system can automatically identify which photometric bands are most poorly modelled (e.g., F200W for sage due to [OIII] emission), pinpointing physical processes (nebular emission, starburst timescales) that need improvement in forward models. -
Physical outlier detection via embedding residuals: Use the most discrepant bins in the UMAP space (e.g., dust-poor starbursts at 2.5 < z < 3.5 with extreme [OIII]) as automated triggers for follow-up analysis. The AI system can now identify rare but physically important galaxy populations that standard SED fitting would miss, and trace their absence to specific model limitations (e.g., snapshot cadence in DMO simulations).
-
Systematic robustness scoring: Implement the systematic variation testing (source extraction parameters, codes, noise scaling, UMAP hyperparameters) as an automated validation layer. The AI system can now certify that model rankings are stable under reasonable methodological choices, with absolute agreement varying <0.1 dex, making results trustworthy for publication.
-
Template-free vs. template-based model discrimination: Leverage the finding that template-based models (spritz) fail to occupy large regions of the embedding space due to limited SED shapes. The AI system can automatically classify models as
template-limited
if their UMAP density shows large empty regions, guiding developers toward more flexible forward modelling. -
Hierarchical model ranking with causal priors: Use the distinction between causal (sc-sam, sage) and semi-empirical (jaguar) models to build an AI system that prioritizes physically motivated models while still identifying the best empirical fit. The system can report both
best overall
andbest causal
rankings, aiding theoretical interpretation. -
Scalable simulation-based inference pipeline: Combine the UMAP approach with CAMELS-style simulations to create an AI system that can rapidly explore parameter spaces of galaxy formation models. The >100x speedup over SED fitting enables iterative model refinement in near-real-time, allowing researchers to test new physics (e.g., feedback prescriptions) within minutes rather than days.
-
Cross-survey transferability: The system can be retrained on different survey footprints (e.g., COSMOS, CEERS) with minimal overhead, using the same UMAP embedding framework. It will automatically adapt to different filter sets and depth, providing consistent model constraints across surveys without re-deriving complex likelihoods.
Abstract
Comparisons between observations of galaxies and theoretical predictions are regularly performed using physical properties, which are inferred by the often slow and biased process of SED fitting. Forward modelling facilitates a reliable alternative, whereby models are evaluated using direct observables alone. However, these datasets become high-dimensional when collating observations from multiple telescopes, leading to sparse sampling, memory intensity and visualisation difficulties. We show that 2D embeddings of JWST and HST photometric fluxes, constructed using the non-linear dimensionality reduction algorithm UMAP, preserve sufficient information to differentiate between five models. Using a simple chi squared-like metric, we show that JAGUAR reproduces the population of bright galaxies (m AB<26) in GOODS-S six times as well as SC-SAM and twelve times as well as SAGE. By adjusting the hyperparameters, we quantify how well each model replicates the distribution of SED shapes. The template SED approach of SPRITZ and the lack of photoionisation in SAGE cause significant discrepancies, highlighting the importance of comprehensive forward modelling. The embedded position of each galaxy can be identified >100 times faster than inferring its properties with Bayesian SED fitting, making this approach an ideal alternative for deriving statistical model constraints from large surveys such as LSST and Euclid, and performing simulation-based inference with CAMELS.
Sources
- Optimizing Photometric Redshift Training Sets I: Efficient Compression of the Galaxy Color-Redshift Relation with UMAP
- The Spectral Energy Distributions of Galaxies
- MEGATRON: Reproducing the Diversity of High-Redshift Galaxy Spectra with Cosmological Radiation Hydrodynamics Simulations
- UMAP: Uniform Manifold Approximation and Projection for Dimension Reduction
- First Light and Assembly of GalaxieS (FLAGS) I: The JWST/NIRCam Number Counts and IGL as Constraints on Galaxy Formation Models
Related papers
- Apparent Stability in Self-Gravitating Turbulence and the Evolution of Molecular Clouds
- Two sets of potential-density basis pairs for the study of radial perturbations in collisionless spherical stellar systems
- Constraining reionization-era Ly alpha escape with JELS-MUSE: a highly complete H alpha-selected sample at z about6.1
- Deriving volume density profiles of filaments from observed surface densities
- Little Red Dots and Supermassive Black Hole Seed Formation in Ultralight Dark Matter Halos
- MEGATRON: how the first stars can create an iron metallicity plateau in the smallest dwarf galaxies