Angular clustering and bias of photometric quasars in the Kilo-Degree Survey Data Release 4

arXiv:2511.17311 · astro-ph.CO, astro-ph.GA · Submitted 2026-08-20 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Astrophysics Radio. Generated commentary on the latest astrophysics papers.

Vera: Next we'll be talking about the paper "Angular clustering and bias of photometric quasars in the Kilo-Degree Survey Data Release 4".

Jocelyn: The paper was written by Anjitha John William, Maciej Bilicki, Wojciech A. Hellwing, Szymon J. Nakoneczny and Priyanka Jalan from Center for Theoretical Physics, Polish Academy of Sciences, al. Lotników 32/46, 02-668 Warsaw, Poland.

Vera: Stay tuned as we take you through the paper and discuss its implications.

Summary and Implications: Vera: The summary of the findings shows that these one hundred fifty-seven thousand quasars are definitely not spread out evenly across the sky; they exhibit distinct clustering.

Jocelyn: It’s a powerful confirmation that gravity is at work, pulling matter together into large structures over billions of years.

Subrahmanyan: The paper provides very concrete results by quantifying this clustering and directly linking it to the characteristic scale of dark matter halos at different epochs.

Vera: What’s particularly striking is how they measure the effective bias, showing that this property changes depending on the redshift of the objects we observe.

Jocelyn: It’s a powerful validation point because it gives us empirical evidence that quasars act as excellent cosmological tracers, confirming decades of theory about structure formation.

Subrahmanyan: This allows us to essentially see a timeline of structure formation, providing crucial data points that help validate N-body simulations by checking if the observed clustering matches the simulated growth.

Vera: The data isn't just reporting statistics; it’s offering a clear signal about how this clustering evolves over time.

Jocelyn: We can see a clear evolution in how strong this clustering is, which we track across those four defined tomographic bins.

Subrahmanyan: It’s like having a cosmic clock for growth, providing critical data points that help us understand the physics of structure formation without relying on single snapshots of how things look today.

Vera: So, these results aren't just academic numbers; they are providing a deep validation for the prevailing cosmological paradigm by showing that the quasar distribution adheres to predictable patterns governed by gravity across timescales.

Jocelyn: That leads us naturally into how they achieved such reliable measurements, which we’ll discuss next when dealing with all those potential errors in identifying and measuring those objects.

Improvements and Methodology: Vera: Now we are looking at the methodology because, as exciting as the results are, the methods truly make this work rigorous and dependable.

Jocelyn: It’s a massive undertaking to ensure that they rigorously clean up every source of error, from systematic noise to external contamination that could skew our results.

Subrahmanyan: I found their modeling for foreground contaminants particularly impressive, as they address issues like local artifacts or interference that are purely terrestrial, preventing us from accidentally attributing instrumental noise to cosmic phenomena.

Vera: Think about the complexity of running a survey over such a massive area; we are battling everything from local variations in calibration to different patches of sky.

Jocelyn: And beyond just cleaning up the background, they introduced this highly advanced machine learning framework called Hybrid-z for estimating the redshift of those objects. This is far more sophisticated than traditional methods that can fail spectacularly when the data quality drops off across a survey area.

Subrahmanyan: By demanding consistency across all these different data types—the images, the magnitudes, and other physical properties—they drastically reduce what are called spectral degeneracies, which is a huge win for any constraint-based study.

Jocelyn: The result of this high precision is that when we calculate the clustering strength in different redshift bins, we can be much more certain that those bins truly represent narrow, distinct epochs of cosmic history. It boosts our confidence immensely in what the sky is showing us at specific times.

Vera: So, these methodological refinements are not just technical details; they are the backbone that allow us to translate complex theoretical predictions into highly reliable empirical measurements.

Jocelyn: That leads directly into how this reliability allows us to interpret the results and what it all boils down to—the overall implications for the future of cosmic surveys.

Discussion of Key Results: Vera: We’ve covered an immense amount of ground, looking at the data and methodology from "Angular clustering and bias of photometric quasars in the Kilo-Degree Survey Data Release four" and it’s clear this has been a foundational piece of work for KiDS.

Jocelyn: It really showcases how far we've come with wide-field imaging, using advanced AI to get a reliable picture of the universe that spans millions of light-years.

Subrahmanyan: The results are incredibly powerful because they confirm that quasars are indeed excellent tracers for understanding the physics of how dark matter halos grow and cluster over cosmic time.

Vera: We have seen how their bias increases significantly with redshift, from about one point six at low z to nearly four point zero at high z, which is a clear signal that we’re seeing these objects in progressively more massive structures as time goes on.

Jocelyn: The paper's use of the Hybrid-z model and its extensive analysis allows us to build a robust statistical picture that gives us incredible confidence in the distribution of these sources across the sky.

Subrahmanyan: Looking ahead, this work paves the way for future extensions into larger surveys like LSST or 4MOST, which is exciting because we are laying a critical groundwork for what's coming next.

Vera: We’ve also seen how they corrected for stellar contamination and studied how the choice of redshift distribution affects bias estimates, so that’s crucial to keep in mind as we move forward.

Jocelyn: It is definitely a rigorous approach, ensuring that the observations reveal not just where the quasars are, but what those patterns tell us about the overall structure of their environment.

Subrahmanyan: We are essentially getting a detailed map of gravity's influence on cosmic evolution using objects identified by advanced color analysis rather than perfect spectral measurements.

Vera: So, to wrap up this discussion—based on our entire conversation today—the "Angular clustering and bias of photometric quasars in the Kilo-Degree Survey Data Release four" provides a definitive, robust framework for studying how these bright objects trace the gravitational scaffolding of the early universe.

Jocelyn: It's an exciting step forward, truly showing that we are better equipped to handle large-scale surveys than ever before.

Subrahmanyan: We can’t wait to see how this work scales up in future projects, building on the foundation they’ve established today.

Final Wrap-up: Vera: If we have to distill everything we’ve discussed into one core takeaway, it is that this work has provided an incredibly robust and precise tool for mapping the gravitational architecture of the early universe.

Jocelyn: Exactly. What’s truly remarkable isn't just *what* they found—the confirmation of cosmic structure growth—but the sheer methodological sophistication required to achieve such high confidence in those measurements. They essentially built a gold standard for wide-field photometric surveys.

Subrahmanyan: From a theoretical standpoint, it’s incredibly satisfying because it confirms that the physics we've been modeling using general relativity and dark matter scaffolding are not just mathematical constructs; they are observable patterns imprinted on the sky.

Vera: It really underscores how critical large, deep surveys are in modern astrophysics. They allow us to test our most ambitious models against real cosmic reality across time and space simultaneously.

Jocelyn: And that reliability—the ability to trust that measured clustering strength at a specific redshift bin—is what opens up entirely new avenues of research for the next generation of telescopes and simulations.

Subrahmanyan: In essence, they’ve provided a detailed blueprint of how mass accumulated over time, making this study a crucial anchor point for our entire understanding cosmic evolution.

Vera: So, as we wrap up our discussion on "Angular clustering and bias of photometric quasars in the Kilo-Degree Survey Data Release four" we leave with immense appreciation for the power of rigorous data analysis paired with massive survey capabilities.

Jocelyn: It's a phenomenal piece of work that solidifies the methodology for future efforts, leaving us all incredibly excited about what other cosmic structures await discovery.

Subrahmanyan: We are now ready to hear what’s next in the field, building on this foundation they’ve established today.

Anjitha John William, Maciej Bilicki, Wojciech A. Hellwing, Szymon J. Nakoneczny, Priyanka Jalan

Center for Theoretical Physics, Polish Academy of Sciences, al. Lotników 32/46, 02-668 Warsaw, Poland

astro-ph.CO, astro-ph.GA

Submitted: 2026-08-20

Updated: 2026-08-24

Comments: 13 pages, 7 figures

Journal ref: A&A, 712, A196 (2026)

DOI: 10.1051/0004-6361/202558249

Code: https://github.com/Anjithajm/Hybrid-z

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 100/100

The gist: We investigate "the angular clustering and effective bias of photometrically selected quasars in the Kilo-Degree Survey Data Release 4 (KiDS DR4)." The study utilizes a deep learning framework called

Key concepts

Quasar Clustering
The finding that quasars are not spread evenly across the sky. This clustering demonstrates that gravity is actively pulling matter together into large structures, providing empirical evidence for how the universe formed.
Effective Bias
A property measured by the paper showing how the clustering of quasars changes depending on their redshift (distance). Measuring this bias helps scientists understand how quasars act as reliable tracers of cosmic structure formation over time.
Photometric Quasars
Quasars identified using advanced color analysis and imaging data rather than perfect spectral measurements. This method is crucial for large-scale surveys, allowing researchers to map the universe's structure across vast areas.
Tomographic Bins
The technique of dividing the observed sky into multiple redshift bins. This allows scientists to track how the strength of clustering changes across different epochs of cosmic history, effectively creating a 'cosmic clock' for structure growth.

Terminology

Summary

We investigate the angular clustering and effective bias of photometrically selected quasars in the Kilo-Degree Survey Data Release 4 (KiDS DR4). The study utilizes a deep learning framework called Hybrid-z, which is designed to update the photometric redshifts (photo-zs) of KiDS quasars. This framework combines four-band KiDS images and nine-band KiDS+VIKING magnitudes. Hybrid-z was trained on a combined dataset of the latest Dark Energy Spectroscopic Instrument (DESI) DR1 and Sloan Digital Sky Survey (SDSS) DR17 quasars matching with KiDS. The performance of the model is characterized by achieving "average bias delta z < 0.01 and scatter about 0.04(1 + z) on a test sample."

The updated catalog, comprising about 157k quasars over 777 deg squared, is divided into four tomographic bins spanning 0.1 z phot 2.7. In these bins, the researchers measure the angular two-point correlation function and compare it with theoretical predictions for dark matter clustering.

The results of the clustering analysis show that the quasar bias increases from b about 1.6 at z about 0.6 to b about 4.0 at z about 2.2, and is well matched by a quadratic relation in redshift. Furthermore, the analysis indicates that KiDS quasars reside in dark matter halos of mass 10 (M eff / h-1 M) in the range about 12.7-12.9 and effective peak heights nu eff rising from about 1.5 to 2.9 over our redshift span.

The study also examines two potential systematic effects: stellar contamination and the redshift distribution assumed in the theoretical modeling. The findings regarding these systematics are as follows: The former [stellar contamination] has a negligible effect, whereas the latter [redshift distribution] significantly impacts the derived b(z), emphasizing the importance of redshift calibration.

The methodology for calculating angular clustering involves using the angular two-point correlation function (omega(theta)), which is calculated using the standard Landy & Szalay (1993) estimator: omega obs(theta) = DD(theta) - 2DR(theta) + RR(theta) over 1. The measurements are performed in nine logarithmically spaced angular bins, but the analysis focuses on four tomographic bins defined by z phot = 0.1, 0.8, 1.2, 1.8, 2.7 (Table I).

The derivation of quasar bias (b(z)) is achieved by modeling the observed angular two-point correlation function (omega q(theta)) as a scaled version of the matter clustering omega m(theta), specifically omega q(theta) = b 2(z) omega m(theta). The best-fit bias values are determined through a chi-squared minimization (chi 2) process.

A detailed comparison was made between using the photometric redshift distribution (dN/dz phot) and using the cross-matched spectroscopic redshift distribution (dN/dz spec). The researchers found that using dN/dz spec yields systematically higher bias values compared to dN/dz phot, as the amplitudes of the theoretical predictions for dark matter 2PCF are lower in the former case than in the latter. Despite these differences, they concluded that the choice of the underlying redshift distribution has a non-negligible impact on the inferred quasar bias estimation.

The final conclusions are that "this work is the first cosmological application of quasars selected from KiDS and paves the way for future extensions in the final KiDS DR5, the Legacy Survey of Space and Time, or the 4-metre Multi-Object Spectroscopic Telescope."

Improvements for AI systems

As an expert in AI research, I have analyzed the methodology presented in this paper to derive several critical, high-impact improvements that can be implemented across various advanced AI systems. These changes move beyond simple classification and into robust multi-modal inference and systematic error mitigation.

The core innovation of the Hybrid-z model is its successful integration of two distinct data modalities: high-resolution visual features (via CNN processing of images) and statistical photometric features (via ONN processing of magnitudes).

  • Improvement: Generalizing the Dual-Branch Architecture for Multi-Modal Input. The system should be designed to accept and process multiple, heterogeneous input types simultaneously—not just image/magnitude, but potentially time-series data, spectral fingerprints, or environmental context (e.g., local density maps).

  • What the Improved System Can Do: This allows the AI to perform robust classification and estimation even when one modality is degraded or missing. For instance, in astronomical surveys with limited resolution (like KiDS), if visual features are ambiguous due to stellar contamination, the system can rely on the statistical magnitude features (ONN branch) for a reliable output, achieving superior fault tolerance compared to single-input models.

The paper significantly improved its training set by incorporating data from DESI DR1 and SDSS DR17, which are more comprehensive than the original SDSS DR14 used in previous work.

  • Improvement: Implementing Adaptive, Stratified Training Set Augmentation. The AI system must be trained not only on a single dataset but on a dynamically expanded set that covers multiple, diverse observational regimes (e.g, combining data from different telescopes or different epochs). This requires an automated cross-matching and weighting strategy.

  • What the Improved System Can Do: Accurately model complex parameter spaces. By training on a wider range of input conditions, the the AI can handle edge cases (like high-z quasars) with greater confidence. The system is now capable of generalizing its knowledge to unseen data that falls outside its initial distribution, leading to significantly higher predictive accuracy than models constrained by limited original training sets.

The paper dedic significant effort to modeling two key systematics: stellar contamination and the uncertainty in the assumed redshift distribution (dN/dz).

  • Improvement: Integrating Self-Correction via Probabilistic Systematics Modeling. The AI system should incorporate a Bayesian framework that allows it to estimate not just a single answer, but also the probability of failure due to known physical or observational biases (e.g., contamination). This requires calculating the expected impact of systematic errors on the final output.

  • What the Improved System Can Do: Quantify uncertainty in cosmological inference. Instead of simply reporting a result (e.g, b about 4.0), the system will report a confidence interval that explicitly accounts for both statistical noise and systematic uncertainties (e.g, b = 4.0 plus or minus 1.29 at z=2.15, with an added plus or minus delta uncertainty due to modeling assumptions). This allows researchers to determine if the observed signal is a true physical phenomenon or merely a byproduct of imperfect data processing.

The paper successfully relates angular clustering (omega q) to dark matter clustering (omega m) via an effective bias b(z) and derives host halo mass (M eff).

  • Improvement: Developing a Nested Inference Framework for Bias and Parameter Estimation. The AI system should use the observed clustering statistics not just as a final check, but as an integrated component of its prediction model. Instead of performing sequential steps (1. Estimate z, 2. Calculate omega q, 3. Solve for b), the system performs a simultaneous, iterative minimization across multiple variables (z, b, M eff).

  • What the Improved System Can Do: Simultaneously constrain multiple physical properties. The AI can provide a single optimized solution that links the observed spatial distribution of objects (clustering) directly to their host environments (halo mass). This allows for rapid testing of cosmological models and enables the identification of features (like b(z) increasing with redshift) that would be missed by traditional, decoupled analysis.


Summary of System Capability: The improved AI system is no longer merely a classifier or a correlation calculator. It is a Multi-Modal, Self-Correcting Inference Engine capable of processing diverse data inputs, quantifying systematic uncertainties in its predictions, and providing statistically robust estimates for complex physical properties like bias and host halo mass simultaneously.

Abstract

We investigate the angular clustering and effective bias of photometrically selected quasars in the Kilo-Degree Survey Data Release 4 (KiDS DR4). We update the previous photometric redshifts (photo- z s) of the KiDS quasars using Hybrid-z, a deep learning framework combining four-band KiDS images and nine-band KiDS+VIKING magnitudes. Hybrid-z is trained on the latest Dark Energy Spectroscopic Instrument (DESI) DR1 and Sloan Digital Sky Survey (SDSS) DR17 quasars matching with KiDS, and achieves average bias δz < 0.01 and scatter about 0.04(1 + z) on a test sample. The updated catalog of about 157k quasars over 777 deg squared is divided into four tomographic bins spanning 0.1 at most z phot at most 2.7. In each bin, we measure the angular two-point correlation function and compare it with theoretical predictions for dark matter clustering. We estimate the best-fit scale-independent quasar bias, which increases from b about 1.6 at z about 0.6 to b about 4.0 at z about 2.2, and is well matched by a quadratic relation in redshift. Our clustering analysis indicates that KiDS quasars reside in dark matter halos of mass 10(M eff/h-1M) in the range about 12.7 -- 12.9 and effective peak heights ν eff rising from about 1.5 to 2.9 over our redshift span. We study two systematics that could affect the bias derivation: stellar contamination and the redshift distribution assumed in the theoretical modeling. The former has a negligible effect, whereas the latter significantly impacts the derived b(z), emphasizing the importance of redshift calibration. Our work is the first cosmological application of quasars selected from KiDS and paves the way for future extensions in the final KiDS DR5, the Legacy Survey of Space and Time, or the 4-metre Multi-Object Spectroscopic Telescope.

Sources

Related papers