Systematic Effects of Hydrogen and Helium Atmosphere Mismatch on Radius Inference in PSR J0740+6620-like Synthetic NICER Data
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: Astrophysics Radio. Generated commentary on the latest astrophysics papers.
Vera: Next we'll be talking about the paper "Systematic Effects of Hydrogen and Helium Atmosphere Mismatch on Radius Inference in PSR J0740+6620-like Synthetic NICER Data".
Jocelyn: The paper was written by Isiah M. Holt, M. Coleman Miller, Alexander J. Dittmann and Frederick K. Lamb from Department of Astronomy, University of Maryland, College Park, MD 20742-2421, USA and NASA Goddard Space Flight Center, Greenbelt, MD USA and Joint Space-Science Institute, University of Maryland, College Park, MD 20742-2421 USA and Institute for Advanced Study, 1 Einstein Drive, Princeton, NJ 08540 USA and Illinois Center for Advanced Studies of the Universe and Department of Physics, University of Illinois at Urbana-Champaign, 1110 West Green Street, Urbana IL 61801-3080 USA and Department of Astronomy, University of Illinois at Urbana-Champaign, 1002 West Green Street, Urbana IL 61801-3074 USA.
Vera: Stay tuned as we take you through the paper and discuss its implications.
Paper discussion segment 1: Vera: We've got a fascinating one today called "Systematic Effects of Hydrogen and Helium Atmosphere Mismatch on Radius Inference in PSR J0740+six thousand six hundred twenty-like Synthetic NICER Data." Jocelyn, the title alone makes me think about all those hours we spend staring at noisy X-ray data from NICER.
Jocelyn: It sounds pretty intimidating at first glance, but it's basically asking if we're being fooled by what we think a neutron star's atmosphere is made of. If the model assumes hydrogen but it's actually helium, are our measurements of the star's size actually correct?
Vera: Exactly, and looking at the authors like Isiah Holt and M. Coleman Miller, this team really knows their way around these specific pulsars. They aren't just guessing; they're using high-mass pulsars as their benchmark to test these errors.
Jocelyn: I'm wondering if our current observations are actually robust enough to handle these mismatches, or if we're building our whole understanding of dense matter on a potentially shaky foundation.
Subrahmanyan: That is the critical question because the radius tells us everything about the equation of state. If we get the radius wrong by even a couple of kilometers because we picked the wrong gas for our model, our entire theoretical picture of what's happening inside those cores gets shifted.
Vera: It’s a massive scale error for something that happens at such a microscopic level in the atmosphere.
Jocelyn: So, they aren't looking at real stars yet, but using synthetic data to simulate how much we might be messing up?
Subrahmanyan: Precisely, they are running controlled experiments to see exactly where our modeling breaks down. They used the properties of PSR J0740+six thousand six hundred twenty because it's a heavy-hitter at about two point one solar masses, which puts extreme pressure on these models.
Vera: We should probably look at how they actually set up these "fake" stars to see if their simulation matches what we see in the sky.
Paper discussion segment 2: Jocelyn: Now that we know they're using synthetic data based on PSR J0740+six thousand six hundred twenty let's look at what they actually found in their experiments. Vera, the way they manipulated the signal strength is really clever.
Vera: It was a smart move to keep the total count at about five hundred fifty thousand but change how much of that was "hot spot" versus background noise. They found that if the signal is mostly background, like what we see with J0740+six thousand six hundred twenty you can get away with using the wrong atmosphere model without a huge radius error.
Jocelyn: Wait, so the noise actually hides our mistakes? That sounds like a dangerous way to do science.
Vera: It is! When the "spot-to-background" ratio is low, the mismatch doesn't show up in the radius calculation significantly at first. But as soon as they made that signal much stronger—increasing that ratio by a factor of twenty—the errors started to creep in.
Subrahmanyan: This is where the physics gets interesting because there's a direction-dependent bias happening here. If you have hydrogen-generated data but fit it with a helium model, you might actually get an acceptable fit, but your radius estimate ends up being wrong.
Jocelyn: How much "wrong" are we talking about? Is it just a tiny nudge or something that changes the whole conclusion?
Subrahmanyan: In the case where they fit hydrogen-generated data with a helium model, the radius can be pushed up by more than two-sigma. That’s enough to make a researcher think they've discovered something new about neutron star density when really they just picked the wrong gas.
Vera: And yet, the paper says that even when the radius is biased, the "goodness-of-fit" tests might still say everything is fine.
Jocelyn: That’s what scares me; if the residuals look okay, we might never know we're wrong unless we use better math.
Paper discussion segment 3: Vera: It really comes down to how we validate these models, and the authors are pushing for something more robust than just looking at residuals. Jocelyn, they mentioned that the Bayesian evidence is actually much smarter than standard goodness-of-fit tests.
Jocelyn: Right, because while a "wrong" model can be tweaked to look okay on paper—meaning it has acceptable chi-squared values—it shouldn't be able to fool the Bayesian approach. The paper shows that the Bayesian evidence consistently points toward the correct atmosphere model even when the radius is being biased.
Vera: It’s like a more rigorous judge that looks at how much "fine-tuning" a model needs to work. A wrong model might fit, but it has to be very specific and unlikely to do so, so the Bayesian evidence penalizes it.
Jocelyn: So the suggestion is that we shouldn't just stop once our chi-squared values look good?
Vera: Exactly, they want us to use Bayesian model comparison as a standard part of the toolkit. If you aren't comparing hydrogen vs. helium using evidence, you might be missing a huge systematic error in your radius measurement.
Subrahmanyan: This has massive implications for the next generation of X-ray missions and high-precision surveys. As we get better data with higher signal-to-noise ratios, these atmosphere mismatches are going to become much more obvious and potentially very problematic if we aren't prepared.
Jocelyn: So, for the really bright pulsars where the signal is clear, we have to be extra careful about what kind of gas we're assuming is on the surface?
Subrahmanyan: Yes, because in those high-fidelity cases, the mismatch doesn't hide in the noise anymore; it shows up clearly as a physical error. The paper essentially provides a roadmap for how to avoid these traps by being much more disciplined with our statistical comparisons.
Vera: We should probably wrap this up before we get lost in all the math of Bayesian priors!
Conclusion: Jocelyn: This has been such an eye-opener regarding how much we rely on our assumptions about what these stars are made of. We've covered a lot, from how noise can hide errors to why the Bayesian evidence is our best friend for finding the truth.
Vera: It really highlights that even with amazing data from NICER, our interpretation is only as good as our models. If we want to know what's inside a neutron star, we have to be certain about what's on its surface.
Jocelyn: We've been discussing "Systematic Effects of Hydrogen and Helium Atmosphere Mismatch on Radius Inference in PSR J0740+six thousand six hundred twenty-like Synthetic NICER Data."
Subrahmanyan: It’s a vital reminder that in astrophysics, the "best fit" isn't always the true one, and we must always ask if there's a better model waiting to be tested.
Vera: Well said, Subrahmanyan. Thanks for joining us! We'll see everyone next time with another deep dive into the latest from arXiv. Goodbye!
Jocelyn: Bye everyone! See you at the next paper!
Subrahmanyan: Until then, keep looking up and questioning your models! Goodbye.--- END OF SCRIPT ---
Department of Astronomy, University of Maryland, College Park, MD 20742-2421, USA · NASA Goddard Space Flight Center, Greenbelt, MD USA · Joint Space-Science Institute, University of Maryland, College Park, MD 20742-2421 USA · Institute for Advanced Study, 1 Einstein Drive, Princeton, NJ 08540 USA · Illinois Center for Advanced Studies of the Universe and Department of Physics, University of Illinois at Urbana-Champaign, 1110 West Green Street, Urbana IL 61801-3080 USA · Department of Astronomy, University of Illinois at Urbana-Champaign, 1002 West Green Street, Urbana IL 61801-3074 USA
astro-ph.HE, astro-ph.IM
Submitted: 2026-09-17
Updated: 2026-09-17
Comments: 23 pages, 10 figures, 6 tables. Accepted to ApJ
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 89/100
The gist: This paper investigates how assuming an incorrect atmospheric composition (hydrogen vs.
Terminology
Summary
This paper investigates how assuming an incorrect atmospheric composition (hydrogen vs. helium) affects the inference of neutron star radii using synthetic pulse-profile modeling (PPM) data. Because neutron star radii provide unique insights into the nature of dense matter,
understanding these modeling systematics is critical for accurately constraining the equation of state (EOS) of matter in extreme environments.
The Research Objective
The study explores whether a mismatch between the true atmospheric composition and the assumed model can produce significantly biased radius estimates while remaining hidden inside a fit with statistically acceptable residuals.
Using synthetic data based on the 2.1 M⊙ pulsar PSR J0740+6620, the researchers test if an incorrect model can mimic a good fit through parameter compensation. The experiment specifically examines:
** The consequences of assuming a hydrogen atmosphere when the data is helium-based, and vice versa. **
** How the detectability of such mismatches depends on the relative strength of the pulsed signal
by varying the spot-to-background ratio. **
** Whether Bayesian evidence can successfully identify the correct model even when standard goodness-of-fit tests fail to reject an incorrect one. **
Methodology and Simulation
The researchers generated synthetic NICER-like pulse waveform data using fully ionized hydrogen or helium atmospheres. They adopted a two-uniform-temperature spot geometry
based on the best-fit solution for PSR J0740+6620. To test different signal strengths, they applied a multiplier (m) to the baseline spot-to-background ratio, ranging from 1 (the realistic, heavily background-dominated regime
) to 100 (a strongly spot-dominated scenario
). The analysis involved:
-
An input-parameter evaluation where the 12 primary parameters were held fixed.
-
A full Bayesian posterior exploration using the pocoMC sampler to allow all parameters to vary.
-
The calculation of both phase-channel and bolometric χ2 statistics to assess goodness-of-fit.
-
The computation of Bayesian evidence (Z) to perform model comparison via the log evidence ratio (∆ ln Z).
Key Findings on Radius Bias
The study reveals that the incorrect atmosphere model can yield statistically acceptable phase-channel and bolometric χ2 values
while simultaneously producing significant biases. The effect is notably asymmetric:
** Fitting hydrogen-generated data with a helium model (He, m → H, d) is the dangerous case,
as it can increase the inferred radius by more than 2σ in spot-dominated regimes. **
** Fitting helium-generated data with a hydrogen model (H, m → He, d) produces smaller increases that remain within 2σ for the cases we examined.
**
The authors suggest this asymmetry may be linked to the high compactness of PSR J0740+6620, noting that the amount of light bending grows rapidly as a star approaches the photon sphere,
making the modulation sensitive to radius changes.
The Role of Bayesian Evidence
A central conclusion is that χ2 acceptability is not a sufficient safeguard against atmospheric composition bias.
While an incorrect model can successfully absorb the mismatch
by adjusting other parameters to pass χ2 tests, the Bayesian evidence consistently identifies the correct atmospheric model.
Because the evidence integrates over the entire posterior volume, it penalizes models that require fine-tuned parameter combinations
to achieve a fit. Consequently, the authors argue that Bayesian evidence is an essential complement to goodness-of-fit statistics
for reliable radius inference in pulse-profile analyses.of fit tests. and the Bayesian evidence consistently identifies the correct atmospheric model. Our findings reinforce the importance of using the Bayesian evidence for model comparison and goodness-of-fit tests.
Keywords: Millisecond pulsars (1062); X-ray stars (1823); Neutron stars (1108); Neutron star cores (1107)
- INTRODUCTION
The cores of neutron stars consist of matter that
is theorized to be catalyzed to its ground state at
densities up to several times nuclear saturation density (ns ≈ 0.16 fm−3, equivalent to a mass density
ρs ≈ 2.6 × 1014 g cm−3). Laboratories are unable to
replicate the high densities, low temperatures, and extreme neutron–proton asymmetries present in neutron
star cores. Thus, observations of neutron stars provide unique insights into the nature of dense matter.
Most of the macroscopic properties of a neutron star,
such as its mass, radius, and tidal deformability, depend on the equation of state (EOS) of the matter in its
interior, i.e., the pressure as a function of the
energy density. Over the past decade, numerous ob∗ NASA Einstein Fellow
2
Isiah M. Holt et al.
servations have helped constrain the EOS. These include measurements of high neutron star masses such as
PSR J1614−2230 (M = 1.97 ± 0.04 M⊙, Demorest et al.
2010), PSR J0348+0432 (M = 1.806 ± 0.037 M⊙, Saffer et al. 2024), PSR J0740+6620 (M = 2.08 ± 0.07 M⊙,
Fonseca et al. 2021), and PSR J0952−0607 (M = 2.35 ±
0.17 M⊙, Romani et al. 2022). They also include constraints on the tidal deformability of neutron stars from
gravitational-wave observations of the binary neutron
star merger GW170817 (Abbott et al. 2017; Abbott et al. 2018; De et al. 2018), and estimates of neutron star
radii from thermal X-ray observations. Among these,
measurements of the radii of neutron stars with precisely
determined masses are especially valuable.
Fits of nonmagnetic helium model atmospheres to the
spectra of cooling neutron stars in quiescent low-mass Xray binary systems (qLMXBs; Guillot et al. 2011, 2013; Lattimer & Steiner 2014; Bogdanov et al. 2016) com--- Page 2 ---
2
monly yield radius estimates ∼ 50% larger than fits of
nonmagnetic hydrogen model atmospheres to the same
data (Servillat et al. 2012; Catuneanu et al. 2013; Heinke et al. 2014; see also Miller & Lamb 2016 for a review).
The origin of this discrepancy could lie in a spectral
shape degeneracy. Hydrogen and helium atmospheres
produce emergent spectra of different hardness at the
same effective temperature, so fitting a hydrogen atmosphere model to helium atmosphere data (or vice
versa) forces a trade-off between the inferred temperature and emitting area, and hence the radius. The soft X-ray spectra of these sources are roughly thermal and lack strong composition-dependent features, so
the data are often insufficient to distinguish between a
smaller, hotter neutron star with one composition and a
larger, cooler neutron star with another (see, e.g., Miller & Lamb 2016). Similarly, hydrogen and carbon model
atmospheres can give equally good fits to the spectra of
isolated cooling neutron stars while returning very different radii (see, e.g., Klochkov et al. 2015). In these
spectral fitting methods, the composition ambiguity is
difficult to resolve because the data lack the spectral
features or independent geometric constraints needed to
distinguish atmospheric models.
The Neutron Star Interior Composition Explorer
(NICER; Gendreau et al. 2012) conducts phase- and
energy-resolved measurements of the thermal X-ray
pulses produced by some rotating neutron stars, which
can be used to constrain the masses and radii of these
stars. Mass and radius constraints have been derived from NICER data, sometimes supplemented with
X-ray Multi-Mirror (XMM-Newton) data, for several
neutron stars. These include the ∼ 1.4 M⊙ pulsar
PSR J0030+0451 (Miller et al. 2019; Riley et al. 2019;
Vinciguerra et al. 2024; Kini et al. 2026), the ∼ 2.1 M⊙
pulsar PSR J0740+6620 (Miller et al. 2021; Riley et al.
2021; Salmi et al. 2022; Dittmann et al. 2024; Salmi et al.
2024a; this pulsar is especially important because of
its high mass and thus high central density), the ∼ 1.4 M⊙
pulsar PSR J0437–4715 (Choudhury et al. 2024; Miller
et al. 2026), the ∼ 1.04 M⊙ pulsar PSR 1231–1411
(Salmi et al. 2024b; Qi et al. 2025), the ∼ 1.4 M⊙
pulsar PSR J0164–3329 (Mauviard et al. 2025), and the
∼ 1.8 M⊙ pulsar PSR J2124–3358 (González-Caniulef et al. 2026). In all such analyses, the atmosphere of the
neutron star plays an important role, as it determines
the angular distribution of the thermal X-ray
radiation
leaving the stellar surface.
In pulse-profile modeling (PPM), the predicted X-ray
pulse waveform depends not only on the stellar compactness, observer geometry, and emission region properties, but also on the beaming of the emergent radia-
tion, which is influenced by the atmospheric composition (see, e.g., Pechenick et al. 1983; Poutanen & Gierliński
2003; Lo et al. 2013; Miller & Lamb 2015).
Published PPM analyses of NICER data have predominantly assumed partially and/or fully ionized hydrogen atmospheres when modeling the X-ray
emission (see, e.g., Miller et al. 2019; Riley et al. 2019; Miller et al. 2021; Riley et al. 2021; Salmi et al. 2024a; Dittmann et al.
2024; Choudhury et al. 2024; Miller et al. 2026).
Salmi et al. (2023) investigated atmosphere model
systematics in NICER pulse-profile analyses of
PSR J0030+0451 and PSR J0740+6620, including fully
ionized hydrogen and helium models. They found that
none of the atmosphere cases significantly changed the
inferred radius of PSR J0740+6620, which they attributed potentially to the source’s X-ray faintness,
tighter external constraints, and/or view geometry. Our
synthetic experiment is complementary to that result.
We confirm their result at low source-to-background
ratios, but demonstrate that the modeling assumption of
atmospheric composition can introduce overwhelming
systematic biases when analyzing higher-fidelity data.
In this paper, we check whether atmospheric composition mismatch can bias the inferred radius of a
PSR J0740+6620-like neutron star while remaining hidden inside a statistically acceptable fit. We generate
synthetic NICER-like pulse-profile data assuming fully ionized hydrogen or fully ionized helium
atmospheres, and fit the hydrogen-generated data with both atmosphere
models and the helium-generated data with both models. In each case, we vary the ratio of hot spot
counts to background counts while keeping the total number of counts fixed at ∼ 5.5 × 10
5, which is the number
observed from PSR J0740+6620, in order to explore how
the detectability of the composition mismatch depends
on the relative strength of the pulsed signal. For each
configuration, we perform input-parameter evaluations,
full posterior sampling, goodness-of-fit assessments, and
Bayesian evidence comparisons.
We find that in both directions of the atmospheric
mismatch, the incorrect atmosphere model can yield statistically acceptable phase-channel and bolometric χ2
values after posterior exploration, yet still produce biased radius estimates. The effect is direction-dependent.
Fitting hydrogen-generated data with a helium model
can increase the inferred radius by more than 2σ,
whereas fitting helium-generated data with a hydrogen
model produces smaller increases that remain within
2σ for the cases we examined. However, in these cases, the Bayesian evidence consistently identifies the
correct atmospheric model. Our findings reinforce
the importance of using the Bayesian evidence for
model comparison and goodness-of-fit tests.
The remainder of this paper is organized as follows:
Section 2 describes the effects of atmospheric beaming
of the emergent radiation on the pulse profile. Section 3
describes our methods for generating and analyzing the
synthetic data. Section 4 presents our results. We discuss the implications of these results in Section 5 and
summarize our conclusions in Section 6.
- ATMOSPHERIC BEAMING AND ITS EFFECTS
ON THE PULSE PROFILE
PPM is less susceptible to atmospheric uncertainty
than spectral fitting methods because the pulse waveform carries geometric information via phase-dependent
flux modulation. The time-varying projection of a heated region (“hot spot”) as the star rotates encodes the observer inclination, the spot colatitude, and
the stellar compactness through the relativistic lightbending that maps surface emission angles to the observer’s line of sight. Even if the beaming function is incorrect, the phase at which the spot appears and disappears and the overall depth of the modulation are primarily determined by the spot geometry and the star’s
compactness, reducing the freedom available for an incorrect atmosphere to absorb the mismatch. Indeed,
analyses of synthetic data have shown that PPM radius estimates are robust against several other classes of
systematic error. Lo et al. (2013) and Miller & Lamb
(2015, 2016) demonstrated that incorrect assumptions about the shapes and temperature distributions of the
hot spots do not significantly bias the inferred radius,
provided the fit is statistically acceptable. Holt et al.
(2025) showed that in joint NICER/XMM-Newton analyses, even a factor-of-five underestimate of the XMM-Newton background shifts the radius posterior by only
∼ 1σ. These results have strengthened the case for PPM
as a method for reliable radius inference. To our knowledge, however, no prior synthetic data study has systematically checked whether atmospheric composition
mismatch, specifically, data generated with one atmosphere using a model that assumes another, can produce
a comparably small or potentially larger bias.
In fully ionized atmospheres at the effective temperatures relevant to NICER (kTeff ∼ 0.1 keV), the
dominant sources of continuum opacity are free–free (inverse bremsstrahlung) absorption and electron scat-
tering. The free–free absorption coefficient scales as
αff ∝ Z 2 ni ne /ν 3 for ion charge Z, ion and electron
number densities ni and ne, and photon frequency ν
(see, e.g., Rybicki & Lightman 1979). The composition
dependence of this expression is more subtle than the
explicit Z 2 factor suggests. At a given mass density
ρ, a fully ionized hydrogen atmosphere has ni = ne =
ρ/mH, whereas a fully ionized helium atmosphere has
ni = ρ/4mH and ne = ρ/2mH. A helium nucleus
is four times the mass of a proton and supplies only two electrons. As a result, helium offers fewer targets and fewer
absorbers per gram. The reduction in ni cancels the
factor Z 2 = 4 exactly, and the accompanying reduction in ne leaves αff a factor of two lower for helium than for
hydrogen at the same density and temperature.
The beaming pattern, i.e., the intensity as a function of the
angle θ between the outgoing photon direction
and the surface normal, is shaped by how the radiation
decouples from the atmosphere, and the difference between the two compositions is therefore not a simple consequence of the ionic charge. On the basis of the scalings above, we expect it to be modest rather than dramatic.
Figure 1 shows the emergent beaming patterns for both of the model atmospheres. At every energy, the helium
atmosphere retains more intensity toward grazing angles than hydrogen, meaning that its beaming is closer to
isotropic, whereas hydrogen is more limb-darkened, consistent with previous calculations (see, e.g., Bogdanov et al.
2021; Salmi et al. 2023). The two patterns are
nonetheless similar in overall shape, and the difference
between them is small. Its magnitude depends on the effective
temperature, surface gravity, and photon energy.
For PPM, the consequence of this beaming difference is that a helium atmosphere directs relatively more radiation toward grazing angles compared to a hydrogen
atmosphere. When the star rotates and the hot spot
sweeps across the observer’s line of sight, the angular
distribution of the emitted radiation determines how
steeply the observed flux rises and falls with pulse phase.
A more limb-darkened (hydrogen-like) pattern produces a sharper pulse profile with deeper modulation for a given geometry, whereas a more isotropic (helium-like) pattern produces a broader, shallower pulse. We illustrate this in Figures 2 and 3, which compare synthetic pulse profiles generated with hydrogen and helium atmospheres for PSR J0740+6620-like geometry. At the
spot-to-background ratio of PSR J0740+6620, the two
compositions yield nearly identical waveforms, with fractional modulations of fH = 0.034 and fHe = 0.033 (Figure 2). Amplifying the ratio by a factor of 60 to suppress
the background exposes the shape difference, yielding
fractional modulations, fH = 0.484 and fHe = 0.471 (Figure 3). In both regimes, the hydrogen
atmosphere
produces the slightly deeper, sharper pulse expected
from its stronger limb darkening.
This is particularly relevant for PSR J0740+6620,
a ∼ 2.1 M⊙ pulsar (Fonseca et al. 2021) whose high
mass makes it valuable for constraining the EOS of
cold, dense matter at densities above those accessible
through observations of ∼ 1.4 M⊙ neutron stars (see, e.g., Miller et al. 2021; Riley et al. 2021; Salmi et al. 2024a; Dittmann et al. 2024; Salmi et al.
2024a; this pulsar is especially important because of its
high mass and thus high central density), the NICER observations
of PSR J0740+6620 are severely background-dominated
(Wolff et al. 2021), with the estimated ratio of hot spot
counts to total counts in the NICER band being of order
a few percent. This means that the pulsed signal
from the spots is superimposed on a large background
of counts from other X-ray sources in the field. This
low spot-to-background ratio has two competing implications for the composition mismatch problem. It may
make the waveform less sensitive to the beaming pattern, because the beaming-dependent modulation constitutes a smaller fraction of the
total signal and the statistical uncertainties in the pulse shape are large. On
the other hand, if the beaming mismatch does shift the
inferred parameters, the large background means that
the shift may be harder to detect through standard
goodness-of-fit diagnostics.
- METHODS
In this section we describe our procedure for generating synthetic NICER-like data for fully ionized hydrogen and helium atmospheres, and for fitting each dataset with both atmosphere models. We first describe the baseline configuration (§3.1) and then the
spot-to-background scaling experiment (§3.2), the synthetic data generation procedure (§3.3), the
likelihood and normalization treatment (§3.4), the two-stage analysis (§3.5), the fit diagnostics
(§3.6), the Bayesian evidence calculation (§3.7), and the radius bias diagnostic
(§3.8).
3.1. Baseline J0740-like Configuration
We adopt the two-uniform-temperature spot geometry and best-fit 12-parameter
solution from Miller et al. (2021) as our reference configuration. The parameters are the gravitational mass M, equatorial circumferential radius Re, observer inclination i, hydrogen
column density NH, source distance d, and the colatitudes, angular radii, and effective temperatures of the two hot spots. Table 1 lists these 12 parameters, the priors
Miller et al. (2021) used when they fit this pulse waveform model to the actual NICER and XMM-Newton data, and the best-fit values. As noted in Table 1, we adopted slightly wider priors for the spot angular radii ∆θ1 and ∆θ2 than were used by Miller et al. (2021). The purpose of choosing a specific reference configuration is not to exhaust all possible pulsar geometries, but to check the consequences of atmospheric composition mismatch in a realistic setting.
We set the total expected number of counts in the synthetic datasets to Ncounts ∼ 5.5×105, consistent with the
NICER data on PSR J0740+6620 used in Miller et al. (2021). The baseline ratio of hot-spot counts to background counts, R0 ≡ S0 /B0, reflects the background-dominated conditions of the actual NICER observations, in which only a few percent of the total counts originate from the pulsed hot-spot emission.
3.2. Spot-to-Background Ratio Experiment
To explore how the consequences of atmospheric composition mismatch depend on the relative strength
of the pulsed signal, we systematically vary the ratio of
hot-spot counts to background counts while holding the
total expected count number fixed. We define the multiplier m as
m≡
Rsynth
,
R0
, (1)
where Rsynth is the rescaled spot-to-background
counts S0 /B0 ratio, and R0 ≡ B0/S0 is the baseline ratio. The total num
ber of counts is:
Ncounts = Ssynth + Bsynth, (2)
where Ssynth and Bsynth are the rescaled total spot and background counts, respectively. We require their ratio to satisfy
Ssynth
= Rsynth. (3)
Bsynth
For clarity, the count normalization used in both generation and likelihood evaluations can be written explicitly. For each multiplier we define
am ≡
(m)
(0)
Ssynth, S0, (4)
bm ≡
(m)
(0)
Bsynth, B0. (4)
If Sjk and Bjk are the baseline expected spot and background counts in energy channel j and phase bin k, the
mean counts used to generate the synthetic data are
(m)
(0)
λjk = am Sjk + bm Bjk. (5)
The waveform calculation implements the spot scaling through the exposure time: for each synthetic dataset (m)
(0) we use Tobs = am Tobs. Because the predicted spot counts scale linearly with exposure, evaluation of the same 12 physical input parameters for that dataset predicts am S0 = Ssynth spot counts, rather than the original S0. Thus, the increased spot normalization is not supplied by an additional signal normalization parameter.
The factor bm is used to construct the synthetic background. The treatment of the background during fitting is described in §3.4.
By choosing m values 1, 20, 40, 60, 80, and 100, we explore configurations ranging from the realistic, heavily background-dominated regime of PSR J0740+6620 to strongly spot-dominated scenarios similar to the millisecond X-ray pulsar PSR J0437-4715. The high multiplier cases are controlled experiments and are not intended to represent PSR J0740+6620 itself with otherwise unchanged observational properties. A source with such a large pulsed fraction would generally require the
observational setup and background assumptions to be
reconsidered. We also explored higher multipliers such as 120, 140, 160 and 180, but these multipliers do not add significant information, given the already very high signal-to-noise.
3.3. Synthetic Data Generation
The hot-spot emission spectra and beaming patterns are computed using the NSX code for fully ionized hydrogen or helium
atmospheres (Ho & Lai 2001), following the same approach adopted in Miller et al. (2021). We note that the
surface gravity of a neutron star is large enough that the lightest element present is expected to settle to the top of the atmosphere within seconds to minutes (based on extrapolations of calculations in Alcock & Illarionov 1980).
As hydrogen is the most abundant element in the universe, and PSR J0740+6620 likely underwent prolonged accretion to reach its current spin frequency, a pure hydrogen atmosphere is a well-motivated assumption (Romani 1987; Bogdanov et al. 2019). However, the true
surface composition is uncertain. A helium atmosphere is possible if the accreted material was hydrogen-poor or if subsequent processes altered the composition. We therefore check both fully ionized hydrogen and helium atmospheres.
The NICER data for PSR J0740+6620 are recorded in 94 energy channels and are binned into 32 rotational phase bins. We refer to each of the 32 × 94 = 3008 combinations of a phase bin and an energy channel as a phase-channel bin. For the synthetic hydrogen data, we compute the expected phase-channel count distribution using the fully ionized hydrogen atmosphere model.
The baseline per-bin spot and background distributions are scaled to the new totals Ssynth and Bsynth from §3.2 without changing their shape, and we draw a Poisson realization of the counts in each phase-channel bin. For each multiplier m we perform an independent Poisson realization. The resulting dataset is then fit with both the hydrogen atmosphere model and the helium model. For the synthetic helium data, we repeat the procedure using the fully ionized helium atmosphere model to generate the expected counts. Each multiplier again corresponds to an independent Poisson realization. The resulting dataset is then fit with both the helium and hydrogen models. In both cases, we vary the ratio of hot spot
counts to background counts while holding the total expected count
number fixed at ∼ 5.5 × 105, which is the number
observed from PSR J0740+6620, in order to explore how
the detectability of the composition mismatch depends
on the relative strength of the pulsed signal. For each
configuration, we perform input-parameter evaluations,
full posterior sampling, goodness-of-fit assessments, and
Bayesian evidence comparisons.
We find that in both directions of the atmospheric
mismatch, the incorrect atmosphere model can yield statistically acceptable phase-channel and bolometric χ2
values after posterior exploration, yet still produce biased radius estimates. The effect is direction-dependent.
Fitting hydrogen-generated data with a helium model
can increase the inferred radius by more than 2σ, whereas fitting helium-generated data with a
hydrogen model produces smaller increases that remain within
2σ for the cases we examined. However, in these cases, the Bayesian evidence consistently identifies the
correct atmospheric model. Our findings reinforce
the importance of using the Bayesian evidence for
model comparison and goodness-of-fit tests.
The remainder of this paper is organized as follows:
Section 2 describes the effects of atmospheric beaming
of the emergent radiation on the pulse profile. Section 3
describes our methods for generating and analyzing the
synthetic data. Section 4 presents our results. We discuss the implications of these results in Section 5 and
summarize our conclusions in Section 6.
- ATMOSPHERIC BEAMING AND ITS EFFECTS
ON THE PULSE PROFILE
PPM is less susceptible to atmospheric uncertainty
than spectral fitting methods because the pulse waveform carries geometric information via phase-dependent
flux modulation. The time-varying projection of a heated region (“hot spot”) as the star rotates encodes the observer inclination, the spot colatitude, and
the stellar compactness through the relativistic lightbending that maps surface emission angles to the observer’s line of sight. Even if the beaming function is incorrect, the phase at which the spot appears and disappears and the overall depth of the modulation are primarily determined by the spot geometry and the star’s
compactness, reducing the freedom available for an incorrect atmosphere to absorb the mismatch. Indeed,
analyses of synthetic data have shown that PPM radius estimates are robust against several other classes of
systematic error. Lo et al. (2013) and Miller & Lamb
(2015, 2016) demonstrated that incorrect assumptions about the shapes and temperature distributions of the
hot spots do not significantly bias the inferred radius,
provided the fit is statistically acceptable. Holt et al.
(2025) showed that in joint NICER/XMM-Newton analyses, even a factor-of-five underestimate of the XMM-Newton background shifts the radius posterior by only
∼ 1σ. These results have strengthened the case for
Improvements for AI systems
Based on the systematic vulnerabilities identified in this paper regarding atmospheric composition mismatch in pulse-profile modeling, I propose the following specific improvements for high-stakes AI systems involved in scientific inference and parameter estimation:
- Implement
Model Adequacy
Layers via Bayesian Evidence Integration
The paper demonstrates that standard goodness-of-fit diagnostics (like Pearson's χ2) can fail to detect significant physical bias if the model is allowed enough parameter freedom to absorb
the error.
-
The improved AI system would not merely report a low residual value as a measure of success. Instead, it would perform an automated comparative analysis using Bayesian evidence (marginal likelihood) across a hierarchy of candidate models.
-
This system could detect when an AI is
overfitting
to the wrong physical model (e.g., assuming hydrogen when the data is helium), preventing the output of highly confident but physically incorrect parameters (like a biased neutron star radius).
- Asymmetric Bias Detection and Sensitivity Auditing
The paper reveals that errors are direction-dependent: assuming a helium atmosphere for hydrogen data causes massive radius bias, whereas the reverse does not.
-
The improved AI system would perform
Directional Stress Testing.
For any critical parameter estimation, the AI would systematically swap fundamental assumptions (composition, geometry, or background noise models) to map thebias landscape.
-
If the system detects that a specific modeling error leads to a non-linear or extreme shift in the output (as seen in Figure 6), it would flag the result as
High Risk/Low Reliability,
even if the statistical residuals are acceptable.
- Automated Uncertainty Quantization for
Hidden Systematics
Current AI models often provide a single best-fit value and a confidence interval that only accounts for statistical noise, not model inadequacy.
-
The improved AI system would implement
Systematic Uncertainty Envelopes.
It would calculate the variance in results caused by switching between different valid physical frameworks. -
Instead of reporting a radius of 16 km ± 0.5 km, the system would report:
Radius = 16 km ± 0.5 km (Statistical) ± 2.5 km (Model Systematic),
alerting human researchers that the result is highly sensitive to the underlying physical assumptions.
- High-Fidelity Signal-to-Noise (SNR) Contextualization
The paper shows that mismatches are hidden
in low SNR/high background regimes but become catastrophic in high SNR regimes.
-
The improved AI system would include a
Reliability Scaling Factor
based on the signal's dominance over the background. -
For high-fidelity data (e.g., future X-ray missions), the AI would automatically trigger more rigorous model comparison protocols, recognizing that higher precision in measurement actually increases the danger of being misled by an incorrect physical model.
Abstract
Constraints on neutron star radii provide insight into the properties of the cold, dense matter in their interiors. Previous studies using synthetic Neutron star Interior Composition Explorer (NICER) pulse waveform data have demonstrated that radius inferences derived therefrom are robust against several classes of modeling systematics. Here we explore the consequences of assuming the wrong atmospheric composition, using synthetic data based on the about 2.1 M pulsar PSR J0740 + 6620. We find that the assumption of a hydrogen atmosphere when the synthetic data assumed a helium atmosphere, or vice versa, produces little bias in the inferred radius at the spot-to-background ratio of the actual PSR J0740 + 6620 data. However, when we increase the spot-to-background ratio by a factor of about20 while keeping the total number of counts fixed at the observed about 5.5 times 10 5, we find that composition mismatch can produce significantly biased radius estimates while remaining hidden inside a fit with statistically acceptable residuals. Even in these cases, the Bayesian evidence consistently identifies the correct atmospheric model. Our findings reinforce the importance of using the Bayesian evidence for model comparison and goodness-of-fit tests.
Sources
- A NICER view of the millisecond pulsar PSR J2124 - 3358: evidence for a helium atmosphere
- An Investigation of Systematic Effects from Background Priors on PSR J0740$+$6620 Radius Estimates using Synthetic NICER and XMM-Newton Data
- A Lower Mass Estimate for PSR J0348+0432 Based on CHIME/Pulsar Precision Timing
Related papers
- Numerical Studies of Accretion Flows onto a Neutron Star Engulfed in a Massive Star
- Collisionless Accretion of Finite-Angular-Momentum Plasma onto a Spinning Black Hole
- Impact of Magnetic Field Topology on Electromagnetic and Gravitational Waves from Binary Neutron Star Merger Remnants
- XRISM Resolve Spectroscopy of GX 5-1: Constraints on Iron Spectral Features in a Luminous Neutron-Star Binary
- SN 1006: A Cosmic Laboratory for Investigating Shock Acceleration Physics
- Neutrino Spectral Pinching in 3D Core-Collapse Supernovae: Late-Time Convergence, Failed-Explosion Signatures, and Viewing-Angle Dispersion