page_by_page
This episode discusses a paper by Carteret and Bourrier on observing exoplanet atmospheric escape via metastable helium. The hosts explain that comparing low-resolution JWST data to models introduces bias if stellar lines are ignored, and that NIRPS, NIRISS, and NIRSpec have complementary strengths for detecting and characterizing helium outflows.
Introduction to the show: ident: Astrophysics Radio.
Vera: Next we'll be talking about the paper "Toward a unified framework for helium observations and interpretation of atmospheric escape.".
Jocelyn: The paper was written by Yann Carteret and Vincent Bourrier from Observatoire Astronomique de l’Université de Genève.
Vera: Stay tuned as we take you through the paper and discuss its implications.
Paper summary: Vera: Welcome back, everyone. Today we're looking at a new paper from Yann Carteret and Vincent Bourrier at the Geneva Observatory, accepted for publication in Astronomy and Astrophysics. It's a systematic study of how we observe atmospheric escape from exoplanets using the metastable helium triplet at around ten thousand eight hundred thirty-three angstroms — and, just as importantly, how we compare those observations with theoretical models.
Jocelyn: The central message, I think, is that there's a subtle but serious mistake lurking in the standard way people have been comparing low-resolution observations — mostly from JWST — with model predictions. If you take a theoretical transmission spectrum and simply convolve it with the instrument profile, you introduce a bias that grows as the resolution drops. The paper demonstrates this clearly, and then goes on to map out which instruments are actually best suited for which scientific questions.
Subrahmanyan: The comparison is between ground-based high-resolution spectrographs like NIRPS and JWST's two low-resolution modes, NIRSpec and NIRISS. The interesting result is that NIRPS and NIRISS end up being sensitive to more or less the same helium signatures, while NIRSpec — with its intermediate resolution — extends the detection limit to fainter targets. So the story is not simply that space beats the ground, or that higher resolution is always better. The two approaches are genuinely complementary.
Vera: And that complementarity is the heart of the paper. Low-resolution space-based observations are best for measuring the spatial extent of outflowing tails and for getting long, uninterrupted time coverage, while high-resolution ground-based observations are what you need to probe the actual dynamics of the gas close to the planet.
Jocelyn: They build the whole study on careful simulations of WASP-69 b, one of the best-studied systems in helium, and then generalise to different outflow geometries — from spherical thermospheres to dense streams produced by Roche-lobe overflow. It's a methodological paper, but it has very concrete implications for how future observing programs should be designed.
Subrahmanyan: Indeed, given how many JWST programs are now routinely picking up helium in transmission — sometimes as a secondary product of transit observations — this is exactly the right moment for a careful look at how we interpret those data.
Vera: Let's start at the beginning, then — with the introduction and the historical context that motivates the whole study.
Page 1 of the paper: Vera: The paper opens with the history of how we came to observe atmospheric escape at all. Escape is thought to be a key process shaping exoplanet demographics — it can strip a sub-Neptune down to a bare core, or transform a hot Jupiter over time. But observing the escaping gas directly is hard, and the first detections were only possible from space, through the Lyman-α line of hydrogen, back in 2003 with Vidal-Madjar and colleagues.
Jocelyn: Lyman-α is a wonderful tracer because hydrogen is everywhere in these outflows, but it sits in the ultraviolet, which is blocked by Earth's atmosphere. So for years, the only window into atmospheric escape was a space-based one, and a narrow window at that. That changed with the metastable helium triplet: helium in its metastable state absorbs at ten thousand eight hundred thirty-three angstroms, in the near-infrared, which means ground-based telescopes can see it.
Subrahmanyan: And there's a beautiful physical reason why it works so well. The metastable state decays back to the ground state within a few hours, so the helium you detect has to be freshly produced in the extended upper atmosphere — it traces the hydrodynamical regime of the outflow, the thermosphere, where the gas is being actively accelerated. It's not a long-lived tracer like the hydrogen exosphere; it's a snapshot of the escape region itself.
Vera: So the ground-based community jumped on it with high-resolution spectrographs that can resolve the individual lines of the triplet — instruments like CARMENES, ESPRESSO, and NIRPS. And then JWST arrived, and suddenly people were detecting helium at low resolution, sometimes around small planets where ground-based searches had come up empty, and sometimes catching extended outflows that ground-based observations had missed entirely because the absorption lasted longer than the night.
Jocelyn: That's the tension the paper wants to resolve. You have two very different strategies — high-resolution ground-based spectroscopy and low-resolution space-based spectroscopy — both observing the same helium signature, but the community has never really established a unified framework for comparing them.
Subrahmanyan: And the key thing, as the paper states, is that these two strategies probe different aspects of the outflow: high resolution gives you the line cores and the velocity structure, while low resolution gives you the broad spectral shape and the temporal extent. But there's a catch in how you model the data, and that catch is where the methodology comes in.
Vera: Exactly. To test all of this properly, the authors had to build very careful synthetic observations — which is where the paper goes next.
Page 2 of the paper: Vera: So now we're into the machinery of the paper. The authors generated mock observations using two codes: EvE, which simulates the transit in full three-dimensional geometry, and p-winds, which computes the thermospheric density and temperature profiles. Everything is anchored to WASP-69 b, a hot Saturn that is essentially the gold standard for helium studies — it has one of the highest signal-to-noise detections known, around sixty-five, and had been predicted theoretically to be the best system for this kind of measurement.
Jocelyn: And the crucial thing about EvE is that it treats the star properly. It doesn't just use a disk-integrated spectrum; it builds a two-dimensional grid of local spectra across the stellar surface, accounting for center-to-limb variations, limb darkening, and the Rossiter-McLaughlin effect. These are the so-called POLDs — planet-occulted line distortions — and they can create features in the transmission spectrum that masquerade as atmospheric absorption if you don't model them.
Subrahmanyan: The thermosphere is generated with p-winds using a helium-to-hydrogen ratio of ninety to ten, and the exosphere is treated as a Monte Carlo particle simulation shaped by the stellar XUV flux. Even the tails are parametrized — tube-like structures in the orbital plane, defined by an escape velocity and an angle. So the models cover the full chain from the deep thermosphere out to the escaping gas.
Vera: Then comes the instrumental step. They convolve the simulated flux time series with a Gaussian profile matching each instrument — NIRPS at a resolving power of seventy-five thousand, NIRSpec at about two thousand, and NIRISS at six hundred fifty — and resample onto the actual wavelength grids. And here is where the paper makes a seemingly trivial choice that turns out to matter enormously: they compute the absorption as one minus the ratio of the in-transit to the out-of-transit spectra, both processed identically.
Jocelyn: That's equation one in the paper, and the reason it matters is that the out-of-transit spectrum is the star. The star has its own absorption lines, and when you convolve the flux ratio, the stellar lines and the planetary lines interact in a nonlinear way. If you instead convolve the absorption spectrum directly — which is what many model codes do — you're implicitly assuming the stellar continuum is flat.
Subrahmanyan: And we should mention that the error bars are realistic too. For NIRPS, they calibrate against actual observations from the instrument's guaranteed time program, which includes real weather and sky conditions, while for JWST they use Pandexo and the exposure-time calculator. So the mock data are not idealized noise; they reflect what observers would actually get.
Vera: With that machinery in place, the paper can now ask the question at the core of the study: what happens when you compare models to data the way the literature typically does? And as we're about to see, the answer at low resolution is a systematic bias.
Page 3 of the paper: Vera: So the paper now demonstrates the bias in a very concrete way. They take a synthetic helium absorption signal — a Gaussian with a peak of three percent and a full width at half maximum of zero point seven angstroms, typical values for atmospheric escape — and they compute what NIRSpec would actually observe, properly, using that ratio of convolved fluxes, with two different stellar spectra: WASP-69, which has deep and narrow stellar lines, and WASP-121, which has broader and shallower lines.
Jocelyn: Then they fit the resulting mock observation using the standard literature approach — directly convolving the absorption spectrum without including the star. For WASP-69, they recovered a peak absorption of two point five two percent instead of the true three percent. That's a sixteen percent underestimate, and it's about four sigma at the peak of the absorption. For WASP-121, the bias was much smaller — two point nine three percent, nearly consistent with the input.
Subrahmanyan: The physical explanation is quite elegant. At low resolution, the instrumental profile averages the stellar flux over a broad wavelength range. When the star has deep, narrow lines, the planetary absorption near the line center produces a smaller relative change in flux, because you're comparing against a deep stellar line rather than a smooth continuum. So the apparent absorption is diluted. Directly convolving the absorption spectrum ignores this entirely and overestimates the signal.
Vera: And the deeper point is that this bias can create apparent inconsistencies between ground-based and space-based measurements — which could easily be misinterpreted as atmospheric variability. The authors also note that the effect becomes worse when the helium line is broader than the assumed zero point seven angstroms, which happens in extended outflows where dynamical broadening dominates, or when the deepest atmospheric layers saturate.
Jocelyn: So that's for narrow atomic lines. But most JWST low-resolution science is about broadband molecular features, so they checked whether the same bias affects those. They simulated a full transmission spectrum of WASP-69 b with the SCARLET code and computed the NIRISS observation three different ways: properly using the stellar spectrum, resampling only, and convolution plus resampling.
Subrahmanyan: And there the differences were much smaller — of order fifty to two hundred fifty parts per million — because molecular bands are broad blends of many lines, wider than both the instrumental kernel and the stellar lines. But the authors still recommend that retrieval frameworks at least apply the instrumental convolution, and ideally compute the model exactly as the data are computed, using the stellar spectrum.
Vera: So the bias is quantified, the mechanism is understood, and the prescription is clear. The next question is: given all this, how do the three instruments actually compare when it comes to detecting helium in the first place?
Page 4 of the paper: Vera: For the detection question, the paper uses a clean and simple metric: the equivalent width of the helium absorption divided by its uncertainty, with a detection threshold of three sigma. They generated absorption signatures of different amplitudes, placed them around stars of different brightness — encoded in the J magnitude — and computed the signal-to-noise for a single transit with each instrument.
Jocelyn: And the result is laid out in those detection maps. The striking finding is that NIRPS and NIRISS end up with essentially the same detection thresholds, even though NIRISS has a vastly larger collecting area and much finer time sampling. The reason is the convolution: at NIRISS resolution, the helium absorption gets smeared over very broad pixels, the peak signal is diluted, and that cancels out the sensitivity advantage.
Subrahmanyan: NIRSpec, however, sits in a sweet spot. Its resolving power of about two thousand is low enough to be stable and space-based, but high enough to preserve the helium signal. It also avoids a specific problem that NIRISS has: there is a stellar silicon line only about three angstroms from the helium triplet, and at NIRISS resolution — with a kernel of roughly seven angstroms — the two are blended, which suppresses the apparent helium absorption. NIRSpec separates them cleanly.
Vera: The maps also define three regimes, which the authors label A, B, and C. Case A is high absorption and high signal-to-noise; case B is low absorption and high signal-to-noise; case C is high absorption but a faint target. And in that third case, NIRSpec is the only instrument that reaches a clear detection. That has real consequences for planning: if you want to push helium measurements toward fainter planets, NIRSpec is the tool.
Jocelyn: There is also a subtle point about the continuum — the contribution from the planet's opaque body, which the helium absorption is measured against. Both JWST instruments constrain that continuum comfortably for essentially all targets brighter than magnitude fourteen, whereas from the ground with NIRPS it becomes poorly constrained for faint planets. But for the helium line itself, the high spectral resolution of NIRPS compensates for the small three-point-six-meter telescope aperture.
Subrahmanyan: So the detection story is nuanced: NIRPS and NIRISS are roughly equivalent, NIRSpec is somewhat better — especially for faint targets — and the sensitivity alone does not tell you which instrument gives you the best physical constraints. That's the next step in the paper.
Vera: Exactly — moving from detection to characterization. And to do that, the authors simulate four very different outflow structures and ask what each instrument can actually learn from them.
Page 5 of the paper: Vera: So now we move to characterization. The four outflow scenarios are inspired by hydrodynamical simulations. Three of them represent outflows confined by stellar winds of increasing strength — weak, medium, and strong — following the models of MacLeod and Oklopčić. The fourth is a more extreme case: a dense stream formed by Roche-lobe overflow, in which the planet is effectively spilling material directly into space.
Jocelyn: When you look at the time-averaged absorption spectra at high resolution, these four cases are clearly distinguishable. The stream case shows two distinct peaks — a blueshifted one from the trailing tail and a redshifted one from the leading tail, because the gas is no longer comoving with the planet. The medium case, with a very extended thermosphere, shows substantially broader wings, because different regions along the line of sight have very different projected velocities.
Subrahmanyan: At low resolution, however, most of that dynamical information vanishes. The weak, medium, and strong wind cases become essentially indistinguishable in the NIRISS spectrum — the thermal broadening is simply insufficient to separate them. Only the stream case shows a slight difference in the relative amplitudes of the spectral bins. So low-resolution observations are nearly blind to the kinematics of the outflow.
Vera: But then the light curves tell a complementary story. The helium light curves from JWST have much finer temporal sampling and, crucially, continuous coverage across the entire transit, including before and after. From the ground, you're limited by the night, and in the stream case the helium absorption can last longer than the night itself. The paper is quite blunt about this: a typical ground-based observing window would simply miss the stream structure entirely.
Jocelyn: There's also the perennial problem of the baseline — the reference spectrum you divide your in-transit observations by. From the ground, with telluric contamination and instrumental instabilities, identifying where the absorption actually starts and ends is genuinely hard. JWST sidesteps most of that, and the paper points out that JWST has already shown its value by re-observing known systems and catching pre- and post-transit absorption that ground-based campaigns had missed.
Subrahmanyan: So a clear division of labor emerges: high resolution for the dynamics and velocity structure, low resolution for the spatial extent and duration of the outflow. Both are needed for a complete picture — which, naturally, brings us to the question of the baseline itself and the dangers of getting it wrong.
Vera: Yes — and this is a subtle issue that can undermine both strategies if you're not careful.
Page 6 of the paper: Vera: This section addresses something that sounds mundane but is actually a trap: the choice of the baseline, the reference spectrum that defines the unabsorbed star. In principle, the baseline should be pure starlight. In practice, for extended outflows, the helium absorption can already be present in the exposures you are using as a baseline — especially from the ground, where you're constrained by the night and by the cadence.
Jocelyn: The paper derives the effect mathematically. If the baseline exposures contain some absorption, then the reference spectrum is contaminated, and that contamination propagates through the entire absorption time series. The measured absorption at every time is biased downward — and because the optical depth is wavelength dependent, the bias is not a constant scaling. It distorts the shape of the absorption spectrum itself, which is far more dangerous.
Subrahmanyan: They illustrate this with their four outflow models, constructing biased baselines from exposures two to three hours before or after mid-transit. For the fully spherical, compact atmosphere, the baseline is clean and nothing changes. But for the extended cases — especially the stream — the averaged absorption spectrum is visibly underestimated, and the distortion depends on the line-of-sight velocity of the escaping gas. If you compared such a spectrum to models without accounting for this, you would infer the wrong mass-loss rate and the wrong geometry.
Vera: The practical warning is stark: in the stream case, a ground-based observing window from three hours before to three hours after mid-transit would completely miss the stream structure. The recommendation is to cover the longest possible baseline, and ideally to coordinate with space-based observations to anchor the true out-of-transit level.
Jocelyn: And then the conclusions tie everything together. At high resolution, the stellar distortions — the POLDs — need to be modeled once your signal-to-noise is high enough, and the paper expects that bias to be stronger for fast rotators and cool host stars. At low resolution, the convolution bias must be avoided by computing the model as a ratio of convolved fluxes, using the stellar spectrum. Instrument-wise, NIRPS generally gives tighter constraints on temperature and mass-loss than the JWST instruments, NIRSpec is competitive for faint targets, and the two approaches are genuinely complementary.
Subrahmanyan: The paper also explicitly notes that its conclusion differs from an earlier study by Dos Santos and colleagues, which had argued that JWST provides tighter constraints on escape parameters. The difference comes down to methodology — and in particular, to properly accounting for the stellar spectrum when computing the models.
Vera: It's a strong ending for a paper that's really about the craft of measurement as much as the physics. Let's wrap up our discussion now.
Conclusion: Vera: So, stepping back: this paper gives the helium community something it badly needed — a unified way to think about high- and low-resolution observations of atmospheric escape. The core message is twofold. First, the way you compare models to data matters, and at low resolution, ignoring the stellar spectrum biases your results by more than the noise. Second, the instruments are not in competition; they do different jobs.
Jocelyn: High-resolution ground-based spectroscopy is irreplaceable for the dynamics — the velocity shifts, the line shapes, the distinction between different wind-confinement geometries. Low-resolution space-based spectroscopy is irreplaceable for the long, uninterrupted time coverage that lets you measure the extent of the outflow and catch absorption that extends well beyond the optical transit.
Subrahmanyan: And the practical guidance is clear: compute your model the same way you compute your data, choose your baseline with care, and whenever possible, observe simultaneously from the ground and from space. For a field that is now producing a steady stream of JWST helium detections, that guidance is genuinely timely.
Vera: For me, the most striking number is that a three percent helium signal, seen with NIRSpec, would be underestimated by sixteen percent under the standard modeling assumption — enough to create a several-sigma inconsistency with a ground-based measurement. That's not an esoteric detail. It's the difference between interpreting a mismatch as atmospheric variability and recognising it as a methodological artifact.
Jocelyn: And the fact that WASP-69 b — one of the best-characterised helium systems in the sky — was used as the testbed gives the results extra weight. If the bias matters for WASP-69 b, it certainly matters for fainter systems.
Subrahmanyan: I'm also glad they included those four outflow geometries. The weak, medium, and strong stellar-wind cases, plus the Roche-lobe overflow stream, give us a useful vocabulary for thinking about what these outflows actually look like — which is far from the simple spherical picture that older models assumed.
Vera: And with that, we'll say goodbye to this paper. It's a methodological cornerstone for helium observations, and I suspect it will be cited in every future observing proposal that mixes ground-based and space-based data. Thank you both for the discussion.
Jocelyn: Thanks, everyone, for listening. Next time, we'll pick up a new paper and see what else the universe has in store.
Subrahmanyan: Until then, keep looking up.