Agentic AI for Gravitational Wave Data Analysis: A Head-to-Head Comparison of Coding Agents Executing a Matched Filter Pipeline on Einstein Telescope Simulated Data
astro-ph.IM, cs.AI, cs.HC
Submitted: 2026-05-27
Updated: 2026-09-13
Comments: Version accepted for publication in IOP Physica Scripta, online abstract version shortened to satisfy arXiv requirements
License: http://creativecommons.org/licenses/by/4.0/
The gist: We report a methodological study of agentic AI in gravitational-wave data analysis: two systems, Claude Code (Anthropic) and Codex (OpenAI), autonomously executed the same simple end-to-end pipeline
Terminology
Abstract
We report a methodological study of agentic AI in gravitational-wave data analysis: two systems, Claude Code (Anthropic) and Codex (OpenAI), autonomously executed the same simple end-to-end pipeline on Einstein Telescope (ET) simulated data, on shared infrastructure and without human intervention. The object of study is the behaviour, reliability and auditability of the agents, not the physics output, used here as a controlled test case. The pipeline comprises power spectral density estimation from simulated ET noise, geometric template bank generation with IMRPhenomD waveforms, matched-filter recovery of 100 binary black hole injections, results generation, and LLM-assisted production of a LaTeX manuscript in Physical Review D style. Both agents received identical specifications and resources. The experiment was run twice: first with unrealistically loud injections, then with signals rescaled to a physically motivated SNR range. In both runs the results converged, with comparable detection efficiency and template bank size. The agents, however, behaved very differently: Claude Code finished in about 3.4 minutes with silent deviations from the specification, while Codex needed about 16 minutes across explicit self-correcting restarts, including an unsolicited optimization of the matched-filter inner loop. In the second run, a subtle difference in interpreting the SNR-range instruction produced a genuine scientific divergence: Claude Code silently raised the SNR floor to 8 (100% efficiency), while Codex followed the specification literally down to SNR 7 and recorded one missed detection. We discuss the implications - speed versus auditability, silent deviation versus explicit self-correction, instruction interpretation, and intermediate data representations in multi-model pipelines - for agentic AI in scientific workflows, within the limits of a single-pipeline, two-run benchmark.
Sources
- AI Agents Can Already Autonomously Perform Experimental High Energy Physics
- A mock data challenge for next-generation detectors
- The GstLAL Search Analysis Methods for Compact Binary Mergers in Advanced LIGO's Second and Advanced Virgo's First Observing Runs
- Designing a template bank to observe compact binary coalescences in Advanced LIGO's second observing run
Related papers
- A signal dedispersion algorithm for imaging-based transient searches
- AVICA: A fully automated CASA pipeline for large volume VLBI data calibration
- Spectral Map Making with SPHEREx
- Long-Integration Magnetar Burst Observatory (LIMBO): Instrument Summary and Early FRB Rate Constraints
- Towards independent event horizon imaging of the supermassive black holes in M87 and the Milky Way
- A PINK update: Improvements to the CELEBI fast radio burst data reduction and analysis pipeline