AVICA: A fully automated CASA pipeline for large volume VLBI data calibration

arXiv:2604.17448 · astro-ph.IM, astro-ph.GA · Submitted 2026-04-19 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Astrophysics Radio. Generated commentary on the latest astrophysics papers.

Vera: Today's paper: "AVICA: A fully automated CASA pipeline for large volume VLBI data calibration".

Jocelyn: Calibrating large volumes of Very Long Baseline Interferometry (VLBI) data is traditionally a time-consuming process requiring significant human intervention, but this work introduces AVICA,

Vera: First, who's behind it and why it matters.

Paper summary: Vera: We’ve discussed how AVICA automates the calibration of large volumes of Very Long Baseline Interferometry data, focusing on handling heterogeneous formats and automated antenna selection to remove manual input from the process. The authors have presented this fully automated CASA pipeline as a way to make large-scale VLBA data processing accessible.

Jocelyn: It’s clear that the work by A. Kumar et al., titled "AVICA: A fully automated CASA pipeline for large volume VLBI data calibration," aims to address the practical limitations of current pipelines, particularly concerning the manual intervention needed when dealing with decades of evolving correlator formats and file structures.

Subrahmanyan: From a theoretical perspective, the significance lies in proving that complex calibration tasks can be managed robustly across highly varied datasets without needing bespoke configurations for every single run.

Vera: It’s about taking a traditionally time-consuming process and making it runnable for a wider community of researchers who might not have the expertise to manually tune all those parameters.

Jocelyn: I think the implication is that we can start leveraging much larger, older VLBA archives more effectively than previously possible because the barrier to entry for processing them has been significantly lowered.

Subrahmanyan: This capability means we can conduct more comprehensive studies on phenomena like Supermassive Compact Objects using these diverse datasets, as it removes calibration as a primary bottleneck in our analysis.

Vera: So, in simple terms, AVICA provides a tool that handles the heavy lifting of calibration automatically so we spend less time babysitting the software and more time looking at the actual astronomical signals.

Jocelyn: And that’s what makes it impactful for pulsar and sky surveys too; having reliable, automated pipelines means we can deliver cleaner results from those long-term observations more consistently.

Subrahmanyan: Ultimately, this paper demonstrates a way to scale VLBI data calibration using existing tools like CASA in a way that doesn't require extensive manual input at every stage of the workflow.

Conclusion: Vera: So, we've just been walking through the technical details of AVICA, and now we need to wrap up by talking about what this paper actually is and why it matters for us as a community.

Jocelyn: I think focusing on the title, "AVICA: A fully automated CASA pipeline for large volume VLBI data calibration," really gets to the heart of what they did—it’s about taking a massive manual chore and turning it into something that runs itself.

Subrahmanyan: From my end, it speaks to the feasibility of applying standard software like CASA to tackle datasets that were previously too big or too messy for routine analysis. It shows a path forward for handling the sheer volume we're seeing in modern VLBI surveys.

Vera: Exactly, and when you look at the authors, they’ve managed to integrate several complex components—Python libraries, workflow managers like ALFRD—into one cohesive system that handles everything from file formats to source selection automatically.

Jocelyn: That automation aspect is what really excites me; it means we can process archival data from decades ago without needing a specialist just to set up the initial calibration parameters for every single run.

Subrahmanyan: That level of automation has real implications for theoretical work because it lowers the barrier to entry for using these deep historical datasets, which feeds into our models about galactic structure and compact objects.

Vera: It really means that the bottleneck isn't just having a big telescope or a large archive anymore; it’s about how efficiently we can extract and calibrate that data, and AVICA seems to solve that extraction problem.

Jocelyn: For pulsar surveys specifically, this pipeline suggests we can get much more consistent results across different observational epochs because the calibration steps are performed identically every time.

Subrahmanyan: If we can reliably process these heterogeneous datasets at scale, we gain a much richer statistical sample to test our astrophysical theories about how matter behaves in extreme environments.

Vera: So, looking ahead, I think the real impact of AVICA is making large-scale VLBI analysis something that the general research community can actually do without needing a dedicated calibration team for every project.

Jocelyn: It’s a big step toward democratizing high-quality VLBI data processing; we're moving away from highly customized manual pipelines toward robust, automated systems.

Subrahmanyan: This capability means we can finally analyze those faint, distant sources more thoroughly because the calibration noise isn't introducing systematic errors due to human input variability anymore.

Vera: And that’s the big picture—we’re not just calibrating images; we’re unlocking the potential of vast historical radio observations in a much more practical way.

A. Kumar, C. Casadio, M. Janssen, D. Álvarez-Ortega, F. M. Pötzl

Institute of Astrophysics, Foundation for Research and Technology – Hellas · Department of Physics, University of Crete · Institute for Mathematics, Astrophysics and Particle Physics (IMAPP), Radboud University

astro-ph.IM, astro-ph.GA

Submitted: 2026-04-19

Updated: 2026-09-28

Comments: 12 pages, 9 figures; Published in Astronomy & Astrophysics. Code available at https://github.com/avikhagol/avica

Journal ref: A&A, 713 (2026) A287

DOI: 10.1051/0004-6361/202660469

Code: https://github.com/avikhagol/alfrd

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 92/100

The gist: Calibrating large volumes of Very Long Baseline Interferometry (VLBI) data is traditionally a time-consuming process requiring significant human intervention, but this work introduces AVICA, a fully

Key concepts

AVICA
A fully automated CASA pipeline that handles the entire calibration workflow for VLBA data. It automates complex tasks like data preprocessing, source selection, and calibration, allowing users to process large datasets without needing to manually input parameters.
rPICARD
The core calibration framework that AVICA extends. It executes a sequence of steps including amplitude calibration, fringe-fitting for instrumental effects, atmospheric phase correction using multi-band fitting, and complex bandpass calibration.
ALFRD
Automated Logical Framework for executing Dynamic scripts. This in-house module manages the scheduling and workflow orchestration of AVICA. It allows the pipeline to run dynamically with minimal dependency on the CASA software stack, serving as the backend for automated execution.
Blind Calibration
The ability to calibrate data without manual input of parameters like calibrator sources or reference antennas. AVICA achieves this by automatically ranking and selecting optimal calibrators and reference antennas based on signal-to-noise ratios and antenna availability.

Terminology

Summary

Calibrating large volumes of Very Long Baseline Interferometry (VLBI) data is traditionally a time-consuming process requiring significant human intervention, but this work introduces AVICA, a fully automated CASA pipeline designed to automate the calibration of VLBA datasets for projects like SMILE. This pipeline is significant because it enables fully blind calibration of heterogeneous archival VLBA data without manual parameter input, making large-scale processing practical for community use.

The gist: AVICA demonstrates that fully blind calibration of heterogeneous archival VLBA data is achievable using CASA, without manual parameter input.

Pipeline Architecture and Components

AVICA extends the existing CASA-based rPICARD calibration framework by automating several critical steps in the VLBI workflow. The pipeline is designed with two levels of usability: a fully automated end-to-end pipeline utilizing ALFRD as the workflow backend, and individual components that can be invoked independently for experienced users to maintain control over specific functionalities like calibration methodology or preprocessing steps.

The core of the system relies on several integrated modules:

  1. AVICA Python package, which provides an extensive library for manipulating and inspecting FITS-IDI and Measurement Set (MS) file formats using libraries like CFITSIO and Astropy.

  2. ALFRD (Automated Logical Framework for executing Dynamic scripts), which serves as the scheduling and workflow management module, developed in-house to provide orchestration with minimal dependencies on the CASA software stack.

  3. rPICARD, which acts as the core calibration framework that AVICA extends with automated preprocessing and selection steps.

Data Preprocessing Workflow

The preprocessing stage is designed to handle the heterogeneity of archival VLBA data spanning three decades, addressing issues like evolving correlator formats and file structures. This stage operates on both FITS-IDI and Measurement Set (MS) data formats to reduce data volume before calibration begins. Key preprocessing steps include:

(See Fig. 1 for the complete workflow)

  1. Handling format inconsistencies within the FITS-IDI files, such as binary data in ASCII table extensions and removing duplicate antenna entries. This is managed by the AVICA function, which corrects these issues automatically based on AIPS Memo 1145 standards.

  2. Reducing data loading time for very large files (e.g., those exceeding 100 GB), where only required calibrator and target sources are extracted prior to loading, triggered when source count exceeds 50 and file size is above 100 GB, using the fitsidiutil module.

  3. Cross-matching the sources present in the file against the Radio Fundamental Catalogue (RFC) or a VLBA calibrator list to identify potential calibrators within 100 milliarcseconds of catalogued entries. The top 10 matched sources, ranked by flux density, are retained alongside the target source.

  4. Automatically downloading missing gain curve or system temperature data from the NRAO archive if they are absent from the FITS-IDI file, and generating a flux calibration table in ANTAB6 format to be appended to the FITS-IDI files using the casa-vlbi7 package.

Automated Selection of Calibrators and Reference Antennas

A crucial feature of AVICA is its automated selection process for calibrators and reference antennas, which is designed to eliminate manual parameter input.

  1. The pipeline ranks calibrators and reference antennas automatically using signal-to-noise ratio (S/N) from the FFT-based fringe detection.

  2. The reference antenna selection prioritizes an antenna that balances a high median S/N across all baselines with proximity to the geometric center of the array, while also satisfying conditions like remaining unflagged for the majority of calibrator scans and being available throughout science target scans.

  3. A ranked list of reference antennas is provided, allowing rPICARD to re-reference solutions to a substitute antenna as needed via the refantmode='flex' parameter in the fringefit task, ensuring continuity in fringe-fit solutions without gaps or discontinuities.

  4. The pipeline defaults to selecting the top 5 sources as calibrators, ranked by their median FFT S/N across all baselines and scans, subject to antenna availability constraints.

Calibration Workflow Execution

The calibration stage follows the full rPICARD scheme, executing in sequence amplitude calibration, single-band fringe-fitting for instrumental effects, multi-band fringe-fitting for atmospheric phases, complex bandpass calibration using cross-correlations, and finally phase calibration via a multi-band fringe fit.

  1. Sampler corrections are applied using autocorrelations to account for amplitude errors introduced by digitization in the correlation process.

  2. A scalar bandpass correction is performed to fix the shape of the passband in amplitude as a function of frequency using auto-correlations.

Improvements for AI systems

As a fastidious and diligent AI researcher, I see AVICA as a powerful proof-of-concept for automating complex, heterogeneous data pipelines in radio astronomy. The core innovation lies in its ability to execute a multi-stage calibration workflow—spanning data preprocessing, automated calibrator/reference antenna selection, and the full rPICARD framework—without manual intervention across massive datasets.

Here are specific improvements and capabilities that can be derived from the AVICA system for AI/ML applications:


)

  1. Automated Pipeline Optimization via Reinforcement Learning (RL):

This is a direct extension of AVICA's automated workflow management (ALFRD).

  • The RL agent could be trained to dynamically adjust the pipeline parameters—such as the selection criteria for calibrator ranking (e.g., changing the FFT S/N threshold, or the proximity constraint for reference antenna selection)—based on real-time feedback from previous runs.

  • The goal would be to maximize successful calibration completion or S/N detection percentage while minimizing execution time, effectively optimizing the pipeline's configuration on-the-fly for a new data batch.

  1. Adaptive Data Format Handling and Feature Extraction:

AVICA already handles heterogeneous formats (FITS-IDI vs. MS) and performs automated preprocessing (source extraction).

  • An AI system could be trained to automatically detect and apply format-specific patches or transformation rules when encountering novel, uncatalogued data structures from new arrays, going beyond the current explicit checks in AVICA's preloading step.

  • It could perform advanced feature engineering on the raw visibility data (e.g., predicting instrumental noise characteristics based on ancillary metadata) to inform the subsequent calibration steps before they are even executed.

  1. Predictive Failure Analysis and Error Mitigation:

The paper notes that 22 datasets failed due to corrupted or incomplete input data, and 15 failures stemmed from metadata inconsistencies.

  • An AI model could be trained on the failure signatures (e.g., specific combinations of missing system temperature records, unusual file corruption patterns) to predict which future archival datasets are likely to fail before the pipeline even starts processing them.

  • This predictive capability would allow researchers to prioritize high-risk data for manual inspection or pre-emptive data cleaning, significantly reducing wasted computational time on doomed runs.

  1. Automated Calibration Quality Assessment and Benchmarking:

AVICA provides metrics like fringe solution percentage and compares results against other pipelines (VIPCALs).

  • An ML model could be trained to ingest the calibration outputs (e.g., the final S/N distributions shown in Fig. 5 and Table 2) from a large library of historical datasets, learning to predict the quality metric (like fringe solution percentage) based on input data characteristics (flux density, observation length, frequency band).

  • This would create a Quality Predictor that provides an immediate, automated confidence score for any new batch of data before committing to the full calibration run.

  1. Self-Correcting Fringe Fitting Algorithms:

The paper relies on a sequence of fixed steps (scalar bandpass correction, amplitude calibration, fringe fitting).

  • A deep learning approach could be used to learn the optimal sequence or combination of these steps for a given dataset's noise characteristics. For instance, if an observation is dominated by atmospheric phase noise, the AI could dynamically prioritize the multi-band fringe-fitting step over the scalar bandpass correction step, adapting the calibration strategy to maximize accuracy for that specific data point.

The improved AI system would be a Cognitive VLBI Calibration Engine capable of:

  1. Automatically optimizing its own configuration parameters (thresholds, reference selection) using RL to achieve maximum calibration success rate and speed.

  2. Intelligently interpreting and adapting to novel or corrupted archival data formats autonomously.

  3. Predicting which datasets will fail during processing based on input metadata patterns, allowing for proactive error management.

  4. Providing an instant, data-driven quality assessment score for any calibration output before human review, effectively acting as a highly sophisticated pre-processor and validator for scientific results derived from VLBI data.

Abstract

Calibrating large volumes of Very Long Baseline Interferometry (VLBI) data traditionally requires significant human intervention at every stage. While the Common Astronomy Software Applications (CASA) package is the standard data reduction tool across major radio observatories, no existing CASA-based pipeline operates in a fully automated manner across the heterogeneous data formats produced by the Very Long Baseline Array (VLBA) over three decades of operations. The Search for Milli-Lenses (SMILE) project, requiring the calibration of 5000 VLBA sources, makes such blind automation a practical necessity. We introduce the Automated VLBI pipeline in CASA (AVICA), which automates the calibration of archival VLBA data. AVICA extends the CASA-based rPICARD framework by automating preprocessing of FITS-IDI and Measurement Set data formats, calibrator and reference antenna selection via FFT-based fringe detection, and execution of the full calibration workflow. Progress tracking is handled by ALFRD (Automated Logical Framework for executing Dynamic scripts), which orchestrates pipeline execution and records results in real time. AVICA was validated on 1000 NRAO archival sources spanning 1995-2023, covering 1372 band-separated observations across the S, C, X, U, and K bands. Calibrated output was produced for 978 sources (97.8%), with 22 failures due to corrupted or incomplete data. Mean per-source execution time was 30 minutes using MPI parallelization with up to 20 cores. AVICA demonstrates that fully blind calibration of heterogeneous archival VLBA data is achievable with CASA. The automated calibrator and reference antenna selection will be incorporated into a future rPICARD release, extending blind calibration to any supported array. AVICA and ALFRD are available as open-source Python packages.

Related papers