peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies

arXiv:2508.04803 · cond-mat.mtrl-sci, cond-mat.mes-hall, cond-mat.str-el, cond-mat.supr-con · Submitted 2025-08-06 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: Quantum Radio. Generated commentary on the latest quantum physics and condensed matter papers.

Kai: Today's paper: "peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies".

Mira: Angle-resolved photoemission spectroscopy (ARPES) provides a direct experimental probe of electronic band structures, and this paper introduces 'peaks',

Kai: First, who's behind it and why it matters.

Paper summary: Mira: The paper argues that the current landscape has several issues: existing packages like PyARPES, for example, impose too many fundamental convention choices regarding angular and energy scales which makes them hard to use when you have multiple experimental setups (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Kai: So the paper is essentially saying that because ARPES data is becoming multi-dimensional, we need a tool that can handle this complexity efficiently, not just one that handles one specific setup well (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Lev: From an error correction standpoint, I see the problem as reproducibility; if every lab uses different conventions without clear translation tools, replicating a result across different experimental platforms becomes nearly impossible (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Kai: And 'peaks' claims to solve this by supporting lazy data loading and parallel processing, which directly addresses the increasing data volumes that are making traditional analysis methods too slow (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Mira: Furthermore, the authors emphasize the need for transparency by incorporating extensive metadata handling using models like pydantic to ensure a consistent framework across different data types, which they claim is essential for reproducibility in this field (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Lev: If they can nail the metadata consistency, it means that when we eventually want to run simulations or tests on actual quantum hardware, we'll have a much clearer picture of what inputs are necessary and what assumptions were made in the data preparation stage (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Kai: So, the core message is that ARPES is evolving into a multidimensional spectroscopy, and 'peaks' is presented as the modern Python solution built on xarray to manage this evolution (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Mira: It’s about providing flexibility; they mention supporting data loading into xarray:DataTree structures, which allows users to group data in ways that mirror the actual hierarchy of their experiment (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Lev: That hierarchy grouping is important because it lets you isolate specific slices of the data—like a particular momentum or energy range—which is exactly what we need to do when checking if a physical process happens in a certain region (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Kai: So, this paper sets up 'peaks' as an infrastructure piece for the next generation of ARPES analysis, moving beyond simple visualization to complex data management (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Mira: And they are also looking ahead by including initial functionality for unsupervised machine learning approaches like clustering and denoising right within the package (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Lev: That’s exciting because if we can use ML to clean up noisy ARPES signals, it could potentially help us extract clearer features that are more relevant when mapping those features onto our theoretical models for quantum error correction (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Kai: That’s the big picture, right? It’s about making the data handling robust enough to support those advanced computational methods we're trying to develop.

Conclusion: Mira: Looking at the work, the implication is that by standardizing how we handle the massive, multi-dimensional ARPES datasets, this package helps bridge the gap between experimental measurement and theoretical simulation more effectively (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Kai: I think what this means in practical terms is that researchers can spend less time debugging data loading issues and more time focusing on the actual physics—whether that's understanding band structures or figuring out how to build better quantum error correction protocols (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Lev: For us, the impact is that we can start using these standardized tools to prepare data in a way that is directly usable by our simulation pipelines, which speeds up the iterative process of testing new physical ideas without getting bogged down in pre-processing headaches (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Mira: And because the package includes initial ML tools, it suggests that this standardization isn't just about data handling; it’s about creating a more structured environment where advanced analytical techniques can be applied consistently across different experimental outcomes (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Kai: So, 'peaks' is presented as an essential piece of software infrastructure that supports the next phase of ARPES research by making the data management side less cumbersome (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

Lev: It’s about providing a solid foundation so that when we move from analyzing experimental snapshots to building predictive models for quantum states, we have reliable data pathways (peaks: a Python package for analysis of angle-resolved photoemission and related spectroscopies).

SUPA, School of Physics and Astronomy, University of St Andrews

cond-mat.mtrl-sci, cond-mat.mes-hall, cond-mat.str-el, cond-mat.supr-con

Submitted: 2025-08-06

Updated: 2025-08-06

Journal ref: Journal of Open Source Software, 11(126), 10440 (2026)

DOI: 10.21105/joss.10440

Code: https://github.com/phrgab/peaks

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 81/100

The gist: Angle-resolved photoemission spectroscopy (ARPES) provides a direct experimental probe of electronic band structures, and this paper introduces 'peaks', a Python package designed to facilitate fast

Key concepts

xarray
xarray is a powerful Python library used to create labeled, N-dimensional arrays. It is essential for 'peaks' because ARPES data is inherently multi-dimensional (e.g., momentum and energy). xarray allows the package to handle these complex data structures efficiently, supporting lazy loading and parallel processing of large datasets.
DataHeterogeneity Handling
ARPES experiments often use different conventions, scales, and units across various machines. 'peaks' solves this by using location-specific data loaders that automatically adapt to different starting formats. It ensures that data from multiple experimental setups can be loaded and processed consistently within the package.
Metadata Framework
The package uses pydantic models to enforce a consistent framework for metadata attached to the data. This means every piece of ARPES data comes with standardized information about its origin, conventions, and properties. This consistency is crucial for transparent analysis and reliable interpretation of the results.
Machine Learning Integration
The package includes initial tools for unsupervised machine learning analysis, such as principal component analysis and clustering. These ML methods help researchers perform tasks like denoising or grouping spatially-resolved ARPES data, moving the analysis beyond simple visualization to more advanced pattern recognition.

Terminology

Summary

Angle-resolved photoemission spectroscopy (ARPES) provides a direct experimental probe of electronic band structures, and this paper introduces 'peaks', a Python package designed to facilitate fast visualization and complex analysis of multi-dimensional ARPES datasets.

The gist

'peaks' provides a Python package for advanced data analysis of ARPES and related spectroscopic data.

Statement of need

Over recent years, significant technological improvements have developed ARPES into a truly multidimensional spectroscopy, requiring efficient handling and advanced analysis of 3-, 4-, and higher-dimensional datasets. Furthermore, there is an increasing push to incorporate machine learning (ML) methods into the analysis pipeline, alongside the need for greater transparency through open-source packages with clear metadata handling. This motivates the use of Python for ARPES data analysis due to these requirements.

Package comparison and motivation

The authors note several existing packages and highlight why 'peaks' is necessary:

  1. PyARPES (Stansbury & Lanzara, 2020) makes several fundamental convention choices (regarding angular and energy scales and units, alignments, and sign conventions) which complicates use with multiple experimental setups.

  2. pesto (Polley, 2025) is noted as an excellent easy-to-use alternative, but is heavily oriented towards use with data collected from the Bloch beamline of the Max-IV synchrotron.

  3. ERLabPy (Han, 2025) provides similar functionality to 'peaks', though with differences in the approach to handling data (e.g., co-ordinate systems).

How it works

'peaks' is designed to be run in an interactive notebook environment and supports lazy data loading and parallel processing, reflecting the increasing data volumes. Its core structure is built heavily on the xarray package, which provides a powerful data structure for the N-D labelled data arrays common to ARPES data.

Data handling and metadata

The package manages heterogeneity in existing ARPES setups through several features:

Data is loaded into xarray:DataArray’s using location-specific data loaders to support multiple starting data formats and conventions, reflecting the heterogeneity in existing ARPES setups.

Extensive metadata is included in the DataArray attributes, making use of pydantic (Colvin et al., 2025) models to ensure a consistent metadata framework.

It also supports loading data into xarray:DataTree structures, allowing the user flexibility over grouping data in configurations which reflect the data hierarchy of the underlying experiment, and permitting batch processing or metadata configuration.

Analysis and visualization tools

'peaks' aims to provide a relatively comprehensive suite of tools for ARPES and related spectroscopic data via a modular approach, supporting the experimenter from initial data acquisition, visualisation, and sample alignment through data processing and more advanced analysis. Key capabilities include:

  1. Tools for aiding the experimenter in aligning samples for subsequent measurements, with care taken to handle different conventions to facilitate both standardized analysis and 'on-the-fly' analysis.

  2. ARPES-specific data selection, such as momentum (MDC) and energy (EDC) distribution curve extraction.

  3. Data processing tasks including momentum conversion and Fermi level corrections, data normalisation, and derivative-type methods to aid data visualisation.

  4. Data fitting capabilities, which include parallel processing and fitting of lazily-loaded data - see e.g. Figure 2, building on the lmfit package.

  5. Initial functionality for unsupervised machine learning analysis of spatially-resolved ARPES data, including principal component analysis, clustering, and denoising built around Scikit-learn.

Future plans

The authors indicate that in the future, 'peaks' could be augmented with additional ML approaches tailored to ARPES data analysis, facilitated by the standard xarray-based data structures. The incorporation of additional data structures and functionality for processing spin-resolved ARPES data is also planned. The package is available at https://github.com/phrgab/peaks.

Acknowledgements

The development was supported by discussions, suggestions, and bug reports from various researchers, as well as financial support from the UK Engineering and Physical Sciences Research Council (Grant Nos. EP/X015556/1, EP/T02108X/1, and EP/R025169/1), the Leverhulme Trust (Grant Nos. RL-2016-006 and RPG-23-256), and the European Research Council (through the QUESTDO project, 714193). The code is available at https://github.com/phrgab/peaks.

Improvements for AI systems

Here are specific improvements to AI systems based on the capabilities described in the peaks package:

  1. The improved system can perform automated, high-dimensional data analysis of Angle-Resolved Photoemission Spectroscopy (ARPES) and related spectroscopic data, including 3-, 4-, and higher-dimensional datasets.

  2. It can facilitate fast visualization of multi-dimensional ARPES datasets using extensive inline and pop-out GUI support.

  3. The system can handle the complex data hierarchy typical of ARPES experiments, supporting lazy data loading via Dask arrays for efficient processing of datasets exceeding available memory, and parallel processing capabilities.

  4. It can standardize data handling across heterogeneous experimental setups by utilizing location-specific data loaders to manage different input formats and conventions (angle scales, energy units, sign conventions).

  5. The system can ensure consistent metadata frameworks using Pydantic models during data loading to guarantee reproducibility and transparency in ARPES analysis.

  6. It can maintain a detailed analysis history (data provenance) by recording all applied processing steps, facilitating effective collaborative working on experimental results.

  7. The system can perform automated data selection tasks, including extraction of momentum distribution curves (MDC) and energy distribution curves (EDC), data merging, summation, symmetrization, and necessary preprocessing steps like momentum conversions and Fermi level corrections.

  8. It can execute advanced fitting procedures for large datasets using lazy loading techniques (e.g., fitting individual EDCs or MDCs from across a large dataset in parallel).

  9. It can incorporate unsupervised machine learning methods (built around Scikit-learn) for initial analysis, such as principal component analysis (PCA), clustering, and denoising of spatially-resolved ARPES data.

  10. The improved system can be augmented with future machine learning approaches tailored to ARPES data, leveraging its standard xarray-based data structures for further ML integration.

Abstract

The electronic band structure, describing the motion and interactions of electrons in materials, dictates the electrical, optical, and thermodynamic properties of solids. Angle-resolved photoemission spectroscopy (ARPES) provides a direct experimental probe of such electronic band structures, and so is widely employed in the study of functional, quantum, and 2D materials. peaks (Python Electron spectroscopy Analysis by King group @ St Andrews) provides a Python package for advanced data analysis of ARPES and related spectroscopic data. It facilitates the fast visualisation and analysis of multi-dimensional datasets, allows for the complex data hierarchy typical to ARPES experiments, and supports lazy data loading and parallel processing, reflecting the ever-increasing data volumes used in ARPES. It is designed to be run in an interactive notebook environment, with extensive inline and pop-out GUI support for data visualisation.

Related papers