Optimal Symmetries in Binary Classification
cs.LG, cs.AI, physics.data-an, stat.ML
Submitted: 2024-08-16
Updated: 2026-09-21
Comments: added experiments on Jet tagging for optimal and non-optimal symmetries
License: http://creativecommons.org/licenses/by/4.0/
The gist: We develop a theoretical foundation for designing group-equivariant neural networks that align the choice of symmetries with the underlying probability distributions of the data.
Terminology
Abstract
We develop a theoretical foundation for designing group-equivariant neural networks that align the choice of symmetries with the underlying probability distributions of the data. Utilising the general structure of fibre decompositions on the domain under group equivariant maps and its relation to that of the likelihood ratio, we present a theoretical framework for identifying group actions that maintain optimal classification performance via the Neyman-Pearson lemma. This provides a unified methodology for improving classification accuracy especially in fundamental applications where one has knowledge of the inherent symmetries of the distributions and how they are broken by measurement. As an application to jet classification at the Large Hadron Collider, we find that there can be performance gains when one utilises smaller permutation symmetries within the constituents. This work offers insights and practical guidelines for constructing more effective group equivariant architectures in diverse machine-learning contexts.
Sources
- An Efficient Lorentz Equivariant Graph Neural Network for Jet Tagging
- Explainable Equivariant Neural Networks for Particle Physics: PELICAN
- Lorentz-Equivariant Geometric Algebra Transformers for High-Energy Physics
- Adam: A Method for Stochastic Optimization
- Approximating Likelihood Ratios with Calibrated Discriminative Classifiers
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks