Bidirectional Neural Networks for Global Nucleon-Nucleus Optical Model Calculations
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "Bidirectional Neural Networks for Global Nucleon-Nucleus Optical Model Calculations".
Tom: A neural network emulator based on Bidirectional Liquid Neural Networks (BiLNN) provides a differentiable mapping from an optical potential to scattering wave functions,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: Speaking of structure, let’s talk about what these authors are calling "Bidirectional Neural Networks for Global Nucleon-Nucleus Optical Model Calculations" and why that title matters. It tells us immediately that they aren't just looking at a single interaction; they are aiming for a global view across different nuclei and energies.
Jane: That’s right, Tom; the focus on "Global" suggests the network is intended to generalize beyond just one specific scenario, which is huge when we talk about nuclear data evaluation where you need consistency everywhere. The authors are addressing the fact that traditional methods struggle with needing accurate predictions for systems far from stability or at very high energies.
Lu: The paper shows how they solve this by using a specific coordinate system, rho = kr, which they claim normalizes the oscillation wavelength regardless of the projectile energy, allowing one network to cover a massive range from one to two hundred MeV <ref:2512.22500#pg0,the oscillation wavelength regardless of>. This normalization is what gives them that broad applicability.
Meng: Normalizing wavelengths across such a huge energy span sounds mathematically complex; how do they ensure that this universal coordinate system still accurately captures the physics of different nuclei, like comparing twelve C to two hundred eight Pb <ref:2512.22500#pg0>?
Lalam: The methodology seems to suggest that by using these phase-space coordinates, the network learns a structure inherent in the scattering solutions themselves rather than just memorizing data for each specific nucleus or energy point.
The paper's summary: Tom: So, summarizing what this paper actually does, they’ve introduced an emulator based on BiLNN that provides a differentiable link between the optical potential and the resulting scattering wave functions. The key takeaway here is that this allows researchers to do gradient-based optimization for finding better potentials or understanding how small changes in input parameters affect the output predictions.
Jane: That differentiability is critical because it lets us use standard optimization tools, like AdamW with a cosine annealing schedule, to tune these models directly. This means we can search for optimal optical potentials much faster than running iterative numerical methods every time we want to test a new interaction.
Lu: The training involved using Numerov solutions computed with the KD02 optical potential across twelve target nuclei, spanning both protons and neutrons up to partial waves l = thirty <ref:2512.22500#pg0>. This extensive training set is what gives the network enough data to learn this complex relationship.
Meng: That’s a lot of data generation—solving the Schrödinger equation numerically for all those combinations—so how large was this training dataset in practice, and what was the resulting accuracy when they compared it to known solutions?
Lalam: The paper states that the model achieves an overall relative error of zero point six percent across its entire training domain, which shows a high degree of fidelity to the original Numerov solutions used for training.
The paper's improvements: Tom: Now let's look at what they suggest as improvements or key architectural choices within this framework. They highlight that the Bidirectional architecture is particularly important, proving most critical in their ablation study, which is a significant finding in itself about how the structure of the network helps.
Jane: It’s interesting because they also found that certain input features were more impactful than others; specifically, they noted that the Sommerfeld parameter eta and mass encoding A one/three A one/six proved to be more influential than some of the hand-crafted semiclassical features.
Lu: The paper suggests that the phase-space coordinate rho = kr is the most powerful tool, because it demonstrates that scattering wave functions have a universal structure when viewed in those natural units of de Broglie wavelength, which is what allows for generalization across energy variations.
Meng: So, while the bidirectional nature and these phase-space coordinates are key to generalization, from a practical standpoint, is the main limitation something else? What does the paper state as something this method doesn't do well?
Lalam: The authors flag that while they’ve achieved a high relative error of zero point six percent in training, the system still requires careful application because it’s an emulator; it doesn't replace the underlying physical solution entirely but acts as a differentiable surrogate for solving the Schrödinger equation.
Conclusion: Tom: Alright team, we’ve covered a lot about this work on "Bidirectional Neural Networks for Global Nucleon-Nucleus Optical Model Calculations." We’ve seen how they use BiLNNs and phase-space coordinates to create a differentiable solver that can handle a huge energy range and multiple nuclei.
Jane: It really boils down to having an AI system that can rapidly evaluate complex physics without needing slow, iterative numerical integration every single time, which opens up avenues for much faster model development in this area.
Lu: The implications for the field are significant because it validates that we can encode the smooth dependence of solutions on target mass and charge into the network rather than just memorizing specific targets.
Meng: From an engineering perspective, if we can use these gradients to tune potentials efficiently, it means we can speed up the entire pipeline for creating new nuclear models significantly.
Lalam: I think this advance is particularly important because it shows how deep structural properties of physical solutions can be learned through neural networks, potentially improving how we design and train more complex physical AI systems in general.
Tom: Exactly! We’re wrapping up our discussion on the "Bidirectional Neural Networks for Global Nucleon-Nucleus Optical Model Calculations," but I think we’ll save some time to look at those other exciting papers next.
School of Physics Science and Engineering, Tongji University · Southern Center for Nuclear-Science Theory (SCNT), Institute of Modern Physics, Chinese Academy of Sciences
nucl-th, cs.LG
Submitted: 2025-12-27
Updated: 2026-06-16
DOI: 10.1103/qw54-df4l
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 76/100
The gist: A neural network emulator based on Bidirectional Liquid Neural Networks (BiLNN) provides a differentiable mapping from an optical potential to scattering wave functions, enabling gradient-based
Key concepts
- Phase-space coordinates ρ = kr
- This transformation normalizes the oscillation wavelength based on the de Broglie wavelength (kr). This allows one network to accurately describe scattering waves across a very wide range of projectile energies by treating the wave function oscillations as having a universal period in this new coordinate system.
- Bidirectional Liquid Neural Network (BiLNN)
- The BiLNN architecture uses closed-form continuous-time layers that process the radial sequence both forward and backward. This structure naturally incorporates both boundary conditions of scattering problems—the wave function starting at zero at the origin and behaving correctly far away from the nucleus.
- Optical Potential V(r)
- The optical potential represents the combined effect of a projectile interacting with the target nucleus, including both real (attractive/repulsive) and imaginary (absorption) components. The network learns to map this potential, along with physical parameters like energy and mass, directly to the resulting scattering wave function.
Terminology
Summary
A neural network emulator based on Bidirectional Liquid Neural Networks (BiLNN) provides a differentiable mapping from an optical potential to scattering wave functions, enabling gradient-based optimization and uncertainty quantification for nucleon-nucleus scattering. This work is significant because it introduces a single network capable of generalizing across the full parameter space—energies, partial waves, and target nuclei—while maintaining accuracy comparable to traditional Numerov solutions.
How it works
The core innovation is the use of phase-space coordinates, specifically phase-space coordinates ρ = kr,
which normalize the oscillation wavelength regardless of projectile energy.
This transformation allows a single network to span a vast energy range from 1 to 200 MeV by making the wave function oscillate with a universal period in this new coordinate system. The network is trained on wave functions computed using the Numerov method with the KD02 optical potential, spanning twelve target nuclei (12C to 208Pb), both protons and neutrons, and partial waves up to l = 30.
Bidirectional Liquid Neural Network (BiLNN) Architecture
The network architecture is a Bidirectional Liquid Neural Network (BiLNN)
based on Closed-form Continuous-time (CfC) layers. This bidirectional structure processes the radial coordinate sequence in both forward and backward directions, naturally incorporating the two boundary conditions of the scattering problem: the wave function vanishing at the origin and the asymptotic Coulomb behavior at large distances.
The architecture consists of five components: an encoder, a forward liquid layer integrating from r = 0 outward, a backward liquid layer integrating from rmax inward, a combiner that merges information from both directions, and a decoder to produce the real (ψR) and imaginary (ψI) parts of the radial wave function.
Input Features and Learning Mechanism
The network is trained to learn a mapping from the optical potential V(r) and physical parameters (energy E, partial wave l, target mass A) to the radial wave function ul(r). At each radial point, the network receives a 9-dimensional feature vector designed to capture specific physical information. Key features include:
**/ρ/ρmax: The normalized phase-space coordinate. 1) **
VR/E and W/E: The real and imaginary parts of the optical potential, normalized by projectile energy, which determines whether the particle is in a classically allowed or forbidden region. 2)
η/ηmax: The Sommerfeld parameter, characterizing Coulomb interaction strength relative to kinetic energy. 3)
sin ϕWKB and cos ϕWKB: WKB phase information derived from the local wave number, providing guidance on expected oscillation phase. 4)
D/Dmax: The cumulative WKB absorption integral, capturing the amplitude envelope due to the imaginary potential. 5)
l/lmax and A1/3/6: Normalized partial wave quantum number and a scale factor related to the nuclear radius.
Training, Optimization, and Generalization
The network is trained by minimizing a cost function defined as the mean squared error between predicted and true wave functions: L = 1/N Σ h(ψ pred R,i − ψ true R,i) squared + (ψ pred I,i − ψ true I,i) squared.
Optimization is performed using the AdamW algorithm with a cosine annealing schedule. The training data consists of approximately 148,800 examples generated by solving the Schrödinger equation numerically and interpolating the Numerov solutions onto a 100-point uniform grid. The model achieves an overall relative error of 0.6%
across the training domain and successfully generalizes to nuclei not included in training (24Mg, 63Cu, 184W) with comparable accuracy.
Physical Insights and Hierarchy of Design Choices
The study reveals several physical insights:
-
The effectiveness of the phase-space coordinate ρ = kr demonstrates that
scattering wave functions possess a universal structure when viewed in the natural units of the de Broglie wavelength,
exploiting scale invariance to generalize across energy variations. -
The success in generalizing to unseen nuclei validates that the network has learned
the smooth dependence of the scattering solution on target mass and charge as encoded in the KD02 parameterization rather than memorizing specific targets.
-
The ablation study shows a hierarchy:
The bidirectional architecture proves most critical,
while WKB input features provide only amodest improvement.
Furthermore, the Sommerfeld parameter η and mass encoding A1/3/6 are found to be more impactful than hand-crafted semiclassical features.
Conclusion
The BiLNN serves as a differentiable surrogate Schr¨odinger equation solver,
fundamentally different from traditional numerical integrators because it allows for "end-to-end gradient propagation through the surrogate.
Improvements for AI systems
As a fastidious researcher, I have analyzed the provided paper, Bidirectional Neural Networks for Global Nucleon-Nucleus Optical Model,
and identified several high-impact areas where this methodology can be leveraged to significantly improve existing AI systems.
Here are the specific improvements and capabilities of an improved AI system:
)
-
Improve gradient-based optimization in classical/physical simulations (e.g., fluid dynamics, molecular modeling, quantum chemistry).
-
Enhance uncertainty quantification in high-dimensional parameter spaces (e.g., Bayesian inference for complex physical models).
-
Enable rapid, differentiable exploration of large model parameter sets without requiring expensive numerical differentiation at every step.
)
Improving AI Systems via BiLNN Integration: Specific Improvements and Capabilities
The core innovation presented—a Bidirectional Liquid Neural Network (BiLNN) that learns a mapping from an optical potential to wave functions while maintaining differentiability—offers transformative capabilities for scientific AI systems, particularly those dealing with complex, high-dimensional physical parameter spaces.
Here are the specific improvements and what the improved AI system can achieve:
-
Enhanced Gradient-Based Optimization for Inverse Problems:
-
Differentiable Surrogate Modeling for High-Dimensional Parameter Sweeps:
-
Robust Uncertainty Quantification via Gradient-Informed Sampling:
)
Improving AI Systems via BiLNN Integration: Specific Improvements and Capabilities (Detailed)
- Enhanced Gradient-Based Optimization for Inverse Problems:
The BiLNN acts as a differentiable surrogate Schrödinger equation solver, replacing opaque numerical integrators (like Numerov). This allows the AI system to compute exact gradients of physical observables (cross sections, S-matrix elements) with respect to the optical model parameters.
-
Specific Capability: End-to-end gradient propagation through the surrogate. The AI can use this for efficient parameter optimization of complex potentials where traditional adjoint methods fail or are computationally prohibitive.
-
Benefit: Enables rapid and precise tuning of physical models (e.g., optimizing Woods-Saxon parameters in nuclear physics, or interaction strengths in fluid dynamics) by using standard gradient descent algorithms (like AdamW) directly on the objective function, avoiding the linear scaling cost of numerical differentiation.
- Differentiable Surrogate Modeling for High-Dimensional Parameter Sweeps:
The system learns the underlying functional dependence of a complex physical solution across a vast parameter space (e.g., energy, target mass, partial wave). The phase-space coordinate transformation is key here, as it normalizes oscillation wavelengths universally.
-
Specific Capability: Rapid evaluation of millions of potential configurations. Instead of running a computationally expensive forward model for every point in the parameter space to estimate an observable (like cross sections), the AI system can query the trained BiLNN surrogate instantaneously via its differentiable mapping.
-
Benefit: Enables efficient Bayesian inference and frequentist uncertainty quantification across high-dimensional spaces (e.g., nuclear data evaluation, materials science). This allows for systematic exploration of parameter uncertainties with high throughput, crucial for training robust machine learning models in physics.
- Robust Uncertainty Quantification via Gradient-Informed Sampling:
The differentiability of the network provides exact gradients at a cost comparable to a single forward evaluation.
-
Specific Capability: Gradient-informed sampling and sensitivity analysis. The AI system can use these exact gradients to guide sampling strategies (e.g., using techniques like gradient-based importance sampling) to efficiently explore the posterior distribution of model parameters, rather than relying on slow, brute-force methods (like parametric bootstrap).
-
Benefit: Provides systematic uncertainty propagation in complex physical models with high fidelity. This is vital for determining which parameters have the most significant impact on a physical outcome, leading to more reliable and trustworthy scientific predictions.
)
Summary of AI System Improvements:
The BiLNN transforms the AI system from a black-box predictor into an active, gradient-aware solver surrogate. The improved system can now perform:
-
Optimization of complex physical models using efficient gradient descent on the learned mapping, bypassing costly numerical differentiation.
-
High-throughput Bayesian uncertainty quantification across vast parameter spaces by leveraging instantaneous forward evaluations and exact gradients for efficient sampling.
-
Physics-guided discovery: The AI can learn universal functional relationships (like scale invariance via the phase-space coordinate) encoded in the solution structure, allowing it to generalize effectively to unseen physical configurations (e.g., new nuclei or energy regimes).
Abstract
Modern nuclear data evaluation increasingly requires not only accurate scattering calculations, but also efficient methods for uncertainty quantification and parameter optimization, tasks that benefit from differentiable solvers amenable to gradient-based algorithms. I present a neural network emulator based on Bidirectional Liquid Neural Networks (BiLNN) that provides a fully differentiable mapping from optical potential parameters to scattering wave functions. The key innovation enabling generalization across the parameter space is the use of phase-space coordinates ρ= kr that normalize the oscillation wavelength regardless of projectile energy, allowing a single network to span 1 to 200 MeV. Trained on Numerov solutions for twelve target nuclei (C to Pb), both protons and neutrons, and partial waves up to l=30, the network achieves an overall relative error of 1.2%. The predicted wave functions yield accurate S-matrix elements and elastic scattering cross sections, reproducing diffraction patterns spanning four orders of magnitude. Importantly, the model extrapolates successfully to nuclei not included in training (Mg, Cu, W) with comparable accuracy, demonstrating that it has learned the physics of the optical model rather than memorizing specific targets. The differentiable nature of the trained model opens the door to gradient-based optimization of optical model parameters and efficient uncertainty quantification.
Related papers
- BRST quantization for the restoration of broken symmetries: a pedagogical example
- FUSION: a skill-based research agent for publicly obtainable nuclear-physics codes
- Sensitivity of Neutron Star Observables to Transition Density in Hybrid Equation-of-State Models
- Exterior complex scaling enables physics-informed neural networks for quantum scattering
- Microscopic Insights into the Quarkyonic Hadron--Quark Crossover: Lessons from Ultracold Fermi Gases
- An Effective Upper Bound on the Pressure-to-Energy Density Ratio in Neutron Stars