Programmable k-local Ising interactions and shallow optical Kolmogorov--Arnold networks through repeated data encounters

arXiv:2508.17440 · physics.optics, cs.ET, cs.LG · Submitted 2026-08-22 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Programmable k-local Ising interactions and shallow optical Kolmogorov--Arnold networks through repeated data encounters".

Jane: The paper was written by Nikita Stroev and Natalia G. Berloff from Department of Physics of Complex Systems, Weizmann Institute of Science and Department of Applied Mathematics and Theoretical Physics, University of Cambridge.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Jane: The initial discussion is really about how they are tackling the core problems that have historically required two different devices. You know, Ising machines are designed for quadratic problems, but real-world problems often require interactions of order k greater than two.

Tom: And KANs need hundreds of independent, per-ridge nonlinearities to function properly in a machine learning context. So how do we get both high-order discrete coupling and continuous function approximation in the same space?

Lu: The title suggests they are solving this by leveraging a structural nonlinearity that is normally linear, which is quite creative. They aren't forcing the hardware to behave nonlinearly through physical materials like Kerr media.

Meng: That’s good news for fabrication, because relying on material nonlinearities often means dealing with alignment issues and power constraints. I hope this method scales well beyond small demonstration sizes for "Programmable k-local Ising interactions and shallow optical Kolmogorov-Arnold networks on Photonic Platforms.

Jane: Meng makes a practical point; we need to ensure that the complexity doesn' is manageable, but it sounds like the way they addressing both discrete search and continuous function approximation is quite elegant.

Tom: It really feels like they are creating a unified framework, Jane. They aren't just building two systems side by side; they are finding a common language between these distinct optimization and learning paradigms.

Lalam: The implication for me is that this could mean we no longer have the distinction between a pure "solver" machine and an "inference" model, because they can share the same physical hardware.

Paper discussion summary: Tom: So, what is the central mechanism that makes this possible? It's not just one component; it’s how they arrange the light path.

Jane: The key idea centers around a folded 4f relay system, which is essentially a two-pass mechanism. It takes the initial spin field and re-images it back onto the same spatial light modulator, effectively allowing them to isolate clique sums.

Lu: This two-pass relay is where the magic happens because it transforms that linear scatterer into a computational resource by giving every single window its own dedicated second pass phase patch.

Meng: And that’s crucial for us engineers: each selected clique or channel gets its own independent pass, which means we aren't fighting signal crosstalk between different parts of the input data.

Jane: Exactly, Meng; and because it is a programmable polynomial of the clique sum, this allows us to achieve both native k-local couplings and the many independent nonlinearities required for KAN layers in "Programmable k-local Ising interactions and shallow optical Kolmogorov-Arnold Networks on Photonic Platforms."

Tom: This isn't just about being able to do it, but doing it without needing messy nonlinear materials. The structure itself is doing the work.

Lalam: I love that the entire process relies on structural manipulation of light rather than forcing chemical reactions; it feels like a very clean way to achieve complex computations.

Improvements suggested by the paper: Tom: Let's talk about what this means for improving upon previous work, because previous systems were either global in their nonlinearity or they were limited to quadratic interactions.

Jane: The paper is showing that we can achieve k-local interactions without resorting to quadratization, which is a huge algorithmic improvement. We are getting more "spins" and less complexity in the problem formulation.

Lu: And by eliminating the need for external nonlinear materials, we are making a hardware change of only one extra lens and a fold, which is remarkably minimal overhead for such massive functionality.

Meng: From an engineering standpoint, that minimal footprint is fantastic; it suggests these "Programmable k-local Ising interactions and shallow optical Kolmogorov-Arnold Networks on Photonic Platforms" can be implemented on existing SPIM or even injection-locked VCSEL arrays.

Jane: That applicability across different platforms is a big advantage; we aren' not limited to one type of hardware, which greatly increases the potential deployment options.

Tom: And it’ also allows for in-situ physical gradients using the two frames—forward and adjoint—which is a very efficient way to train these KANs without the slow process of electronic backpropagation.

Lalam: The ability this suggests we can train things directly on the photonic substrate feels like a massive step toward creating self-contained, self-optimizing AI systems for society.

Conclusion: Tom: We've covered so much ground today, from the initial title to how we actually build these machines using "Programmable k-local Ising interactions and shallow optical Kolmogorov-Arnold Networks on Photonic Platforms."

Jane: It’s clear that this is not just a marginal improvement; it's a fundamental shift in how we approach both optimization and machine learning.

Lu: The fact that the mathematical structure for the optimal polynomial response is locally lower-triangular means that finding the solution isn't intractable, which provides strong confidence in the scalable nature of these designs.

Meng: My biggest takeaway is that this architecture offers a clear, repeatable blueprint for engineering, meaning we can transition from theory to actual fabrication much faster now.

Lalam: It’s a powerful vision where discrete problem-solving and continuous function learning are not just coexisting but are inherently unified on the same photonic platform.

Tom: I think we've given our listeners a real idea of what this means for "Programmable k-local Ising interactions and shallow optical Kolmogorov-Arnold Networks on Photonic Platforms." It’s truly impressive work by the authors, Nikita Stroev and Natalia G. Berloff.

Jane: It sounds like the future is here, Tom. We hope to see these concepts implemented in real applications very soon.

Lu: I'm definitely going to be watching how this enables hybrid optimization-learning systems in my own work on AI architecture.

Meng: Hopefully, the next engineering phase will involve scaling this up to millions of variables using metasurface SLMs as the authors suggest.

Lalam: It’s an exciting moment for me because I believe that it can lead to a more efficient and elegant way for humanity to solve complex problems.

Nikita Stroev, Natalia G. Berloff

Department of Physics of Complex Systems, Weizmann Institute of Science · Department of Applied Mathematics and Theoretical Physics, University of Cambridge

physics.optics, cs.ET, cs.LG

Submitted: 2026-08-22

Updated: 2026-08-25

Importance score: 84/100

The gist: Photonic computing promises "energy-efficient acceleration for optimization and learning," yet historically, "discrete combinatorial search and continuous function approximation have largely required

Key concepts

Programmable k-local Ising interactions
This refers to the ability to model complex discrete problems that require interactions of order k greater than two. The system achieves this by using a structural method that allows for higher-order coupling without needing messy nonlinear materials.
Kolmogorov-Arnold Networks (KANs)
KANs are machine learning models requiring many independent, per-ridge nonlinearities for proper function. The platform enables these continuous function approximations by utilizing the structural manipulation of light paths.
Folded 4f relay system
This is the central mechanism described, functioning as a two-pass light path. It re-images the initial spin field back onto the spatial light modulator, which allows for isolating clique sums and transforming a linear scatterer into a computational resource.

Terminology

Summary

Photonic computing promises energy-efficient acceleration for optimization and learning, yet historically, discrete combinatorial search and continuous function approximation have largely required distinct devices and control stacks. This paper addresses this limitation by unifying k-local Ising optimization and optical Kolmogorov-Arnold network (KAN) learning on a single photonic platform.

The core of the approach is the introduction of an SLM-centric primitive that realizes, in one stroke, all-optical k-local Ising interactions and fully optical KAN layers. This is achieved through a two-bounce relay mechanism. The key idea is to transform the structural nonlinearity of a nominally linear scatterer into a per-window computational resource: a folded 4f relay re-images the first Fourier plane onto the SLM so that each selected clique or channel occupies a disjoint window with its own second pass phase patch.

This mechanism provides two simultaneous capabilities. First, it yields native, per clique k-local couplings without nonlinear media. The target energy contribution for a k-local Ising Hamiltonian is sum i in q s i. Within each clique q, this product depends only on the clique sum S q = sum i in q s i. By using "one-dimensional interpolation on the discrete set S k with the parity constraint implied by k, there exists a unique univariate polynomial of degree at most k and with the same parity as k such that sum i in q s i = q(S q). The optical measurement for this is defined as the window-integrated intensity, which, due to the two-pass transfer, admits an even- or odd-parity polynomial expansion in S q, I q(S; theta q) = sum r=0 r k (mod 2) r,q(theta q) S q r." The measured energy is then E meas(s; theta q) = sum clique w q I q(S; theta q) + const.

Simultaneously, the same primitive supplies the many independent, univariate nonlinearities required by KAN layers. The KAN architecture uses linear projections z j,m = w j,m T x. The two-pass relay implements this in a per-window manner: The small-phase linearization of the perwindow map is r,jm(theta j,m) = A j,m[r, n] theta n (j, m) + O theta j,m squared. This allows for a fully optical KAN layer where the ridge functions are implemented as a calibrated map from z to I j,m(theta j,m; z).

The training process is also unified. The authors adopt the two-frame adjoint protocol for in-situ physical gradients. This method allows for measuring all parameter gradients from two optical frames per update (forward + adjoint) without electronic backpropagation. The loss L is quadratic, meaning gradient descent reaches the global minimizer.

The hardware implementation of this unified primitive is minimal and platform-agnostic. "We outline implementations on spatial photonic Ising machines, injection-locked vertical cavity surface emitting laser (VCSEL) arrays, and Microsoft analog optical computers; in all cases the hardware change is one extra lens and a fold (or an on-chip 4f loop), enabling a minimal overhead, massively parallel route to high-order Ising optimization and trainable, all-optical KAN processing on one platform."

The results of the the study show that diagonal calibration reduces the initial error by about one decade, and subsequent gradient refinement drives this error to 10-3 - 10-4 once n iter 100. The paper concludes that this architecture reduces overhead, improves reuse of photonic components, and opens a path to field-programmable photonic co-processors that span discrete and continuous workloads.

Improvements for AI systems

The following improvements detail the technical innovations derived from the paper and articulate precisely what a system utilizing these advancements can achieve.

  1. Native K-Local Interaction Implementation:
  • Improvement: The architecture eliminates the need for quadratization (the process of reducing high-order k-local Hamiltonians to quadratic forms). Instead, it natively represents k-spin products (k>2) as a function of the clique sum (S q).

  • Mechanism: A dedicated, folded 4f relay is used. The first pass performs Fourier transformation, and the second pass re-images that specific clique window onto the SLM with a dedicated phase patch (phi q(y; theta q). This creates a localized, per-window structural nonlinearity.

  • ** Improvement:** This avoids the exponential increase in variables and complexity associated with quadratization, maintaining native high-order interactions.

  1. Decoupled Per-Window Computational Resource:
  • Improvement: The computational resource (the structural nonlinearity) is localized and modularized into discrete windows (cliques or ridges).

  • Mechanism: Each window q operates independently, applying a dedicated, parity-matched polynomial response (q(S q)). This ensures that the k-local coupling weights are highly programmable without interference from other cliques.

Improvement: Massive parallelism and minimal cross-talk between localized computational units.

  1. In-Situ Physical Gradient Training (Adjoint Protocol):
  • Improvement: A two-frame (forward and adjoint) optical protocol is used to calculate the physical gradients (d L / d theta) directly from the measured intensity (I q).

  • Mechanism: By measuring the change in intensity following a forward pass and an adjoint pass, electronic backpropagation is entirely bypassed. This provides direct physical feedback for training the phase-depth parameters (theta q or theta j,m).

Improvement: Real-time training capability (in-situ learning) without the latency and power overhead of traditional electronic gradient computation.

  1. ** Unification of Paradigms:**
  • Improvement: The same physical primitive (the 4f relay + dedicated phase patches) simultaneously provides:

a) A K-Local Solver: A set of independently weighted, high-order polynomial responses for discrete optimization.

b) A KAN Engine: A massive array of independent, trainable univariate nonlinear functions (j,m) acting on linear projections.

Improvement: The ability to combine continuous function approximation (KAN) and discrete combinatorial search (Ising/k-SAT) on a single, unified hardware platform.


The improved system is capable of solving and modeling complex problems that require both high-order discrete search and continuous function approximation, with significantly enhanced efficiency:

  1. High-Order Combinatorial Optimization:
  • The system can solve large instances of k-SAT, graph partitioning, and multi-spin interaction problems (e.g., in condensed matter physics) that are intractable for traditional quadratic solvers (k>2). This achieves 3 orders of magnitude reduction in required variables compared to quadratization methods.
  1. Efficient Continuous Function Approximation:
  • The system can implement and train complex functions (e.g., solutions to Poisson PDEs or sophisticated financial risk models) using the KAN architecture, achieving orders-of-magnitude parameter reduction (up to 100x fewer parameters than MLPs) while maintaining superior approximation accuracy.
  1. Hybrid Algorithmic Acceleration:
  • The system can perform interleaved optimization and learning tasks. For example, a KAN layer could be trained to predict the optimal coupling coefficients (J q) for a specific k-local problem, and then running the k-local solver on that prediction—a process previously impossible to unify in hardware.
  1. System Scalability and Reconfigurability:
  • The modular, per-window design allows for massive scaling (millions of spins/ridges) while maintaining high fidelity due to the independence of each computational unit. The entire system can be reconfigured by simply reprogramming the phase settings (theta q) on the SLM or VCSEL array.

Sources

Related papers