Streamlined optical training of large-scale modern deep learning architectures with direct feedback alignment

arXiv:2409.12965 · cs.ET, cond-mat.dis-nn, cs.LG, physics.app-ph, physics.optics · Submitted 2025-04-02 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Streamlined optical training of large-scale modern deep learning architectures with direct feedback alignment".

Jane: The paper was written by Ziao Wang, Kilian Müller, Matthew Filipovich, Julien Launay, Ruben Ohana et al. from École Normale Supérieure - Université PSL and Sorbonne Université and Collège de France and CNRS and LightOn and Welinq and University of Oxford and Flatiron Institute and École Polytechnique Fédérale de Lausanne (EPFL).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title and Authors: Tom: Welcome back to the show, everyone! Today we are digging into a paper that has got me genuinely fired up. It’s called “Streamlined optical training of large-scale modern deep learning architectures with direct feedback alignment.” Jane, this title is a mouthful, but it’s basically about training giant neural networks using light instead of just electricity.

Jane: That’s exactly right, Tom. And I love that the authors are from a bunch of different places — ENS in Paris, LightOn, Oxford, EPFL. It’s a real mix of academic labs and a company that builds optical processors. That tells me this isn’t just a theoretical idea; they actually built something.

Tom: Yeah, and that’s the part that gets me. They’re not just simulating this. They’ve got a physical box that does math with lasers and cameras. I mean, the paper says it can do random matrix multiplications at speeds up to one thousand five hundred TeraOPS under thirty Watts. That’s a huge deal for energy efficiency.

Jane: Let’s unpack that for a second. Normally, when you train a neural network, you use backpropagation. That’s the standard way of sending the error signal backwards through the network to update the weights. But backpropagation is sequential — each layer has to wait for the next one. This paper uses something called direct feedback alignment, or DFA, which sends the error directly to every layer at once using a random projection.

Tom: And the clever part is that random projection is exactly what light does naturally when it scatters through a random medium. So they let the optics do that heavy lifting. It’s like the physical hardware is perfectly matched to the algorithm. That’s the kind of co-design that makes me think we’re on the edge of something big.

Jane: Exactly. And they didn’t just test it on toy problems. They trained a Transformer with over a billion parameters on language, a Vision Transformer on climate data, and even a diffusion model. That’s a serious range of modern architectures.

Tom: A billion parameters, Jane. Let’s just sit with that for a second. That’s a record for non-conventional hardware training. I mean, we’re used to GPUs doing this, but they did it with a box of mirrors and a camera.

Jane: And that’s what makes this paper so exciting. It’s not just about doing the same thing faster. It’s about opening a different path for how we scale up AI. Tom, I think we’ve got a lot to unpack here, so let’s get into the actual method next.

Paper Summary: Tom: So, Jane, we’ve got the title and the authors. Now let’s talk about what they actually did in “Streamlined optical training of large-scale modern deep learning architectures with direct feedback alignment.” I want to make sure our listeners understand the core idea.

Jane: Right. So the standard way to train a neural network is backpropagation. It works, but it’s got a bottleneck. The error has to travel backwards layer by layer, and that creates a kind of traffic jam. Direct feedback alignment, or DFA, bypasses that. Instead of passing the error backwards, it takes the error from the output and projects it through a fixed random matrix directly to every layer at the same time.

Tom: And that random matrix multiplication is the perfect job for their optical processor. They call it an OPU. Light goes through a scattering medium, and the speckle pattern that comes out is basically a random projection of the input. It’s fast, it’s parallel, and it uses almost no power.

Jane: Exactly. And the paper shows that this works across different types of models. They trained a GPT-like Transformer on movie dialogues, a Vision Transformer on climate projections, and a diffusion Transformer on images. The performances were comparable to standard backpropagation, especially on the climate task.

Tom: I was really struck by the climate example. They trained a Vision Transformer with one hundred thirty-four million parameters to predict global temperature changes. The predictions looked almost identical to the ground truth. And they also trained a fully connected network with one point three billion parameters on the same task. That’s a massive model.

Jane: And here’s the thing — on that huge fully connected network, the optically trained model actually did better than the backpropagation one. That’s surprising. The authors think it’s because the optical signal is naturally normalized, which helps with very wide layers.

Tom: That’s a really interesting twist. Usually we think of backpropagation as the gold standard. But this paper shows that for certain architectures, especially very wide ones, the optical approach can actually be more stable.

Jane: And they didn’t stop there. They also looked at scaling. They showed that as you make networks deeper and wider, the training time for their optical method grows more slowly than for backpropagation. There’s a crossover point where the optical method becomes faster.

Tom: That’s the part that makes me think this could really matter for the future. We’re constantly hitting walls with GPU memory and energy. If optics can break through that, we might be able to train models that are currently impossible.

Jane: Exactly. And that’s what we’re going to dig into next — what this means for the future of training huge models.

Improvements Suggested: Tom: So Jane, we’ve covered the basics. Now let’s talk about what this paper actually improves. In “Streamlined optical training of large-scale modern deep learning architectures with direct feedback alignment,” they’re not just saying “look, it works.” They’re showing how it can be better than what we have.

Jane: Right. The big improvement is in scaling. Backpropagation has a fundamental problem — the backward pass is sequential. Each layer has to wait for the next one. That creates a bottleneck, especially for deep networks. DFA, on the other hand, sends the error to all layers at once. So the backward pass becomes parallel.

Tom: And the optical processor makes that parallel projection nearly free. The paper shows that the training time for their method scales much better with depth and width. In one experiment, they trained a network with ninety-six layers and three thousand eighty neurons per layer, and the optical method was actually faster than backpropagation on a GPU.

Jane: That’s the crossover point we mentioned. And they pushed it even further. With a memory offloading technique, they trained networks up to two point seven billion parameters, and the optical method maintained its speed advantage.

Tom: But here’s what I find really clever — they also showed that the optical method can go beyond what DFA can do on a GPU alone. The random matrix is stored optically, so it doesn’t take up GPU memory. That means you can train larger models without running out of memory.

Jane: That’s a huge practical advantage. The paper shows that as you scale up, the ratio of training time between optical DFA and regular DFA drops from about fifty down to almost one. And beyond a certain size, regular DFA just runs out of memory, but the optical version keeps going.

Tom: So it’s not just about speed. It’s about breaking through the memory wall that limits how big we can make these models. And that’s exactly what the industry needs right now.

Jane: And there’s another improvement I want to highlight. The paper shows that their optical training is robust to noise. They simulated different kinds of noise — like the random matrix drifting over time — and the training still worked fine. That’s important because real optical hardware isn’t perfect.

Tom: That’s a good point. If the system were too sensitive to noise, it wouldn’t be practical. But they showed that moderate noise doesn’t hurt performance, and in some cases it even helps a little, like a regularizer.

Jane: So the improvements are threefold: parallel training, better scaling, and robustness to noise. That’s a pretty compelling package.

Tom: It really is. And I think the next question is — what does this mean for the real world? Let’s bring in Lu and Meng to talk about that.

Conclusion: Tom: Alright, we’ve covered the method, the results, and the improvements. Now let’s wrap up our discussion of “Streamlined optical training of large-scale modern deep learning architectures with direct feedback alignment.” Jane, what’s the big takeaway for our listeners?

Jane: The big takeaway is that we don’t have to be stuck with backpropagation and GPUs forever. This paper shows a real, working alternative that uses light to train massive neural networks. It’s not just a simulation — they built it, and it works.

Tom: And it’s not just a lab curiosity. They trained a billion-parameter Transformer. That’s a real milestone. It means this approach can handle the kinds of models we actually use today.

Jane: And the scaling results suggest that for future, even larger models, this approach could be faster and more energy-efficient than what we have now. That’s a big deal when we’re talking about the environmental cost of training AI.

Tom: Absolutely. And I think the most exciting part is the idea of hardware-software co-design. They didn’t just take an existing algorithm and run it on new hardware. They chose an algorithm that perfectly matches what the hardware is good at. That’s the kind of thinking that could lead to breakthroughs.

Jane: Exactly. And that’s why this paper feels different. It’s not just an incremental improvement. It’s pointing toward a different path for scaling up AI.

Tom: Well said, Jane. So we’re going to say goodbye to this paper and get ready for the next one. Thanks to everyone who tuned in. We’ll be back soon with more exciting research.

Jane: Bye, everyone!

Ziao Wang, Kilian Müller, Matthew Filipovich, Julien Launay, Ruben Ohana, Gustave Pariente, Safa Mokaadi, Charles Brossollet, Fabien Moreau, Alessandro Cappelli, Iacopo Poli, Igor Carron, Laurent Daudet, Florent Krzakala, Sylvain Gigan

École Normale Supérieure - Université PSL · Sorbonne Université · Collège de France · CNRS · LightOn · Welinq · University of Oxford · Flatiron Institute · École Polytechnique Fédérale de Lausanne (EPFL)

cs.ET, cond-mat.dis-nn, cs.LG, physics.app-ph, physics.optics

Submitted: 2025-04-02

Updated: 2026-08-18

Comments: 20 pages, 4 figures; Additional experiments conducted;

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 62/100

Key concepts

Direct Feedback Alignment (DFA)
Instead of the standard sequential backpropagation where error signals travel layer by layer, DFA sends the error directly to every layer at once using a random projection. This bypasses the sequential bottleneck, allowing for parallel training.
Optical Processor (OPU)
The paper uses an optical processor that performs random matrix multiplications. Light passing through a scattering medium naturally creates a speckle pattern, which acts as a fast, parallel random projection, matching the algorithm's needs.
Scaling Improvement
The optical method shows better scaling than backpropagation because its training time grows more slowly with network depth and width. It can train models up to two point seven billion parameters without running out of GPU memory due to storing the random matrix optically.
Hardware-Software Co-design
This concept involves choosing an algorithm that perfectly matches the capabilities of a new hardware system. The paper's success is attributed to this co-design, where the optical hardware was matched precisely to the Direct Feedback Alignment algorithm.

Terminology

Summary

Summary

This paper presents a hybrid electronic-photonic platform for training large-scale artificial neural networks using the direct feedback alignment (DFA) algorithm, termed Optical DFA (ODFA). The authors experimentally implement DFA on an Optical Processing Unit (OPU) that performs large-scale random matrix multiplications—the central operation of DFA—at speeds up to 1500 TeraOPS under 30 Watts of power. The OPU uses free-space optics: a Digital Micromirror Device (DMD) encodes the input vector (ternarized error signal) onto a laser beam, which propagates through a strongly scattering medium with a random transmission matrix, and a camera captures the output intensity. Linear random projections are recovered from intensity measurements without holography using a method involving additional projections with a constant anchor vector. The authors train modern deep learning architectures, including Transformers, with more than 1 billion parameters, achieving good performance on language, vision, and diffusion-based generative tasks.

For the language task, the authors trained a generative Transformer with 1.07 billion parameters (comparable to GPT-2) on the Cornell Movie-Dialogs Corpus. The architecture used 40 decoder blocks with an embedding dimension of 2040 and a context length of 24 tokens. Of the 1.07B parameters, 330M directly received ODFA signals as gradients, while local backpropagation was applied within each block. The training loss trajectories show that ODFA learns consistently, though somewhat more slowly than backpropagation (BP) and DFA. The ODFA-trained Transformer generated increasingly coherent conversational text as training progressed. The authors note that the quality of generated text is limited by the short context length, chosen to minimize the number of optical projections.

For vision tasks, the authors trained a Vision Transformer (ViT) with 134 million parameters on the ClimateBench dataset, mapping global distributions of four anthropogenic forcing factors to surface air temperature. The input and output dimensions were 55296 and 13824, respectively. The ODFA-trained ViT performed comparably to a BP-trained ViT under the same configuration, with similar error magnitudes. The authors also trained a fully connected neural network (FCNN) with 1.3 billion parameters on the same dataset, where 1.03 billion parameters directly received ODFA signals. Notably, the ODFA-trained FCNN outperformed the BP-trained FCNN, attributed to the natural normalization of the optical projected signal.

The authors also explored diffusion-based generative methods, specifically the Diffusion Transformer (DiT). On MNIST, ODFA achieved convergence comparable to BP, producing recognizable digits. On the more challenging Animal Faces (AFHQv2) dataset, ODFA maintained stable training and produced animal faces with recognizable species identity, though with some artifacts.

The scaling analysis focused on training time for FCNNs with synthetic data. The authors measured total wall-clock time for complete training processes. For small models, BP is faster (microseconds per sample), while ODFA has a minimum training time of a few milliseconds due to the DMD and camera frame rates. However, as model size increases, ODFA's advantage becomes apparent. For a 96-layer FCNN with 3080 neurons per layer, ODFA (13.09 ms/sample) was faster than BP (13.39 ms/sample). Using memory offloading to bypass GPU memory limits, ODFA maintained a speed advantage over BP for models up to 2.7 billion parameters (100 layers, 5200 neurons each). The ratio of ODFA to DFA training time decreases from 50 at the scale of 10 cubed × 10 cubed to 1.9 at 10 4 × 10 4, and ODFA can scale beyond DFA's dimensions because the random matrix coefficients are stored optically.

The authors conclude that this symbiotic pairing of non-standard hardware (optics) and training algorithm (DFA) provides a pathway outside the existing hardware lottery. The OPU can process input and output vectors with N 10 6 entries, an order of magnitude larger than current frontier model requirements, and could be increased further with off-the-shelf components. The approach may provide a sustainable pathway for training more capable models through increased parameter counts and architecture adaptation to ODFA's strengths.

Improvements for AI systems

Based on the paper, here are specific improvements I can implement in AI systems:

Improvement: Replace or supplement backpropagation with DFA for training deep neural networks, particularly for models with many layers.

What the improved system can do:

  • Train all layers in parallel rather than sequentially, eliminating backward locking

  • Reduce training time by up to 24% for ultra-deep networks (96+ layers)

  • Enable training of models with 2.7 billion parameters that would otherwise exceed GPU memory limits

  • Achieve comparable performance to backpropagation on vision tasks (ViT with 134M parameters showed similar error levels)

Sources

Related papers