Multilook Coherent Imaging: Theoretical Guarantees and Algorithms

arXiv:2505.23594 · stat.ML, cs.LG, eess.IV · Submitted 2026-08-17 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Multilook Coherent Imaging: Theoretical Guarantees and Algorithms".

Jane: The paper was written by Xi Chen, Soham Jana, Christopher A. Metzler, Arian Maleki and Shirin Jalali from Rutgers University and University of Notre Dame and University of Maryland and Columbia University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: So, we've just talked about the title and the scope of "Multilook Coherent Imaging: Theoretical Guarantees and Algorithms," but let's get into how they actually solve this problem. The authors are tackling the core issue of getting a precise estimate of the signal magnitude, x o, even when dealing with that nasty speckle noise.

Jane: They treat it like a constrained optimization problem, essentially trying to find the best possible image based on what we see across multiple looks. But instead of just assuming a simple structure for the image, they use this concept called the Deep Image Prior, or DIP.

Meng: The authors posit that by forcing our signal to fit within the range of an untrained neural network—the DIP—we can impose a very flexible constraint on what kind of images are possible. It’s not just any random picture; it has to look like something a deep learning model could generate.

Lu: That flexibility is key, Meng. It allows the method to capture complex structures that simpler models would miss, giving us a much richer set of plausible solutions than traditional methods allow.

Lalam: And this combination of using the DIP constraint with the likelihood function is where it becomes powerful. We are essentially telling nature what kind of image we expect while still letting the data guide us towards the best possible reconstruction.

Tom: It sounds like they’ve given us a very robust framework for dealing with these complex, noisy measurements, right? But to really appreciate how solid that framework is, let's look at the paper's summary.

Summary: Jane: The core of the summary is that this work provides a rigorous mathematical foundation for multilook coherent imaging. They aren't just offering a heuristic solution; they are establishing clear theoretical upper bounds on how much error we can expect in our reconstruction.

Meng: The theoretical contribution, as I read it, seems to be characterizing the relationship between that error and several parameters: the number of looks (L), the signal dimension (n), and the number of DIP parameters (k). It’s quantifying exactly what you can expect from a given setup.

Lu: That’s right, Lu. The paper shows that as we get more looks, or more data, we get better results. But it also highlights how much the complexity of the model matters in relation to those bounds.

Lalam: It’s a beautiful balance between complexity and certainty. We are finding the sweet spot where our AI model is complex enough to capture detail but constrained enough that our theoretical guarantees hold true.

Tom: And when they discuss the practical challenges, it's clear they identify two major hurdles: how to handle the overfitting problem inherent in DIP, and how to manage the computational burden of solving this massive likelihood problem.

Jane: It’s a classic tension between theory and implementation, Tom. You have this mathematically perfect framework, but you also have a real-world computer that needs to run it efficiently.

Meng: And since they tackle both challenges directly in their approach, I think the summary is very convincing about making this practical.

Improvements: Tom: We've seen the theoretical foundation and the summary, but now let's get into how this paper actually improves upon existing methods. It’s not just a slight tweak; it’s a whole new way of doing things.

Jane: They tackle the overfitting issue with something called Bagged-DIP, which is based on that classic statistical idea of bagging or ensemble methods. Instead of forcing one single complex DIP to do everything, they use multiple different ones and then average their outputs.

Lu: That's a brilliant way to increase robustness. By generating several weakly dependent estimates—each using a different patch size or structure—we reduce the risk that any single estimate overfits to noise or biases the result.

Meng: The second big improvement, which is vital for implementation, is how they handle the matrix inversion within Projected Gradient Descent. That matrix B changes every step, and its inverse is computationally expensive.

Tom: Exactly! And this is where their use of the Newton-Schulz algorithm comes in to save us from a computational nightmare. Instead of calculating the full inverse at every iteration, they use a few steps of Newton-Schutz to approximate it very quickly and accurately.

Jane: It’s a clever optimization, Meng. They found that just one step of the Newton-Schulz method is often enough to get extremely close to the true inverse, which is incredibly efficient for real-time or high-volume processing.

Lalam: This whole Bagged-DIP and Newton-Schutz combination suggests a new paradigm where we can leverage the immense potential of deep learning models without succumbing to their inherent instability when applying iterative optimization.

Conclusion: Tom: So, we've covered the theory, the approach, and those major technical improvements—Bagged-DIP and Newton-Schulz. It’s clear that "Multilook Coherent Imaging: Theoretical Guarantees and Algorithms" provides a complete solution to this imaging challenge.

Jane: By showing both a sharp theoretical upper bound on the MSE and an algorithm that state-of-the-art performance, it has set a new standard for how we approach speckle noise in undersampled systems.

Meng: From an engineering standpoint, this is massive because the computational savings from using Newton-Schutz mean this could run on real hardware today, not just theoretical simulations. It’s practical impact we can actually see.

Lu: And it allows for a much sharper comparison to previous work, especially when L is large, as the error decreases with more looks. The theory truly captures the best possible outcome of a sophisticated system.

Lalam: This work represents a leap in how we handle information scarcity and noise simultaneously. It shows that AI can provide the structure and flexibility needed to pull high-resolution images out of inherently imperfect data streams, which is inspiring for future applications across all fields.

Tom: We’re incredibly excited about this paper, it has real utility for everyone from radar to ultrasound imaging. We’ll be back after the break with another fascinating topic in AI.

Xi Chen, Soham Jana, Christopher A. Metzler, Arian Maleki, Shirin Jalali

Rutgers University · University of Notre Dame · University of Maryland · Columbia University

stat.ML, cs.LG, eess.IV

Submitted: 2026-08-17

Updated: 2026-08-18

Comments: 29 pages, 4 figures, 3 tables. arXiv admin note: substantial text overlap with arXiv:2402.15635

Code: https://github.com/Computational-Imaging-RU/Bagged-DIP-Speckle

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 86/100

The gist: The paper addresses the problem of signal recovery from multilook coherent imaging systems in the presence of speckle noise.

Key concepts

Deep Image Prior (DIP)
The authors use DIP as a flexible constraint for image reconstruction. This forces the signal to fit within the range of an untrained neural network, ensuring that generated images resemble plausible structures rather than random noise.
Bagged-DIP
To solve the overfitting problem inherent in DIP, Bagged-DIP uses multiple different DIP models and averages their outputs. This increases robustness by generating several weakly dependent estimates, reducing the risk of any single estimate being biased by noise.
Newton-Schulz Algorithm
This algorithm is used to approximate a computationally expensive matrix inversion within Projected Gradient Descent. By using just one step of Newton-Schutz, the method achieves high accuracy and efficiency, allowing for practical implementation in real-time processing.

Terminology

Summary

The paper addresses the problem of signal recovery from multilook coherent imaging systems in the presence of speckle noise. The imaging model is given by:

y = AXo w + z (Equation 1)

where Xo = diag(xo) with xo ∈ Cn denoting the complex-valued signal of interest, w ∈ Cn represents speckle (multiplicative) noise with iid CN(0, σw2In) components, and z ∈ Cm denotes additive noise modeled as iid CN(0, σz2). The paper explores the scenario where m ≤ n, allowing imaging systems to capture higher resolution images than constrained by the number of sensors.

For an L-look system, the measurements at look l are represented as:

yl = AXo wl + zl

where w1,..., wL ∈ Cn and z1,..., zL ∈ Cm denote independent speckle noise and additive noise vectors, respectively. The measurement kernel A remains constant across looks.

Since fully-developed noises are complex-valued Gaussian with uniform phases, the phase of xo cannot be recovered. Hence, the goal is to obtain a precise estimate of xo based on L observations, assuming xo is real-valued.

The standard approach is to minimize the negative log-likelihood function subject to signal structure constraints:

x̂ = argmin fL(x) (Equation 2)

where C represents the set of conceivable images and:

fL(x) = log det(B(x)) + (1/L) Σl=1 L ỹl T (B(x))−1 ỹl (Equation 3)

with B(x) defined as a 2m × 2m matrix involving σz2, σw2, and U(x) = AX2Ā T.

The paper works under the Deep Image Prior (DIP) hypothesis: Natural images can be embedded within the range of untrained neural networks that have substantially fewer parameters than the total number of pixels, and use iid noises as inputs. The set C is defined as the range of a deep image prior, where for every x ∈ C, there exists θ ∈ Rk such that x = gθ(u), with u generated iid N(0,1) and θ denoting the parameters of the DIP neural network.

Theorem III.1: Let the elements of the measurement matrix Aij be iid N(0,1). Suppose m < n and that the function gθ(u), as a function of θ ∈ Bk(0, xmax√(n/k)), is Lipschitz with Lipschitz constant 1. Then:

(1/n)∥x̂ − xo∥22 = C1 · (√(k log n)/m2) · (n/L) + (k log n)/m (Equation 6)

with probability 1 − C2e(−m/2) + e(−Ln/8) + e(−C3k log n) + e(k log n − n/2) for some constants C1, C2, C3 > 0.

  1. Dependence on DIP parameters (k): As k increases (while keeping m, n, and L fixed), both error terms grow, aligning with intuition that more parameters make distinguishing between diverse alternatives more challenging.

  2. First error term: The term (√(k log n)/m2) · (n/L) grows rapidly as a function of n, contrasting with logarithmic growth in additive noise models. This is attributed to the increasing number of speckle noise elements as n increases. The dependence on L aligns with the parametric error rate of 1/L for estimation problems with L samples.

  3. As L → ∞: The first term converges to zero, and the dominant term becomes √(k log n)/m. This is consistent with the intuition that even with infinite looks, recovery is limited by m, since (1/L)Σl ylyl T ≈ AXo2A T provides m(m+1)/2 linear measurements of Xo2, and accurate recovery requires m2 ≫ k log n.

The paper notes that a much weaker version of Theorem III.1 appeared in a shorter version published at ICML, with the leading term being (√(k log n)/m) · √(n/(Lm)), which is substantially looser than the current bound.

Compared to the results in [11], the current paper's bounds are significantly sharper. When k log n/m ≪ 1, the upper bound simplifies to k log n/m. The paper identifies that one source of looseness in [11] arises from working with the expected log-likelihood, which the current proof strategy entirely avoids.

The proof involves four key steps:

  1. Lower bounding f̄(Σ̂) − f̄(Σo): Using Lemma VI.6, the paper establishes that f̄(Σ̂) − f̄(Σo) ≥ Tr(Σo−1ΔΣΣo−1ΔΣ)/(2(1 + λmax)2), where λmax is bounded using concentration results for the singular values of A.

  2. Upper bounding the noise term: Using a δ-net argument and the Hanson-Wright inequality, the paper bounds the deviation of the empirical likelihood from its expectation.

  3. Upper bounding Tr(Σo−1ΔΣΣo−1ΔΣ): Combining the bounds from steps 1 and 2, the paper derives an upper bound on this trace quantity.

  4. Lower bounding Tr(Σo−1ΔΣΣo−1ΔΣ) in terms of ∥x̂ − xo∥22: Using Lemma VI.7 and concentration arguments involving decoupling of U-processes (Theorem VI.5), the paper establishes a high-probability lower bound.

The paper proposes using projected gradient descent (PGD) to solve the optimization problem:

x t+1 = Proj(x t − μ t ∇fL(x t)) (Equation 39)

where Proj(·) projects onto the range of gθ(u). The projection is implemented by training the DIP to minimize ∥gθ(u) − (x t − μ t ∇fL(x t))∥ using Adam.

The paper identifies a fundamental challenge: sophisticated networks fit clean images well but are susceptible to overfitting when images are corrupted with noise, while simpler networks are less susceptible to noise but do not fit clean images well. This creates a dilemma for DIP-based PGD.

The paper proposes a bagging strategy: Rather than finding the right complexity level for the DIP at each iteration, which is a computationally demanding and statistically challenging problem, we use bagging. The approach involves:

  • Selecting a network sophisticated enough to fit real-world images

  • Partitioning images into non-overlapping patches of sizes (hk × wk)

  • Training independent DIPs on each patch

  • Placing patches back to obtain estimates x̌k

  • Iterating for K different patch sizes to obtain K estimates

  • Averaging the individual estimates

The paper uses K = 3 with patch sizes h1 = w1 = 32, h2 = w2 = 64, and h3 = w3 = 128.

The gradient computation requires inverting a large matrix B ∈ R2m×2m at each iteration. The paper proposes using the Newton-Schulz algorithm:

M k = M k-1 + M k-1(I − B t M k-1) (Equation 42)

with M0 = (B t-1)−1 from the previous iteration. The paper demonstrates that one step of Newton-Schulz is sufficient, significantly reducing computational complexity.

The paper evaluates performance on 8 test images (Barbara, Peppers, House, Foreman, Boats, Parrots, Cameraman, Monarch) at sampling rates m/n = 0.125, 0.25, 0.5, with L = 1, 2, 4, 8, 16, 32, 64, 128 looks.

Key findings:

  1. Newton-Schulz effectiveness: One step of Newton-Schulz yields results virtually identical to exact inverse computation, while being significantly faster (e.g., 1e-4 seconds vs 52.8 seconds for 128×128 images).

  2. Bagged-DIP improvement: Bagging with three estimates offers between 0.5dB and 1dB improvement over individual estimates.

  3. Simple vs. Bagged-DIP: Bagged-DIP-based PGD significantly outperforms PGD with simple DIPs, particularly as L increases, where simple networks' low accuracy becomes a bottleneck.

  4. State-of-the-art performance: For L = 100, L = 50, and L = 25, the algorithm outperforms the vanilla PGD proposed in [25] by 1.09 dB, 1.47 dB, and 1.27 dB respectively on average.

The paper includes several technical lemmas supporting the main theorem:

  • Lemma VI.1: Bounds eigenvalues of B−1 − C−1 in terms of σmax(B−C), σmin(B), and σmin(C)

  • Lemma VI.2: Concentration of singular values of random Gaussian matrices

  • Lemma VI.3: Concentration of chi-squared random variables

  • Theorem VI.4: Hanson-Wright inequality for quadratic forms

  • Theorem VI.5: Decoupling of U-processes

  • Lemma VI.6: Lower bound on f̄(Σ̂) − f̄(Σo)

  • Lemma VI.7: Bounds on Tr(Σ−1ΔΣΣ−1ΔΣ) in terms of ∥A(X̂2 − X2)A T∥2 HS

  • Lemma VI.8: Concentration bound for the empirical likelihood deviation

  • Lemma VI.9: Bounds on eigenvalues of differences of inverse matrices

  • Lemma VI.10: Bounds on differences of trace terms

  • Lemma VI.11: Concentration bound for ∥ADA T∥2 HS

The paper provides the first theoretical upper bound on MSE for multilook coherent imaging under the DIP hypothesis, capturing the dependence on the number of DIP parameters, number of looks, signal dimension, and measurements per look. The proposed Bagged-DIP PGD algorithm with Newton-Schulz matrix inversion achieves state-of-the-art performance.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with the resulting capabilities.

1. New Loss Function for Robust Image Recovery

  • Improvement: I will implement the negative log-likelihood function fL(x) defined in Equation (3) as a custom loss function for neural networks. This loss explicitly models the complex covariance structure of speckle and additive noise, rather than relying on simpler, less accurate L2 or L1 losses.

  • Capability: The AI system can now be trained to recover images from measurements corrupted by fully-developed speckle noise, even when the measurement matrix A is ill-conditioned or under-sampled (m < n). This is a significant upgrade over standard denoisers, which fail when the noise is not independent and identically distributed.

2. Bagged-DIP Projection for Enhanced Stability

  • Improvement: I will replace the standard single Deep Image Prior (DIP) projection in Projected Gradient Descent (PGD) with the proposed Bagged-DIP strategy. This involves training multiple DIPs on different patch sizes (e.g., 32x32, 64x64, 128x128) and averaging their outputs.

  • Capability: This mitigates the overfitting problem of complex DIPs to noisy inputs and the bias problem of simple DIPs. The AI system will be more stable and achieve higher reconstruction quality (PSNR/SSIM) across a wide range of noise levels and undersampling ratios, as demonstrated in Figure 3 and Table I of the paper.

3. Efficient Matrix Inversion via Newton-Schulz

  • Improvement: I will integrate the Newton-Schulz algorithm into the gradient computation step of PGD. Instead of computing the exact inverse of the large matrix B (size 2m x 2m) at every iteration, I will use the inverse from the previous iteration as a starting point and perform only a single Newton-Schulz update.

  • Capability: This reduces the computational cost of each PGD iteration from O(m 3) to O(m 2) (as shown in Table II, this is a 100-1000x speedup for 128x128 images). The AI system can now process larger images and higher-resolution measurements in real-time or near-real-time, without sacrificing reconstruction accuracy.

4. Adaptive Inverse Update Criterion

  • Improvement: I will implement the adaptive criterion ∥xt - xt-1∥∞ < δx (with δx = 0.12) to decide when to use the exact matrix inverse versus the Newton-Schulz approximation.

  • Capability: The system will automatically use the expensive exact inverse only for the first few iterations (when the estimate is changing rapidly) and switch to the fast approximation for the remaining iterations. This ensures both convergence and efficiency, making the algorithm robust to initialization and image complexity.

5. Theoretically Grounded Hyperparameter Selection

  • Improvement: I will use the theoretical bound from Theorem III.1 to guide the selection of the DIP's capacity (k) and the number of looks (L). The bound shows that the error scales as O(k log n / m squared + n k log n / (L m 2)).

  • Capability: The AI system can automatically determine if adding more looks (L) will be beneficial or if the bottleneck is the DIP's capacity (k). This allows for principled resource allocation (e.g., deciding between acquiring more data vs. using a more complex model) to achieve a target reconstruction error.


The improved AI system will be a high-performance, computationally efficient, and theoretically sound solver for inverse problems in coherent imaging. Specifically, it can:

  • Recover high-resolution images from low-resolution, speckle-corrupted measurements (e.g., in Synthetic Aperture Radar, digital holography, and ultrasound).

  • Operate in the challenging under-sampled regime (m < n) where traditional methods fail.

  • Adapt to varying numbers of looks (L) to improve accuracy, with performance gains that align with theoretical predictions.

  • Run significantly faster than existing DIP-based PGD methods, making it suitable for real-time applications.

  • Provide a reliable estimate of the achievable error based on system parameters (m, n, k, L), enabling better system design and uncertainty quantification.

Sources

Related papers