PURe: A Plug-and-Play Product-Unit Residual Module for Vision Networks

arXiv:2505.04397 · cs.CV, cs.AI, cs.LG, eess.IV · Submitted 2026-08-06 · Read on arXiv

Ziyuan Li, Uwe Jaekel, Babette Dellen

Department of Mathematics, Informatics and Technology, University of Applied Sciences Koblenz · Technical University of Munich

cs.CV, cs.AI, cs.LG, eess.IV

Submitted: 2026-08-06

Updated: 2026-08-10

Comments: Accepted to the GCPR 2026

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 66/100

The gist: The paper "PURe: A Plug-and-Play Product-Unit Residual Module for Vision Networks" proposes a new architectural component designed to incorporate explicit multiplicative local interactions into deep

Terminology

Summary

The paper PURe: A Plug-and-Play Product-Unit Residual Module for Vision Networks proposes a new architectural component designed to incorporate explicit multiplicative local interactions into deep vision networks. The authors observe that modern vision networks are dominated by additive local transformations, whereas explicit multiplicative local interactions remain underexplored. While product units can model structured nonlinear dependencies, their integration into deep architectures has historically been limited by optimization instability because multiplicative transformations are highly sensitive to input scale, can induce extreme gradients, and are notoriously unstable in deep architectures.

To address these challenges, the authors introduce PURe, a product-unit Residual Module for deep vision networks. The core of the module is a 2D product unit with a real-valued log-domain formulation that makes multiplicative local aggregation practical within deep residual hierarchies. The 2D product unit performs multiplicative local aggregation as follows:

y(i, j) = product m=-r r product n=-r r x(i + m, j + n) w(m,n)

For efficient and stable implementation with real-valued feature maps, the authors utilize a log-domain formulation:

y(i, j) = (sum m=-r r sum n=-r r w(m, n) ((x(i + m, j + n), tau theta)))

In this formulation, tau theta = softplus(theta) + 10-7 is a trainable threshold used to clamp inputs, which keeps the computation in the real-valued domain while reducing the destabilizing effect of extremely small inputs.

The PURe module is designed as a plug-and-play residual module which replaces native spatial residual transformations with an explicit multiplicative alternative. It is embedded into a residual design where the shortcut pathway [is kept] unchanged, preserving the optimization benefits of residual learning. Crucially, within the residual branch, the intermediate ReLU activations inside the original residual branch are removed, since interleaved pointwise nonlinearities would interfere with the intended multiplicative transformation.

The module can be instantiated in various architectures:

  • ResNet-style classifiers: In basic-block configuration, the second 3 × 3 convolution in the residual branch is replaced by a 2D product unit. In the bottleneck configuration, the two surrounding 1 × 1 convolutions are preserved, and only the central spatial-mixing layer is replaced by a 2D product unit.

  • Residual U-Net (ResUNet): For slice-based segmentation on volumetric CT data, the replacement principle is applied to the native residual units in the encoder, bottleneck, and decoder.

Experimental Results:

The authors demonstrate the effectiveness of PURe across several benchmarks:

  • Galaxy10 DECaLS (Image Classification): PURe consistently improves residual CNNs and yields a more favorable accuracy–parameter trade-off. Specifically, PURe-ResNet34 achieves the strongest cleanset performance, outperforming all ResNet baselines while using substantially fewer parameters than much deeper models such as ResNet152. The models also show improved robustness under Poisson corruption and faster optimization speed.

  • ImageNet (Image Classification): PURe improves performance across different scales. PURe-ResNet18 exceeds the reported ResNet34 baseline with roughly half the parameters, and PURe-ResNet50 achieves the strongest overall performance, exceeding the much deeper ResNet152 baseline while using less than half the parameters.

  • CIFAR-10 (Image Classification): PURe-based models consistently outperform CIFAR-style ResNet counterparts across all evaluated depths. Notably, PURe-ResNet272 approaches the much deeper ResNet1001 baseline while using less than half the parameters.

  • AMOS (Medical Image Segmentation): In slice-based CT segmentation, PURe-ResUNet consistently outperforms the baseline ResUNet on the AMOS CT validation set. The improvement in case-wise mDice is statistically significant under paired testing, and the module improves all 15 anatomical structures, particularly for structures with complex boundaries, limited spatial extent, or strong contextual ambiguity.

In conclusion, the authors state that explicit multiplicative local interaction is a practical and effective design primitive for deep residual vision networks, and PURe provides a simple plug-and-play way to enrich local feature transformation within established residual architectures.

Improvements for AI systems

1. Parameter-Efficient Model Compression for Edge Deployment

  • Improvement: Replace standard 3×3 convolutional layers in ResNet-style architectures (both Basic-block and Bottleneck configurations) with PURe modules using the log-domain formulation.

  • What the improved system can do: It can run high-accuracy computer vision tasks on resource-constrained edge devices by using significantly smaller models (e.g., a PURe-ResNet18) that match or exceed the accuracy of much larger, parameter-heavy models (e.g., ResNet34 or ResNet50), effectively reducing memory footprint and computational latency.

2. High-Precision Medical Image Segmentation

  • Improvement: Integrate PURe modules into the encoder, bottleneck, and decoder of Residual U-Net (ResUNet) architectures.

  • What the improved system can do: It can perform highly accurate volumetric segmentation of medical data (such as CT scans), specifically improving the delineation of anatomical structures that are difficult to capture due to complex boundaries, limited spatial extent, or high contextual ambiguity.

3. Robustness to Stochastic Sensor Noise

  • Improvement: Implement PURe-based residual blocks in vision pipelines intended for low-light or high-noise environments.

  • What the improved system can do: It can maintain high classification and detection accuracy even when images are subjected to Poisson corruption (common in photon-counting sensors), providing more stable performance in unpredictable real-world lighting conditions.

4. Enhanced Non-linear Feature Modeling

  • Improvement: Reconfigure residual branches by removing interleaved pointwise ReLU activations and replacing the spatial-mixing layer with the 2D product unit.

  • What the improved system can do: It can capture complex, structured non-linear dependencies between local pixels through multiplicative aggregation, allowing the network to model sophisticated spatial relationships and feature interactions that standard additive-only convolutions fail to represent.

Sources

Related papers