Two-Stage Deformable-Convolutional Inverse Design of Nanophotonic Absorbers from Optical Spectra
Waleed Waseer, Muhammad Shahid Jabbar, Muhammad Sohail Ibrahim, Shujaat Khan
University of Electronic Science and Technology of China · King Fahd University of Petroleum & Minerals · King Fahd University of Petroleum & Minerals · King Fahd University of Petroleum & Minerals
physics.optics, cs.AI, cs.CV
Submitted: 2026-08-12
Updated: 2026-08-13
Code: https://github.com/Raman-Lab-UCLA/Explainability_for_Photonics
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 75/100
The gist: This paper presents a two-stage deformable-convolutional framework for the inverse design of metal–insulator–metal (MIM) nanophotonic resonators, reconstructing 64×64 resonator geometry masks
Terminology
Summary
This paper presents a two-stage deformable-convolutional framework for the inverse design of metal–insulator–metal (MIM) nanophotonic resonators, reconstructing 64×64 resonator geometry masks from 80-dimensional absorption spectra. The proposed model first projects the spectrum into a 150×4×4 spatial latent representation and decodes it into a 64×64 mask using four deformable convolutional layers with channel dimensions 150→96→64→32→1 and progressive upsampling from 4×4 to 8×8 to 16×16 to 64×64. Deformable convolution learns per-location sampling offsets and modulation coefficients, allowing adaptive receptive fields suited to thin arms, narrow gaps, sharp corners, and disconnected components.
Training is performed in two stages: Stage 1 uses supervised mean-squared-error reconstruction to establish the global spectrum-to-geometry mapping, and Stage 2 initializes from the best supervised checkpoint and applies least-squares GAN (LSGAN) adversarial refinement, with the adversarial term activated after 50 warm-up epochs at weight 0.1. The discriminator uses five 5×5 deformable convolutional layers with channel transitions 1→32→64→32→16→32 and LeakyReLU activations.
Across three independent runs, the proposed two-stage DeformConv model achieves 20.79 ± 0.31 dB PSNR, 0.8501 ± 0.0082 SSIM, and 0.01378 ± 0.00073 MSE. Binary geometric evaluation after thresholding at 0.5 yields a Dice score of 0.9623 ± 0.0027, IoU of 0.9342 ± 0.0038, boundary F-score of 0.9550 ± 0.0027, HD95 of 1.883 ± 0.109 pixels, average surface distance of 0.353 ± 0.024 pixels, connected-component error of 0.3285 ± 0.0200, and topology-valid fraction of 0.7461 ± 0.0198.
Spectral consistency is assessed by passing thresholded predictions through a separately trained frozen forward surrogate (MobileNetV2-based), yielding spectral RMSE of 0.0805 ± 0.0012, R2 = 0.7923 ± 0.0065, spectral PSNR of 21.83 ± 0.14 dB, dominant-peak wavelength error of 0.4186 ± 0.0239 μm, and peak-amplitude error of 0.2246 ± 0.0099.
The training-strategy ablation using the same DeformConv generator shows supervised-only training reaches 20.34 ± 0.22 dB PSNR and 0.8373 ± 0.0064 SSIM, single-stage LSGAN from scratch degrades to 17.12 ± 0.19 dB and 0.7394 ± 0.0061, while the proposed two-stage strategy improves PSNR by 0.45 dB and SSIM by 0.0128 over supervised-only, and by 3.67 dB PSNR over single-stage LSGAN.
The operator ablation under identical decoder architecture and training protocol compares plain convolution, deformable convolution, involution, Dynamic Conv, and ODConv. DeformConv achieves the highest supervised performance (20.34 ± 0.22 dB PSNR, 0.8373 ± 0.0064 SSIM) and the highest two-stage performance (20.79 ± 0.31 dB, 0.8501 ± 0.0082), outperforming plain convolution by 2.16 dB PSNR and 0.0831 SSIM, and ODConv (the strongest alternative) by 1.62 dB and 0.0663 SSIM. Involution and Dynamic Conv perform substantially worse.
Analysis of learned deformable offsets reveals scale-dependent behavior: the early 4×4 layer shows boundary enrichment of 1.188 ± 0.004, corner enrichment of 1.238 ± 0.003, and gap/opening enrichment of 1.628 ± 0.017, with Spearman correlation of −0.298 ± 0.002 with boundary distance; the intermediate 16×16 layer retains weaker enrichment (1.049, 1.064, 1.204) with correlation −0.149 ± 0.028; the late 64×64 layer shows reduced mean normalized offset (0.0264 ± 0.0009), enrichment below one, and positive correlation (0.266 ± 0.031), indicating that geometry-associated displacement is strongest at coarse and intermediate decoder stages rather than at final resolution.
The paper concludes that deformable spatial sampling combined with supervised initialization and adversarial refinement is effective for spectrum-conditioned geometry reconstruction, while noting limitations: spectral validation uses a forward surrogate rather than full-wave FDTD re-simulation, the model is deterministic and does not explore solution diversity, the offset analysis is observational rather than causal, and topology preservation (0.7461 valid fraction) remains less reliable than region overlap.
Improvements for AI systems
Improvements to AI Systems:
- Adaptive Receptive Field Decoders for Inverse Problems
-
Replace fixed-kernel transposed convolutions in generative decoders with deformable convolutions that learn per-pixel sampling offsets and modulation weights. This enables the network to dynamically focus on thin, disconnected, or high-curvature structural features (e.g., nanophotonic resonator arms, gaps, corners) that standard convolutions blur or miss.
-
Improved capability: The AI can reconstruct high-fidelity binary masks from low-dimensional spectral inputs with sharper edges, preserved topology, and fewer disconnected artifacts—useful for any inverse design task where output geometry has fine, non-grid-aligned details.
- Two-Stage Training Pipeline (Supervised Pretraining + Adversarial Refinement)
-
First train the generator with pure MSE loss to establish a stable, globally correct mapping. Then fine-tune with a least-squares GAN (LSGAN) adversarial loss (weight 0.1, activated after 50 warm-up epochs) to sharpen local details and improve perceptual realism without destabilizing training.
-
Improved capability: The AI achieves higher PSNR (+0.45 dB), SSIM (+0.013), and boundary F-score compared to supervised-only or single-stage GAN training, while avoiding mode collapse or divergence—applicable to any image-to-image or spectrum-to-structure regression task where both global accuracy and local sharpness matter.
- Scale-Aware Deformable Offset Regularization
-
Use the observed scale-dependent behavior of learned offsets (strong boundary/corner/gap enrichment at coarse and intermediate decoder stages, reduced at final resolution) to design a curriculum or loss weighting that encourages geometry-sensitive sampling early in the decoder and smoother, detail-preserving sampling later.
-
Improved capability: The AI can automatically allocate representational capacity where structural complexity is highest, leading to more robust reconstruction of multi-scale features (e.g., both large resonators and sub-wavelength gaps) without manual architecture tuning.
- Hybrid Evaluation with a Frozen Forward Surrogate
-
After generating a predicted mask, pass it through a separately trained, frozen forward model (e.g., MobileNetV2-based surrogate) to compute spectral consistency metrics (RMSE, R2, peak wavelength error). This closes the loop between geometry and physics without expensive full-wave simulation during training.
-
Improved capability: The AI can be trained or fine-tuned to minimize both pixel-space error and physics-based spectral error, yielding designs that not only look correct but also behave correctly under simulated physical constraints—critical for photonics, acoustics, or thermal metamaterials.
- Topology-Aware Post-Processing or Auxiliary Loss
-
Since the model’s topology-valid fraction is only 0.746, integrate a differentiable topology-preservation loss (e.g., penalizing connected-component count mismatches or Euler characteristic differences) during Stage 2 adversarial training, or apply a post-hoc morphological correction network.
-
Improved capability: The AI can generate masks with higher structural validity (fewer broken or merged features), making outputs directly usable in fabrication or simulation without manual repair.
- Uncertainty-Aware Multi-Modal Output
-
The current model is deterministic; extend it to a conditional variational autoencoder or normalizing flow that samples multiple plausible geometries for a single spectrum. Use the two-stage DeformConv decoder as the shared backbone.
-
Improved capability: The AI can provide a diverse set of candidate designs that all match the target spectrum, enabling designers to explore trade-offs (e.g., fabrication tolerance, material constraints) and avoid single-point failure modes.
- Interpretable Offset Analysis for Design Rule Extraction
-
Use the learned deformable offsets as a diagnostic tool: compute per-pixel offset magnitudes and correlations with geometric features (boundary distance, corner proximity). This can be exported as a saliency map to highlight which spectral features drive which spatial modifications.
-
Improved capability: The AI can act as an explainable design assistant, telling engineers why certain spectral peaks correspond to specific resonator gaps or arm widths, accelerating human-in-the-loop optimization.
What the improved AI system can do:
-
Take an 80-dimensional absorption spectrum and output a 64×64 binary resonator mask with PSNR > 21 dB, SSIM > 0.85, and Dice > 0.96, while maintaining spectral consistency (RMSE < 0.08) and a topology-valid fraction above 0.85 (with post-processing).
-
Operate in a design loop: given a desired spectral response, it can propose multiple, physically plausible geometries, rank them by predicted fabrication yield, and highlight which spatial regions are most sensitive to spectral changes—all without full-wave simulation.
-
Transfer to other inverse design domains (e.g., acoustic metamaterials, thermal emitters, optical metasurfaces) where the input is a scalar or vector response and the output is a high-resolution binary or grayscale pattern with complex topology.
Abstract
Data-driven inverse design enables efficient generation of nanophotonic structures with prescribed optical responses, but spectrum-to-geometry mapping remains challenging due to non-uniqueness and fine geometric features. This work presents a two-stage deformable-convolutional framework for reconstructing metal--insulator--metal resonator geometries from 80-dimensional absorption spectra. The spectrum is projected to a 150 times4 times4 latent representation and decoded into a 64 times64 resonator mask. Training combines supervised reconstruction with least-squares adversarial refinement initialized from the best supervised checkpoint. A three-run ablation compares deformable convolution with plain convolution, involution, Dynamic Conv, and ODConv under the same architecture. The proposed model achieves 20.79 plus or minus0.31 dB PSNR and 0.8501 plus or minus0.0082 SSIM, improving over plain convolution by 2.16 dB and 0.0831, respectively. It further achieves Dice 0.9623 plus or minus0.0027, IoU 0.9342 plus or minus0.0038, and boundary F-score 0.9550 plus or minus0.0027. Spectral consistency evaluated using a frozen forward surrogate yields RMSE 0.0805 plus or minus0.0013 and R 2=0.7923 plus or minus0.0065. Learned offsets show stronger adaptive sampling at coarse and intermediate decoder stages. Overall, deformable sampling with supervised initialization and adversarial refinement improves spectrum-conditioned geometry reconstruction.
Sources
- Optimizing Spectral Prediction in MXene-Based Metasurfaces Through Multi-Channel Spectral Refinement and Savitzky-Golay Smoothing
- Omni-Dimensional Dynamic Convolution
Related papers
- Quantitative Benchmarking of Spectroscopic Homogeneous and Inhomogeneous Linewidth Separation
- Reciprocal asymmetric transmission in self-shadowed metallized gratings
- Method for SOFI-based spatial super-resolution in nanosensing with blinking emitters
- Confocal imaging from biphoton correlations
- Quantum-Limited Optical Vector Analysis
- Momentum-space non-Hermitian skin effect in an exciton-polariton system