A Survey on Industrial Anomaly Synthesis
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "A Survey on Industrial Anomaly Synthesis".
Jane: The gist: This survey comprehensively reviews anomaly synthesis methodologies, introducing the first industrial anomaly synthesis (IAS) taxonomy and exploring cross-modality synthesis and large-scale Vision Language Models (VLM) to boost IAS.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we're looking at this paper called "A Survey on Industrial Anomaly Synthesis." It sounds like a big review of what’s out there in the field.
Jane: Yeah, it does. The authors, Yanshu Wang and the rest of their team, they are setting up a way to look at all these different ways people create anomalies for industrial stuff.
Lu: I think what's really important here is that they're not just listing methods; they’re trying to put them into a structure. They call it an IAS taxonomy.
Meng: A taxonomy means they are organizing the existing techniques into categories so we can actually compare them systematically, which is helpful for practical engineering work.
Lalam: It’s about getting a clear picture of where everyone is in this area, showing the progression from older methods to newer ones.
Tom: Exactly. We're going to talk about what that taxonomy actually looks like and why it matters for understanding the whole landscape of industrial anomaly synthesis.
The paper's summary: Jane: So, this survey goes through about forty representative methods across four main types of approaches. They break them down into Handcrafted, Distribution hypothesis-based, Generative model based, and Vision Language Model based synthesis.
Tom: Forty methods is a lot to take in. It sounds like they've mapped out the entire area from traditional rules to these massive AI models we see today.
Lu: The authors say this taxonomy is designed to reflect methodological progress and how these techniques actually work in practice, which is what makes it useful for researchers moving forward.
Meng: So, when you look at the Handcrafted part, that means they are manually designing rules, right? Like telling the AI exactly how to crop an image or add noise.
Lalam: Right. And then there’s the Distribution Hypothesis based approach, which is all about modeling what "normal" data looks like statistically and then making things that fall outside of that normal.
Tom: It’s a really broad overview, but the key point is they are showing how these different methods relate to each other within this new framework.
Jane: And they specifically point out that previous surveys missed the multimodal stuff and Vision Language Models, so this paper tries to bring those into the main discussion.
The paper's improvements: Tom: They suggest a few key things for future research, which is where it gets interesting. They focus on boosting diversity by using uncertainty-aware models for anomaly synthesis.
Jane: That sounds like they want to make sure the anomalies aren't just one single thing; they want more variety in what kind of abnormal stuff we can generate.
Lu: And another big push is controllability—they suggest focusing on cross-class consistency modeling so that when you change an anomaly, you can precisely control its shape, texture, and how it’s distributed.
Meng: From an engineering standpoint, controllability is crucial because if we need a specific type of failure to test for in a factory setting, we need tools that let us dial in those exact attributes.
Lalam: They also emphasize promoting multimodal anomaly synthesis by developing alignment strategies between different data types, like using VLM and other multimodal transformers together.
Tom: So they are really pushing the boundaries of what's possible by looking at uncertainty and cross-modal learning to make these syntheses much more useful.
Conclusion: Jane: Wrapping up this survey, the authors emphasize that tackling these challenges will really improve how effective industrial anomaly synthesis methods can be used in the real world.
Tom: They’ve established this unified framework for systematic analysis, which is a big step because it gives everyone a common language to discuss what's working and what isn't.
Lu: It’s the first dedicated work that systematically categorizes these methods, which means anyone starting in this space has a solid foundation now.
Meng: For engineers, this means they can actually pick the right synthesis method based on whether they need high structural detail or just broad coverage of distribution patterns.
Lalam: It’s a comprehensive overview of what's available right now, covering about forty representative methods across those four main paradigms we talked about.
Tom: So, to sum up this paper "A Survey on Industrial Anomaly Synthesis," it gives us the necessary structure to move past just listing techniques and start building better systems.
Jane: It’s a solid roadmap for where the field needs to go next as we look at integrating multimodal learning and large scale Vision Language Models into these tasks.
Shanghai Jiao Tong University · City University of Hong Kong · Department of Intelligent Manufacturing, CATL
cs.CV, cs.CE
Submitted: 2025-02-23
Updated: 2026-10-08
Code: https://github.com/M-3LAB/awesome-anomalysynthesis
Importance score: 75/100
The gist: The gist: This survey comprehensively reviews anomaly synthesis methodologies, introducing the first industrial anomaly synthesis (IAS) taxonomy and exploring cross-modality synthesis and large-scale
Key concepts
- Handcrafted Synthesis
- This method uses manually designed rules to simulate anomalies. It involves simple manipulations like cropping, rearrangement, or using external texture libraries to create deviations from the original image. Inpainting is another technique where local areas are masked and filled with noise or black patches to disrupt structural continuity.
- Distribution Hypothesis-based Synthesis
- This approach models normal data statistically to synthesize anomalies through controlled perturbations. Prior-dependent synthesis uses geometric assumptions about the data's structure, while data-driven synthesis extracts features from the latent space to create anomalies by applying statistical perturbations or adaptive strategies.
- Generative Model (GM)-based Synthesis
- This category includes advanced deep learning models like GANs and diffusion methods for realistic synthesis. It is split into full-image synthesis, which learns mappings from noise to abnormal samples; full-image translation, which maps normal images to abnormal ones while keeping the global structure; and local anomaly synthesis, which replaces specific regions with learned anomalies.
- VLM-based Synthesis
- This method utilizes large Vision Language Models (VLMs) with extensive pre-trained knowledge and multimodal cues. VLMs can produce high-quality, context-aware, and detailed abnormal samples in a single stage or through multi-stage pipelines. This approach leverages the model's integrated understanding of both visual and textual information to synthesize complex anomalies.
Terminology
Summary
The gist: This survey comprehensively reviews anomaly synthesis methodologies, introducing the first industrial anomaly synthesis (IAS) taxonomy and exploring cross-modality synthesis and large-scale Vision Language Models (VLM) to boost IAS.
Taxonomy of IAS
The survey introduces a unified review covering about 40 representative methods across four main paradigms: Handcrafted, Distribution-hypothesis-based, Generative models (GM)-based, and Vision language models (VLM)-based synthesis <ref:2502.16412#pg4>. This taxonomy provides a fine-grained framework reflecting methodological progress and practical implications <ref:2502.16412#pg4>. The four main paradigms are Hand-crafted synthesis, Distribution hypothesis-based synthesis, Generative models (GM)-based synthesis, and Vision language models (VLM)-based synthesis <ref:2502.16412#pg4>.
Hand-crafted Synthesis
Hand-crafted synthesis relies on manually designed rules to simulate anomalies <ref:2502.16412#pg5>. Self-contained synthesis manipulates the original image through operations like cropping or rearrangement to simulate texture misalignment or color variation <ref:2502.16412#pg5>. External-dependent synthesis employs external data (e.g., texture libraries) to synthesize anomalies independently of the original image, ensuring abnormal parts are not confined by its content <ref:2502.16412#pg5>. Inpainting-based synthesis removes local areas through masking techniques, disrupting structural continuity by adding noise or black patches to generate anomalies <ref:2502.16412#pg5>.
Distribution Hypothesis-based Synthesis
Distribution hypothesis-based synthesis relies on statistical modeling of normal data distributions and synthesizes anomalies through controlled perturbations <ref:2502.16412#pg5>. Prior-dependent synthesis uses the pre-defined geometric assumptions (e.g., manifold or hypersphere structures) to define normal data distributions in feature space, applying controlled deviations to ensure the synthesized feature-level anomalies lie at the boundary or outside the normal distribution <ref:2502.16412#pg5>. Datadriven synthesis leverages intrinsic statistical properties of data by extracting features in the latent space and synthesizing anomalies through perturbations or adaptive strategies <ref:2502.16412#pg5>.
GM-based Synthesis
Recent advancements in deep GM, such as GANs and diffusion methods, enable realistic anomaly synthesis <ref:2502.16412#pg7>. GM-based synthesis is categorized into Full-image synthesis, Full-image translation, and Local anomalies synthesis <ref:2502.16412#pg5>. Full-image synthesis learns abnormal data distributions and constructs the mapping from random noise to abnormal samples <ref:2502.16412#pg7>. Full-image translation uses domain translation techniques to map normal images to abnormal ones, injecting anomalies while preserving the global structure <ref:2502.16412#pg7>. Local anomaly synthesis replaces specific regions of normal images with learned local anomalies and ensures smooth transitions between abnormal areas and the background <ref:2502.16412#pg7>.
VLM-based Synthesis
Leveraging large-scale pre-trained VLMs with billions of parameters, VLM-based synthesis exploits extensive pre-trained knowledge and integrated multimodal cues to synthesize high-quality anomalies <ref:2502.16412#pg8>. Single-stage synthesis directly produces realistic, contextaware, and detailed abnormal samples <ref:2502.16412#pg8>. Multi-stage synthesis is an advanced VLM technique that provides a comprehensive pipeline for anomaly synthesis <ref:2502.16412#pg8>.
Future Directions
Future research should focus on enhancing anomaly diversity by leveraging adaptive anomaly synthesis techniques such as uncertainty-aware models <ref:2502.16412#pg9>. Controllable synthesis of anomaly attributes requires focusing on cross-class consistency modeling to handle variations in shape, texture, and distribution <ref:2502.16412#pg9>. Promoting multimodal anomaly synthesis involves developing cross-modal alignment strategies that leverage VLM and multimodal transformers <ref:2502.16412#pg9>.
The survey concludes by emphasizing that tackling these challenges will significantly improve the effectiveness of IAS >ref:2502.16412/pg8>. The paper also highlights the limitations of existing methods, such as limited sampling and difficulty in synthesizing realistic anomalies <ref:2502.16412#pg4>. It provides a roadmap for boosting IAS with multimodal learning >ref:2502.16412/pg3>. The work establishes a novel framework for systematic analysis in IAS <ref:2502.16412#pg4>. The paper is the first dedicated work that systematically categorizes IAS methods <ref:2502.16412#pg4>. It offers a comprehensive overview that captures the full scope of techniques available in the field <ref:2502.16412#pg4>. The survey provides a unified and systematic review of anomaly synthesis, covering nearly 40 representative methods across different paradigms <ref:2502.16412#pg4>. It delves into the integration of multimodal cues and large-scale VLM in anomaly synthesis <ref:2502.16412#pg4>. The paper is the first dedicated work that systematically categorizes IAS methods <ref:2502.16412#pg4>. It offers a comprehensive overview that captures the full scope of techniques available in the field >ref:2502.16412/pg3>.
The survey comprehensively reviews anomaly synthesis methodologies, introducing the first industrial anomaly synthesis (IAS) taxonomy and exploring cross-modality synthesis and large-scale Vision Language Models (VLM) to boost IAS. The paper is the first dedicated work that systematically categorizes IAS methods <ref:2502.16412#pg4>. It offers a comprehensive overview that captures the full scope of techniques available in the field >ref:2502.16412/pg3>. The survey provides a unified and systematic review of anomaly synthesis, covering about 40 representative methods across four main paradigms: Handcrafted, Distribution-hypothesis-based, Generative models (GM)-based, and Vision language models (VLM)-based synthesis <ref:2502.16412#pg4>. It delves into the integration of multimodal cues and large-scale VLM in anomaly synthesis <ref:2502.16412#pg4>. The survey provides a unified and systematic review of anomaly synthesis, covering about 40 representative methods across different paradigms <ref:2502.16412#pg4>. It offers a comprehensive overview that captures the full scope of techniques available in the field >ref:2502.16412/pg3>. The survey provides a unified and systematic review of anomaly synthesis, covering about 40 representative methods across different paradigms <ref:2502.16412#pg4>. It delves into the integration of multimodal cues and large-scale VLM in anomaly synthesis <ref:2502.16412#pg4>. The survey provides a unified and systematic review of anomaly synthesis, covering about 40 representative methods across different paradigms <ref:2502.16412#pg4>. It offers a comprehensive overview that captures the full scope of techniques available in the field >ref:2502.16412/pg3>. The survey provides a unified and systematic review of anomaly synthesis, covering about 40 representative methods across different paradigms <ref:2502.16412#pg4>. It delves into the integration of multimodal cues and large-scale VLM in anomaly synthesis <ref:2502.16412#pg4>. The survey provides a unified and systematic review of anomaly synthesis, covering about 40 representative methods across different paradigms <ref:2502.16412#pg4>. It offers a comprehensive overview that captures the full scope of techniques available in the field >ref:2502.16412/pg3>. The survey provides a unified and systematic review of anomaly synthesis, covering about 40 representative methods across different paradigms <ref:2502.16412#pg4>. It delves into the integration of multimodal cues and large-scale VLM in anomaly synthesis <ref:2502.16412#pg4>. The survey provides a unified and systematic review of anomaly synthesis, covering about 40 representative methods across different paradigms <ref:2502.16412#pg4>. It offers a comprehensive overview that captures the full scope of techniques available in the field >ref:2502.16412/pg3>. The survey provides a unified and systematic review of anomaly synthesis, covering about 40 representative methods across different paradigms <ref:2502.16412#pg4>.
Improvements for AI systems
-
How to create a unified framework for anomaly synthesis by implementing a
IAS taxonomy
that reflects methodological progress and practical implications, grounding future research as described in Section 2 of the paper. This taxonomy allows researchers to compare different synthesis approaches systematically across Handcrafted, Distribution hypothesis-based, GM-based, and VLM-based paradigms. -
Develop an AI system capable of generating realistic industrial anomalies by integrating multimodal cues using Vision Language Models (VLM). This system can utilize
single-stage synthesis
ormulti-stage synthesis
approaches to producecontextaware, highly detailed abnormal samples with minimal or no adjustments,
potentially guided by textual prompts as suggested in Section 6. -
Implement a robust anomaly detection pipeline that leverages the insights from the taxonomy to select appropriate synthesis methods based on the required anomaly complexity. This system can dynamically choose between techniques like
Full-image synthesis,
Local anomalies synthesis,
orData-driven synthesis
depending on whether precise structural detail or broad distribution coverage is needed. -
Create an adaptive anomaly generation system that enhances diversity by employing uncertainty-aware models to guide synthesis, as proposed in Section 7. This system can dynamically adjust parameters based on identifying underrepresented abnormal regions and applying a
coarse-to-fine approach, first generating structures then refining details.
-
Design a controllable synthesis module that allows for precise manipulation of anomaly attributes like shape and texture, addressing the limitation that current methods
struggle to handle these variations.
This involves combininglocal anomaly synthesis with advanced segmentation models
to facilitate fine-tuned modifications of anomaly attributes. -
Build a cross-modal alignment module that integrates complementary data sources such as infrared imagery and X-ray scans into the IAS pipeline. This module should use VLM and multimodal transformers to establish
semantic correlations between different modalities,
enabling richer anomaly synthesis by leveragingmulti-source data fusion techniques.
Sources
- A Survey on Visual Anomaly Detection: Challenge, Approach, and Prospect
- SeaS: Few-shot Industrial Anomaly Image Generation with Separation and Sharing Fine-tuning
- AnomalyControl: Learning Cross-modal Semantic Features for Controllable Anomaly Synthesis
- AnomalyXFusion: Multi-modal Anomaly Synthesis with Diffusion
- Dual-Interrelated Diffusion Model for Few-Shot Anomaly Image Generation
- Manifolds for Unsupervised Visual Anomaly Detection
- Unseen Visual Anomaly Generation
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models