ScratchSim: A Procedural Synthetic Data Pipeline for Surface Scratch Detection
Paul Julius Kühn, Saptarshi Neil Sinha, Tiago Kleist, Richard Hoffmann, Arjan kuijper, Michael Weinmann
Fraunhofer Institute for Computer Graphics Research IGD · Technical University of Darmstadt · Delft University of Technology
cs.CV, cs.AI
Submitted: 2026-08-15
Updated: 2026-08-18
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 76/100
The gist: This paper presents a procedural rendering pipeline for generating large-scale annotated synthetic training data for surface scratch detection, addressing the challenge of scarce annotated defect
Terminology
Summary
This paper presents a procedural rendering pipeline for generating large-scale annotated synthetic training data for surface scratch detection, addressing the challenge of scarce annotated defect data in industrial quality control. The pipeline is built on BlenderProc, an open-source procedural rendering framework, and supports configurable material appearance, camera modes, and domain randomization, producing automatic COCO-format annotations.
The pipeline generates scratches procedurally on a black canvas of 8192 × 8192 pixels, with each mask containing N ∈ [17, 25] scratches drawn uniformly at random. Each scratch is defined by a bounding box of random size and placed at a random position, with four points sampled at random used as start point, end point, and two intermediate control points of a cubic Bézier curve to approximate irregular, curved geometry of real surface scratches. A Gaussian blur is applied to soften edges and suppress aliasing artifacts, followed by downscaling by a factor of two and normalization.
The material construction comprises three layers: a surface appearance model, procedurally generated scratches, and their interaction.
For matte powder-coated surfaces, two noise textures varying in scale and level of detail are mixed to capture both frequency components, generating a height map from which a normal map is derived. For glossy automotive surfaces, the material parameters are adjusted towards higher specularity, lower surface roughness, and increased reflectance
to approximate the multi-layer paint system.
A viewing-angle restriction addresses the issue of scratches becoming invisible at shallow viewing angles near edges of round objects. This is implemented by computing the dot product of the unit surface normal and the unit camera viewing direction, such that ⃗v · ⃗n > t, where t ∈ [0, 1] is a user-defined threshold. The visibility threshold was set to 0.4 for the industrial grip and 0.55 for the toy car.
Camera positioning varies for each frame, with two camera modes implemented: a close-range mode (randomcam) placing the camera on a spherical shell with radius r ∈ [2, 5] cm, altitude β ∈ [30, 90], and azimuth constrained within ±20 of the surface normal, and a tripod mode with camera height sampled in [12, 16] cm and horizontal position sampled on a disk of radius 7.5 cm. Each candidate pose undergoes validity checks including camera lying within room boundaries, no obstacle within 1 cm of the lens, camera not inside the object mesh, and excluded surface regions not visible.
The environment consists of an empty room constructed from five planes sharing a single PBR material chosen at random from a texture library, with two fixed overhead light sources positioned on opposite ends of the room, each controllable independently in colour and intensity.
Annotations are generated automatically using an Arbitrary Output Variable (AOV) node that exports the applied scratch mask as rendered from the camera's perspective, which is then binarised, labelled via connected components, and passed to a COCO-format annotation writer that computes bounding boxes for each labelled region.
The pipeline follows a nested loop controlled by two parameters, the number of scenes S and frames per scene F, producing S × F images in total. For each scene, the scratch mask, material properties, object pose, background, and lighting are randomised, while within a scene only the camera pose changes. All images are rendered at 640 × 640 pixels using Blender's Cycles engine with PBR shading.
Datasets were created for two objects: a matte industrial grip and a glossy Ferrari toy car. All synthetic datasets contain 10,000 images each, split into 7,000 training, 2,000 validation, and 1,000 test images following a 70/20/10 ratio. Two synthetic datasets were generated for the Ferrari toy car using the tripod camera mode and glossy material configuration, featuring either a randomized or white static background. The real Ferrari dataset contains 116 images captured in front of static background, split into 82 training, 18 validation, and 16 test images. Four synthetic base datasets were created for the industrial grip by combining two color modes (tricolour and randomcolour) with two camera modes (randomcam and tripod). Two real grip datasets were captured: 120 images for training with five different backgrounds, and 120 images for evaluation against a static background.
Three lightweight edge-deployable detectors were evaluated: YOLOX-S for the industrial grip, and YOLO26-n and LW-DETR-Tiny for the toy Ferrari datasets. Four training strategies were compared: synthetic-only, real-only, mixed synthetic and real, and fine-tuning from synthetic weights.
For the toy Ferrari datasets, training on the full real dataset established the baseline with YOLO26 achieving mAP50 = 0.610 and mAP50-95 = 0.402, and LW-DETR reaching mAP50 = 0.535 and mAP50-95 = 0.366. Synthetic-only training yielded substantially lower scores for both models (YOLO26: 0.451 / 0.300; LW-DETR: 0.382 / 0.243), reflecting the domain gap. Under scarce real-data conditions (10% to 25%), both models struggled, with YOLO26 falling below mAP50 = 0.06 and LW-DETR below 0.36. Infusing synthetic images alongside real data provided clear remedy: at 10% real data YOLO26 recovered to mAP50 = 0.507 (from 0.056 real-only) and LW-DETR to 0.387 (from 0.238). At full infusion (100%), both models surpassed their real-only baselines (YOLO26: 0.691 / 0.463; LW-DETR: 0.663 / 0.458), demonstrating that synthetic data acts as a useful regulariser even when all real data is available. The best results were achieved by the two-stage fine-tuning strategy, where Finetune Synth (WB) → Real attained the best YOLO26 scores across all metrics except precision (mAP50 = 0.723, mAP50-95 = 0.507, P = 0.764, R = 0.652), while for LW-DETR, Finetune Synth → Real yielded the highest mAP50-95 (0.480) and the WB variant led on mAP50 (0.691).
The ablation study found that matching the background alone without a real-domain anchor hurt performance, but once real images were present, white background added +0.050 mAP50 for YOLO26 at 25% infusion, with further gains at 50% and 100% (+0.038 and +0.015), while LW-DETR remained largely unaffected. Increasing the proportion of real images in the infused regime consistently improved performance for both models, with the 50% infused configuration already surpassing the 100% real-only baseline for both YOLO26 (0.635 vs. 0.610) and LW-DETR (0.607 vs. 0.535).
For the industrial grip datasets, training on synthetic data alone fell below the real baseline in all configurations, with tricolour randomcam coming closest with AP 0.196, while randomcolour tripod collapsed to AP 0.011. Mixed training recovered performance across all variants and surpassed the real baseline clearly when randomcam variants were used, with tricolour randomcam reaching AP 0.316 and AR 0.396. Fine-tuning proved the most consistent strategy, with all four variants exceeding the real baseline, and randomcolour randomcam achieving the best overall result of AP 0.319 and AR 0.389.
The ablation examining incremental addition of real images showed consistent trends across all variants, with performance increasing with more real data though marginal gains diminished in later stages. The authors hypothesise that dataset balance explains differences in improvement patterns, since tricolour closely resembles real images while randomcolour benefits greatly from early real examples but begins to stagnate once real images exceed roughly 48.
The paper concludes that "synthetic-only training falls short due to the domain gap, yet synthetic data remains highly beneficial when combined with real images. Mixed training effectively recovers performance under scarce real-data conditions, while fine-tuning from synthetic weights proved the most robust strategy, outperforming real-only training across all architectures and datasets." Future work could explore higher-fidelity material simulation, advanced domain adaptation techniques, active learning, and extending the pipeline to other defect types and object geometries, as well as simulating defects under challenging real-world conditions such as dust, mud, and contaminated or weathered surfaces.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to AI systems for surface scratch detection:
-
Improvement: Implement a two-stage training protocol where the detector is first pretrained on procedurally generated synthetic scratch data, then fine-tuned on real annotated images.
-
Result: This consistently outperforms real-only training. For YOLO26, mAP50 improves from 0.610 (real-only) to 0.723 (fine-tuned). For LW-DETR, mAP50-95 improves from 0.366 to 0.480.
-
Improvement: Build a BlenderProc-based pipeline that procedurally generates scratches using cubic Bézier curves with randomized count (17–25 per image), depth, and curvature, applied via UV-mapped normal maps.
-
Result: The system can generate 10,000 fully annotated images per dataset with COCO-format bounding boxes automatically, eliminating manual annotation costs.
-
Improvement: Implement a visibility threshold (dot product of surface normal and camera vector > 0.4–0.55) that suppresses scratch annotations at shallow viewing angles where scratches are physically invisible.
-
Result: Reduces false positives in training data by preventing annotations for scratches that cannot be seen in the rendered image.
-
Improvement: Use two distinct material pipelines: (a) multi-layer noise-based normal mapping for matte/powder-coated surfaces, and (b) specularity-adjusted PBR materials for glossy/automotive finishes.
-
Result: Enables the same detection architecture to generalize across objects with fundamentally different surface reflectance properties.
-
Improvement: When real data is scarce (10–25% of full dataset), mix synthetic images with real ones during training rather than relying on augmentation alone.
-
Result: At 10% real data, YOLO26 mAP50 recovers from 0.056 (real-only) to 0.507 (mixed). At 50% real data, mixed training surpasses the 100% real-only baseline.
-
Improvement: When real images are captured against a static background, render synthetic images with a matching background color to reduce domain shift.
-
Result: For YOLO26, white-background synthetic data adds +0.050 mAP50 at 25% real infusion and +0.043 mAP50 in fine-tuning, though this effect is negligible for transformer-based detectors.
-
Improvement: Implement two camera modes—close-range (radius 2–5 cm, azimuth ±20° of surface normal) and tripod (fixed height, disk sampling)—with validity checks to avoid occlusions and out-of-bounds poses.
-
Result: Random camera viewpoints (randomcam) significantly outperform fixed tripod setups for matte objects (AP 0.196 vs. 0.110), while tripod mode is sufficient for glossy objects.
-
Detect surface scratches on glossy and matte objects with mAP50 up to 0.723 (YOLO26) and mAP50-95 up to 0.480 (LW-DETR) on real test images, outperforming models trained exclusively on real data.
-
Operate with as little as 10% of real annotated data while maintaining usable detection performance (mAP50 > 0.5), making deployment feasible in manufacturing environments where defect samples are rare.
-
Generate unlimited synthetic training data on demand with automatic pixel-level annotations, supporting rapid iteration on new object geometries without manual labeling effort.
-
Deploy on edge devices using lightweight detectors (YOLOX-S, YOLO26-Nano, LW-DETR-Tiny) that achieve real-time inference while maintaining detection accuracy.
-
Adapt to new materials and lighting conditions through domain randomization of background textures, light intensity, and surface properties, reducing the need for object-specific retraining.
-
Avoid false-positive annotations from shallow viewing angles, improving training signal quality and reducing model confusion.
-
Scale to industrial inspection scenarios where manual annotation is impractical, enabling automated quality control for large-volume production lines.
Abstract
While automated defect detection such as the detection of surface scratched is an important aspect in industrial quality control, the scarcity of annotated defect data make this task challenging. This paper presents a procedural rendering pipeline that generates large-scale annotated synthetic training data using BlenderProc, with configurable material appearance, camera modes, and domain randomization, producing automatic COCO-format annotations. To show the potential of our approach, we evaluate four training strategies, namely synthetic-only, real-only, mixed, and fine-tuning from synthetic weights, across two objects with different material properties and three lightweight edge-deployable detectors, YOLOX, YOLO26, and LW-DETR. Our evaluation show that fine-tuning from synthetic weights consistently outperforms real-only training, and that mixed training effectively recovers performance under scarce real-data conditions, with findings validated across both convolutional and transformer-based architectures. The proposed approach enables scalable defect detection without the burden of large real annotated datasets, making it practical for on-device industrial inspection. The pipeline scripts, 3D model, and both synthetic and real annotated scratch datasets for a glossy toy Ferrari car will be made available through the project website upon acceptance.
Sources
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models