AlbumentationsX: One Augmentation Pipeline for Images and Related Annotations

arXiv:2608.11123 · cs.CV, cs.LG · Submitted 2026-08-11 · Read on arXiv

Vladimir Iglovikov

Albumentations LLC

cs.CV, cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

Comments: 8 pages, 1 figure. Source code: https://github.com/albumentations-team/AlbumentationsX

Code: https://github.com/albumentations-team/AlbumentationsX

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: AlbumentationsX is a data augmentation library that stores the transform list, probabilities, annotation settings, and random seed in one Compose object.

Terminology

Summary

AlbumentationsX is a data augmentation library that stores the transform list, probabilities, annotation settings, and random seed in one Compose object. Each call chooses random values once and applies them to every supported part of the training example, preventing corruption caused by different random changes applied to an image and its annotations. The library keeps each object’s mask, box, and label together and lets projects add their own transforms. It can also save the pipeline definition, show what happened in one call, and run that call again.

The paper explains that a crop must use the same coordinates for the image, mask, boxes, keypoints, stereo views, video frames, or volume, and that code paths choosing these values separately can silently misalign the data. AlbumentationsX performs the augmentation step while the image and annotations are still separate named values, placed after files have been decoded into arrays and before PyTorch groups examples into a batch.

The library provides API objects to express policy decisions: Compose for ordered steps, Transform p for optional steps, OneOf for choosing one child at random, SomeOf for choosing a requested number at random, RandomOrder for choosing a subset and applying it in random order, and Sequential for nesting one ordered block. The Compose seed controls the policy’s random sequence, independent of global random and numpy.random seeds, and concurrent calls use separate random generators.

For instance segmentation, AlbumentationsX provides behavior through instances: a list with one dictionary per object, keeping that object’s mask, box, and class label together. The bbox labels dictionary holds the box fields named by BboxParams.label fields, and instance binding=[masks,bboxes] selects masks and boxes for joint processing. After a geometric transform, if an object’s mask becomes empty, AlbumentationsX removes the object’s mask, box, and class label together, as shown in Figure 1 where the mug’s mask is outside the output frame but its axis-aligned box still overlaps by a narrow strip.

For stereo vision, the setting additional targets= right image:image tells Compose to process right image like the main image, so both receive the same geometry and image-only transforms. A depth map passed as mask receives shared geometry but skips image-only transforms, with mask interpolation=cv2.INTER LINEAR selecting linear interpolation for continuous measurements. For video clips, passing the complete clip as images=frames uses the same crop and flip for every frame.

Custom transforms can be added by deriving from ImageOnlyTransform for pixel or sensor effects, DualTransform for spatial warps, Transform3D for 3D volume operations, or CustomTransformsApplyMixin for project-specific input names. The sample parameters method chooses random values once per call using sampling.py random for few random numbers or sampling.random generator for random arrays, while apply only uses the values it receives.

To repeat or inspect a random call, a policy seed repeats the whole sequence when samples arrive in the same order, while invocation seed repeats one sample regardless of call order. Setting save applied params=True records which transforms ran and their settings, and ReplayCompose runs the exact random call again. Table 3 lists the minimum record for each question: policy definition, seed, and sample order for repeating the ordered sequence; policy definition and invocation seed for repeating one sample; saved Compose definition and library version for rebuilding the policy; save applied params=True output and sample identifier for seeing which transforms ran; ReplayCompose record for running the exact random call again; and sample order, framework seed, worker count, and worker persistence for repeating the complete loader run.

In PyTorch, AlbumentationsX is called inside Dataset. getitem after file bytes have become arrays, returning a fixed-size image and mask so PyTorch’s default batching can stack them. Training and evaluation use separate policies, with validation typically containing deterministic resizing, padding, normalization, and tensor conversion.

The paper compares AlbumentationsX to TorchVision v2, Kornia, and NVIDIA DALI, noting that in the cited releases, the public APIs expose 121 concrete 2D transform classes in AlbumentationsX, 61 in Kornia, and 38 in TorchVision v2, excluding base classes, pipeline containers, format-conversion helpers, and 3D or spectrogram transforms. AlbumentationsX also keeps geometric changes aligned across images and related annotations.

Limitations include that AlbumentationsX cannot infer whether a transform preserves the task label, and target support varies by transform—for example, a color transform may accept an RGB image but reject a 9-channel array. A depth map can share the image’s crop and resize, but a camera calibration matrix requires a separate rule to update its numbers. The examples use AlbumentationsX release 2.4.0 (commit e85171bc777f), available under AGPL-3.0-only, with separately negotiated commercial licenses also offered.

Improvements for AI systems

Improvements to AI Systems:

  1. Unified Multi-Modal Data Alignment
  • AI systems can now process image, mask, bounding boxes, keypoints, stereo views, video frames, and 3D volumes with a single random seed per sample, ensuring geometric transforms (crop, flip, rotation) are applied identically across all modalities.

  • This eliminates silent data corruption from independent random draws, improving model training stability for tasks like instance segmentation, object detection, and multi-view 3D reconstruction.

  1. Reproducible and Debuggable Augmentation Pipelines
  • Systems can log the exact transform sequence, parameters, and seed for every training sample (via save applied params and ReplayCompose).

  • This enables exact replication of any augmented sample for debugging, error analysis, or auditing model failures, reducing time spent on irreproducible training runs.

  1. Task-Aware Annotation Consistency
  • For instance segmentation, the system keeps each object’s mask, box, and label bound together. If a geometric transform empties a mask, the corresponding box and label are removed jointly, preventing mismatched training targets.

  • This improves model precision by avoiding false positives from partially visible objects.

  1. Flexible Policy Composition for Complex Data
  • Systems can express augmentation strategies using OneOf, SomeOf, RandomOrder, and Sequential containers, enabling fine-grained control over optional transforms (e.g., apply exactly 2 of 5 augmentations).

  • This allows AI systems to simulate diverse real-world conditions (e.g., varying lighting, partial occlusions) more effectively, improving generalization.

  1. Seamless Integration with Deep Learning Frameworks
  • The library operates inside Dataset. getitem after file decoding and before batching, returning fixed-size arrays compatible with PyTorch’s default collate.

  • This reduces engineering overhead and ensures consistent augmentation across distributed training workers, improving scalability.

  1. Custom Transform Extensibility
  • Projects can add domain-specific transforms (e.g., sensor noise, 3D warps) by deriving from ImageOnlyTransform, DualTransform, or Transform3D, with a standardized sample parameters and apply interface.

  • This allows AI systems to adapt to niche data types (e.g., medical volumes, satellite imagery) without modifying core pipeline logic.

  1. Deterministic Multi-Threading and Concurrency
  • Each Compose object uses an independent random generator, so concurrent data loaders (e.g., multiple workers) produce reproducible results without global seed conflicts.

  • This improves training reproducibility across different hardware configurations and batch sizes.

  1. Stereo and Video Consistency
  • For stereo vision, additional targets= "right image":"image" ensures both views receive identical geometric and photometric transforms. For video, passing all frames as images=frames applies the same crop/flip to every frame.

  • This enables AI systems to learn spatio-temporal consistency in tasks like depth estimation, optical flow, and action recognition.

  1. Interpolation-Aware Depth and Mask Handling
  • Depth maps can be treated as masks with mask interpolation=cv2.INTER LINEAR to preserve continuous measurements during geometric transforms, avoiding artifacts from nearest-neighbor interpolation.

  • This improves accuracy in monocular depth estimation and 3D reconstruction pipelines.

  1. Policy Versioning and Replay for Auditing
  • Systems can save the full pipeline definition (including library version) and replay any exact augmentation call, enabling compliance and reproducibility in regulated AI applications (e.g., medical imaging, autonomous driving).

  • This supports robust model governance and simplifies error attribution.

What the Improved AI System Can Do:

  • Train more accurate models on multi-modal data with zero annotation misalignment.

  • Debug and reproduce any augmented sample exactly, reducing iteration time.

  • Handle complex data types (stereo, video, 3D volumes) with minimal custom code.

  • Scale to large datasets with deterministic, thread-safe augmentation.

  • Extend to novel sensor or annotation types without rewriting core logic.

  • Provide auditable, versioned training pipelines for compliance-sensitive domains.

Abstract

Augmentation can corrupt a training example when an image and its annotations receive different random changes. A crop must use the same coordinates for the image, mask, boxes, keypoints, stereo views, video frames, or volume. Code paths that choose these values separately can silently misalign the data. AlbumentationsX keeps the transform list, probabilities, annotation settings, and random seed in one Compose object. Each call chooses random values once and applies them to every supported part of the training example. The library keeps each object's mask, box, and label together and lets projects add their own transforms. It can also save the pipeline definition, show what happened in one call, and run that call again. The examples place Compose after files have been decoded into arrays and before PyTorch groups examples into a batch. AlbumentationsX executes the declared transforms. Practitioners still decide whether a flip, crop, color change, or other operation preserves the correct label for their task.

Related papers