Multimodal Taxonomic Conditioning for Generative Plankton Imagery

arXiv:2609.11673 · cs.CV, cs.LG · Submitted 2026-09-10 · Read on arXiv

cs.CV, cs.LG

Submitted: 2026-09-10

Updated: 2026-09-10

Comments: European Conference on Computer Vision (ECCV) 2nd Workshop on Marine Vision

License: http://creativecommons.org/licenses/by/4.0/

The gist: Automated plankton imaging produces severely long-tailed datasets, where the rare taxa of greatest ecological interest have too few images to train or evaluate classifiers reliably.

Terminology

Abstract

Automated plankton imaging produces severely long-tailed datasets, where the rare taxa of greatest ecological interest have too few images to train or evaluate classifiers reliably. We generate synthetic plankton imagery conditioned on taxonomy: a CLIP encoder is adapted on a large plankton corpus with a ranked contrastive objective extended to deep, ragged taxonomies, then frozen to condition a parameter-efficient diffusion transformer. We evaluate synthetic sample quality on distributional fidelity and downstream classifier utility.

Related papers