Sample-Conditioned Representation Selection for Audio Few-Shot Learning

arXiv:2609.17076 · cs.AI, cs.SD · Submitted 2026-09-15 · Read on arXiv

cs.AI, cs.SD

Submitted: 2026-09-15

Updated: 2026-09-15

Comments: Submitted to ICASSP27

Code: https://github.com/Cross-Innovation-Lab/SAMPLESELECT

License: http://creativecommons.org/licenses/by/4.0/

The gist: Few-shot audio classifiers may rely on foreground-background co-occurrences and fail when those correlations shift.

Terminology

Abstract

Few-shot audio classifiers may rely on foreground-background co-occurrences and fail when those correlations shift. On SpurAudio, the resulting representation shift is concentrated and class dependent: for ResNet12, the top 10 percent of channels explain 82.80 percent of the null-corrected shift contribution. We propose SAMPLESELECT, which predicts a fixed-budget feature mask independently for each input while keeping the encoder and source classifier frozen. Training uses differentiable Gumbel Top-k selection with foreground classification and cross-background contrastive losses; inference uses deterministic Top-k masks and support-only linear adaptation. Across ResNet12 and Conv64 in 5-way 1-shot and 5-shot evaluation, SAMPLESELECT gives the best OOD accuracy among the compared methods and improves the matched full-representation control by 4.90-8.38 percentage points. Ablations and representation analyses further support the learned selection mechanism. Code is available at https://github.com/Cross-Innovation-Lab/SAMPLESELECT/

Sources

Related papers