Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

arXiv:2608.11587 · eess.AS, cs.CL, cs.LG · Submitted 2026-08-12 · Read on arXiv

Xulin Fan, Jialu Li, Mohammad Nur Hossain Khan, Kexin Hu, Bashima Islam, Mark Hasegawa-Johnson, Nancy L. McElwain

University of Illinois Urbana-Champaign · University of Arizona · Worcester Polytechnic Institute

eess.AS, cs.CL, cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

Comments: Accepted to Interspeech 2026

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: The paper presents a multi-tier, frame-level audio tagging framework for infant-centered audio understanding in naturalistic home recordings.

Terminology

Summary

The paper presents a multi-tier, frame-level audio tagging framework for infant-centered audio understanding in naturalistic home recordings. The task involves jointly performing speaker-aware diarization and vocalization classification across four tiers: child (CHN), female caregiver (FAN), male caregiver (MAN), and sibling (CXN). Each tier has its own mutually exclusive label set (e.g., CHN: INACTIVE, BAB, CRY, FUS; FAN: INACTIVE, ADS, CDS, SNG, LAU), while different tiers may be active simultaneously.

The proposed architecture combines a LoRA-finetuned Whisper-large-v2 encoder, an MLP projector that downsamples the encoder output by grouping consecutive frames into windows of size 5, a target speaker extractor, and per-tier framewise classifiers. The target speaker extractor conditions a lightweight two-layer Transformer encoder on a family-aware speaker token. This token is factorized into a shared tier token (sτ) and a learned family-specific offset (oτ,f), designed to separate tier-level information from family-dependent variability. At inference on unseen families, offsets are set to zero.

The training objective combines standard framewise cross-entropy loss with a sequence-level temporal smoothing loss that penalizes rapid posterior changes across adjacent frames, weighted by λ = 0.2. Experiments were conducted on approximately 17 hours of labeled audio from 52 families using the LittleBeatsTM wearable device, partitioned into training (37 families), validation (5 families), and test (10 families) sets with no family overlap.

Results show the proposed method achieves the best overall performance with 74.88 Macro-F1 and 68.14 κ averaged across tiers, outperforming baselines including TL-TR512 (adapted from Whisper-AT) and W2V-LB (based on wav2vec 2.0 pretrained on family audio). The largest performance gains over W2V-LB occur in adult tiers (FAN and MAN), while the child tier difference is smaller, attributed to Whisper's pretraining on adult-dominated speech corpora versus W2V-LB's pretraining on family recordings containing child vocalizations.

Ablation studies confirm the contribution of each component: removing LoRA fine-tuning degrades performance consistently; removing family-specific offsets primarily affects tiers with higher cross-family variability (especially MAN); removing temporal smoothing reduces temporal stability with larger drops in CHN and MAN tiers. Preliminary test-time adaptation experiments using a SUTA-style objective (entropy minimization and minimum class confusion) to update offsets for unseen families showed only modest improvements, indicating that more effective unsupervised adaptation approaches remain an important direction for future work.

Improvements for AI systems

Improvements to AI systems:

  1. Family-Aware Speaker Conditioning via Factorized Tokens: Implement the target speaker extractor with a shared tier token plus a family-specific offset. This allows an AI system to separate stable, universal speaker-type characteristics (e.g., female adult vs. child) from family-dependent acoustic variability (e.g., unique voice timbre, room acoustics). The system can then generalize to unseen families by zeroing the offset, enabling robust zero-shot speaker diarization and classification without retraining.

  2. Multi-Tier Joint Diarization and Vocalization Classification: Adopt the four-tier, mutually exclusive label sets with simultaneous active tiers. This enables an AI system to output multiple concurrent labels (e.g., a child crying while a female caregiver sings) rather than forcing a single-speaker assumption. The system can thus parse complex, overlapping vocal interactions in naturalistic home environments, improving accuracy in real-world audio understanding.

  3. LoRA-Finetuned Whisper Encoder with Frame-Level Grouping: Use a LoRA-finetuned Whisper-large-v2 encoder followed by an MLP projector that groups consecutive frames into windows of size 5. This reduces computational cost while preserving temporal resolution. The improved system can process long-duration audio (e.g., 17-hour recordings) efficiently and still capture fine-grained vocalization events (e.g., babbles, cries, laughter) with high frame-level precision.

  4. Sequence-Level Temporal Smoothing Loss: Integrate a temporal smoothing loss (λ=0.2) that penalizes rapid posterior changes across adjacent frames. This makes the AI system produce temporally consistent predictions, reducing flickering misclassifications (e.g., alternating between cry and fuss within milliseconds). The system becomes more reliable for downstream applications like automatic annotation of infant behavior or caregiver interaction analysis.

  5. Test-Time Adaptation with SUTA-Style Objective: Incorporate entropy minimization and minimum class confusion to update family-specific offsets at inference. While current gains are modest, this framework allows the AI system to adapt to a new family's acoustic environment in real time using only unlabeled audio, gradually improving performance without manual annotation—useful for long-term deployment in varied home settings.

What the improved AI system can do:

  • Automatically annotate 24/7 wearable audio recordings from infants and caregivers, distinguishing between child, female adult, male adult, and sibling vocalizations with 74.88 Macro-F1 and 68.14 κ across all tiers.

  • Detect overlapping vocal events (e.g., infant cry + adult directed speech) simultaneously, enabling studies of turn-taking, caregiver responsiveness, and language exposure in naturalistic settings.

  • Operate on unseen families without any fine-tuning, thanks to zeroed offsets, making it deployable in new households immediately.

  • Maintain stable, smooth predictions over time, reducing false alarms in real-time monitoring applications (e.g., alerting caregivers to prolonged crying).

  • Adapt incrementally to a specific family's audio characteristics (e.g., a father with a deep voice or a sibling with a high-pitched laugh) through unsupervised test-time updates, improving long-term accuracy as more unlabeled data is collected.

Abstract

Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, target-speaker-aware Transformer for long-context inference and framewise prediction across tiers. To improve temporal coherence, we incorporate a simple sequence-level smoothing loss, and to enhance robustness across households, we introduce a factorized speaker-token design with a shared tier token and a learned family-specific offset, reducing family bias and promoting generalizable representations. Together, these choices enable efficient and effective infant-centered audio tagging of daylong audio recordings in home environments.

Sources

Related papers