Rhythm of the Deep: Two-tier acoustic organization of sperm-whale codas from click waveforms to second-order sequence dependence

arXiv:2606.16084 · cs.AI, cs.CL · Submitted 2026-06-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Rhythm of the Deep: Two-tier acoustic organization of sperm-whale codas from click waveforms to second-order sequence dependence".

Jane: The paper was written by Mudit Sinha and Sanika Chavan from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: We are starting our show with a paper called "Rhythm of the Deep: Two-tier acoustic organization of sperm-whale codas from click waveforms to second-order sequence dependence."

Jane: This one really stands out because the authors, Mudit Sinha and Sanika Chavan, are independent researchers.

Tom: That is quite a feat to produce something this complex without a massive institutional team behind you.

Jane: It shows how much focus they put into the data, especially since they are looking at a "two-tier" organization.

Lu: I find that concept of two tiers so exciting because it suggests a deep, hidden architecture in the ocean.

Jane: To put it simply, it means the whale sounds are built like a ladder, where small building blocks combine to make larger structures.

Lu: We have likely missed this for a long time because we were only looking at the surface rhythms.

Meng: I wonder if they had to develop specialized software to find those specific layers.

Jane: They actually used "frozen audio encoders," which are pre-trained AI models that help identify patterns without needing human labels.

Lu: That is the beauty of it, because the structure was already there in the waveforms.

Meng: If they have found these layers, does that mean we can eventually decode what the whales are actually communicating?

Lalam: The existence of a hierarchy suggests that whale communication has a complexity that mirrors our own linguistic structures.

Tom: That is a massive implication, so let's look at what those two tiers actually look like in the data.

Summary: Tom: We have established that there are layers, so now we need to examine what Sinha and Chavan actually found in the codas.

Jane: For a long time, people thought codas were just patterns of clicks and the timing between them.

Tom: We used to think we were just hearing Morse code, but the researchers found that every single dot and dash actually carries its own unique sound.

Jane: Exactly, and the first tier is where these specific click identities combine with the rhythm to create a "coda unit."

Lu: Then the second tier is where the real magic happens.

Jane: These coda units aren't just random, because they show a sequence-level dependence.

Meng: I noticed they reported a "lag-two dependence" of about zero point one three two bits for those second-tier units.

Tom: That sounds like a very precise way of saying that what happened two sounds ago influences what is happening now.

Jane: Think about how in English, hearing a specific word can give you a strong hint about what is coming next.

Lu: This proves the whales are operating within a structured acoustic space with its own internal logic.

Meng: I am curious if the structure holds up when you change the speed of the clicks.

Lalam: The paper shows that while individual clicks might change with tempo, the identity of the coda remains much more stable.

Tom: That suggests a level of structural integrity that goes way beyond simple timing.

Improvements: Tom: We have seen the findings, but now we have to discuss how they proved this wasn't just a fluke or an AI error.

Jane: That is a huge question because AI can sometimes find patterns that aren't actually there.

Tom: To prevent that, they used a "controlled-induction framework" to audit their own results.

Jane: They didn't rely on just one model, but used eight different families of audio encoders to see if they all agreed.

Meng: I am interested in how they used "destructive waveform counterfactuals" to see if the patterns disappeared when they intentionally messed with the audio.

Lu: That is such a clever way to stress-test the data.

Meng: They also used "matched nulls" to ensure the patterns weren't just caused by background noise or recording conditions.

Jane: They even proved the system works even if you don't know the exact timing of the clicks.

Lu: This framework could be used to verify truth in any kind of complex, messy signal.

Meng: If we can apply this to other species, it would change how we study animal communication entirely.

Lalam: By demanding this level of proof, they are setting a new standard for how we use technology to understand the natural world.

Tom: They have built a way to separate real biological structure from mere noise.

Conclusion: Tom: We are wrapping up our discussion of "Rhythm of the Deep: Two-tier acoustic organization of sperm-whale codas from click waveforms to second-order sequence dependence."

Jane: This has been an incredible look at the complexity hidden in the ocean.

Tom: It really shifts our understanding of how these whales organize their sounds.

Lu: I am left thinking about all the other hidden structures waiting to be found with these tools.

Meng: I am looking forward to seeing how this validation framework improves our bioacoustic models.

Lalam: This research shows that the natural world is far more organized and communicative than we ever imagined.

Jane: It has been a pleasure discussing this with everyone.

Tom: We will catch you next time for another paper from arXiv.

cs.AI, cs.CL

Submitted: 2026-06-15

Updated: 2026-09-15

Comments: 12 pages, 6 figures, with 12 pages of supplementary material. Preprint

License: http://creativecommons.org/licenses/by-sa/4.0/

Importance score: 86/100

The gist: This paper investigates the acoustic structure of sperm-whale codas, challenging the traditional view that these signals are merely recurring click-count and timing patterns.

Key concepts

Two-tier acoustic organization
Whale sounds are structured like a ladder rather than just simple rhythms. The first tier combines specific click identities into "coda units," while the second tier involves these units following a sequence where previous sounds influence what comes next, suggesting a complex communication structure similar to human language.
Frozen audio encoders
These are pre-trained AI models that help identify patterns in sound waves without requiring human labels. They allow researchers to find the inherent structures already present within the acoustic waveforms of whale communication.
Controlled-induction framework
This is a validation method used to prove that AI-detected patterns are real biological structures rather than errors or noise. It involves using multiple different AI models, testing how patterns react to intentional audio changes, and ensuring results aren't caused by background noise.

Terminology

Summary

This paper investigates the acoustic structure of sperm-whale codas, challenging the traditional view that these signals are merely recurring click-count and timing patterns. By revealing a two-tier combinatorial acoustic organization, the study provides a new framework for understanding how complex biological signals are assembled from lower-level units and participate in higher-order structures.

The Methodological Framework

The researchers employ a controlled-induction framework designed to turn high-dimensional audio representations into falsifiable claims about units, combination rules, and cross-tier relations. Rather than relying on expert labels, the study uses eight families of frozen audio encoders to independently propose candidate click and coda units. This approach allows the researchers to separate the acoustic relation carried by an induced unit from shared encoder shortcuts that might otherwise mimic biological recurrence.

To ensure scientific rigor, the framework subjects every interpretation to a series of stringent validation gates:

  • Cross-view consensus and stability checks to retain recurring candidates across model families.

  • Held-out transfer tests to determine if lower-tier composition predicts coda identity.

  • Matched nulls and destructive waveform counterfactuals to distinguish between identity, order, rhythm, and nuisance carriers.

  • Expert-feature baselines to calibrate findings against established acoustic accounts.

The First Acoustic Tier

The study identifies that the lower tier of coda organization is composed of a recurring click inventory combined with interclick rhythm. The researchers explicitly reject the hypothesis that codas are supported by a stable ordered rule or a fixed sequence of click tokens, noting that stable click order is weak. Instead, once the exact click-token multiset is fixed, the inter-click intervals (ICIs) become the primary predictor of coda identity.

The connection between these lower units and the higher tier is demonstrated through several key metrics:

  • Click-token composition predicts induced coda identity with a median normalized mutual-information lift of 0.380.

  • This predictive power remains even when events are detected using an annotation-independent arm without published click times or counts.

  • The relationship is robust across different model families, indicating the inventory is not determined by a single architecture.

The Second Acoustic Tier

At the second tier, the study reveals that recurring coda tokens exhibit additional sequence-level dependence under a different acoustic carrier. Specifically, coda tokens show 0.132 bits of incremental lag-2 dependence under a prespecified categorical estimator. Crucially, this higher-order structure is not captured by traditional rhythm models; two expert-inspired rhythm representations failed to recover the positive component of this association.

The researchers further distinguish the tiers through tempo manipulation, identifying a cross-tier stability gradient. When the audio is resampled at 1.3× speed:

  • Click identity becomes highly unstable, with an Adjusted Rand Index (ARI) of only 0.074.

  • Coda identity remains significantly more robust, with whole-waveform and ICI-only identities showing much higher stability (ARI of 0.428 and 0.516, respectively).

Implications for Acoustic Science

The results suggest that the acoustic description of sperm-whale codas must shift from a single prescribed rhythm inventory to layered waveform organization. The study concludes that codas are not merely timing templates but are assembled from lower acoustic units whose combinations create recurring coda units that themselves participate in higher-order structure. This controlled-induction framework provides a general method for discovering and falsifying combinatorial structure in any under-annotated audio signal.

Improvements for AI systems

1. Implementation of a Controlled Acoustic-Unit Induction Framework

  • Improvement: Integrate a multi-stage validation pipeline for unsupervised unit discovery that moves beyond simple cluster agreement by incorporating held-out transfer, carrier-matched nulls (e.g., spectrum-matched noise, click-order shuffling), and destructive waveform counterfactuals (e.g., replacing tokens within a preserved timing skeleton).

  • Capability: The AI can autonomously distinguish between genuine structural recurrence in under-annotated audio and model shortcuts caused by shared sensitivities to spectral bias, recording conditions, or segmentation artifacts.

2. Two-Tier Hierarchical Combinatorial Tokenization

  • Improvement: Transition from single-tier tokenization (e.g., direct waveform-to-token) to a dual-layer architecture where Tier 1 models the composition of low-level acoustic identities and inter-event rhythm to form mid-level units, which are then processed by a Tier 2 sequence model for higher-order dependence.

  • Capability: The AI can model complex, layered communication systems in environments where the alphabet is unknown, allowing it to capture hierarchical dependencies (such as lag-2 sequence associations) that single-tier N-gram or transformer models miss.

3. Counterfactual Representation Auditing

  • Improvement: Incorporate a training and evaluation loop using destructive acoustic-null gates, where the model's learned representations are tested against manipulated versions of the input (e.g., phase shuffling, time reversal, or envelope-noise filling).

  • Capability: The AI can self-verify whether its latent space is capturing the intended acoustic carrier (e.g., identity vs. rhythm) and provide explicit abstentions when a detected pattern fails to survive rigorous counterfactual testing.

4. Cross-Tier Tempo-Invariant Identity Mapping

  • Improvement: Implement training objectives that maximize representation stability across tempo scaling (resampling/stretching) to decouple acoustic identity from temporal duration.

  • Capability: The AI can achieve robust recognition of semantic or structural units regardless of playback speed, preventing the model from conflating what is being signaled with how fast it is being signaled.

Abstract

Sperm-whale codas are conventionally characterized by click count and inter-click intervals (ICIs), leaving recurring differences in constituent click waveforms unresolved. This study tests whether acoustic organization is nested across two scales: within codas, where recurring click-waveform differences may complement ICI timing, and across codas, where recurring whole-coda forms may themselves carry sequence dependence. Candidate recurring click and whole-coda groupings were identified from 1,483 codas without prespecifying waveform categories, then evaluated with native-rate spectral/envelope measurements, exact nuisance matching, held-out timing contrasts, and sequence controls. At the first tier, recurring click-waveform groups differed in spectral slope, bandwidth, flatness, high/low-band energy, and envelope structure within matched date, social unit, individual, and sample-rate strata. Their composition added information about whole-coda grouping beyond timing, while timing remained informative when click composition was fixed. The richer description also carried held-out social-unit-associated information beyond timing. At the second tier, direct native-rate waveform summaries recovered the recurring whole-coda forms well above context-preserving nulls, whereas conventional timing did not; the forms also cross-cut published timing-defined coda types. The preceding two-coda context then added held-out predictive information beyond the immediately preceding coda, while a third preceding coda provided no reliable further gain. Together, these results support two-tier acoustic organization: recurring waveform differences and ICI timing jointly organize individual codas, and acoustically grounded whole-coda forms in turn show bounded second-order predictive dependence across sequences.

Related papers