BAT: Learning to Reason about Spatial Sounds with Large Language Models

summary

Video file (mp4)

The gist

BAT: Learning to Reason about Spatial Sounds with Large Language Models Abstract: "Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based

This episode discusses

The paper

BAT: Learning to Reason about Spatial Sounds with Large Language Models · Read on arXiv

Zhisheng Zheng, Puyuan Peng, Ziyang Ma, Xie Chen, Eunsol Choi, David Harwath

University of Texas at Austin · Shanghai Jiao Tong University

Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed SpatialSoundQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or Spatial-AST, which by itself achieves strong performance across sound event detection, spatial localization, and distance estimation. By integrating Spatial-AST with LLaMA-2 7B model, BAT transcends standard Sound Event Localization and Detection (SELD) tasks, enabling the model to reason about the relationships between the sounds in its environment. Our experiments demonstrate BAT's superior performance on both spatial sound perception and reasoning, showcasing the immense potential of LLMs in navigating and interpreting complex spatial audio environments.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "BAT: Learning to Reason about Spatial Sounds with Large Language Models".

Jane: The paper was written by Zhisheng Zheng, Puyuan Peng, Ziyang Ma, Xie Chen, Eunsol Choi et al. from University of Texas at Austin and Shanghai Jiao Tong University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s got me genuinely excited — it’s called “BAT: Learning to Reason about Spatial Sounds with Large Language Models.” Jane, you’ve been grinning since we started reading this one.

Jane: I really have, Tom. This paper is about teaching a large language model to actually understand where sounds are coming from in a three dee space. You know, not just “I hear a dog barking,” but “the dog is barking to my left, about three meters away, and there’s a stereo playing behind me at the same time.” That’s a huge leap from what most audio models can do.

Tom: And that’s the kicker, right? Most of the audio LLMs out there — they work on monaural audio, which is just a single channel, like one ear. They can tell you what’s making a sound, but they have no idea where it’s coming from. This paper, BAT, uses binaural audio — two channels, like having two ears — and it actually reasons about space.

Jane: Exactly. And the way they did it is pretty clever. They built a new audio encoder called Spatial-AST, which takes in the two-channel audio and extracts three things at once: what the sound is, which direction it’s coming from, and how far away it is. Then they hook that up to a large language model — LLaMA-two specifically — so the model can answer questions like “Is the dog to the left of the stereo?”

Tom: And it’s not just a toy demo. They created a whole new dataset called Spatial Sound QA, with over eight hundred thousand question-answer pairs. They simulated realistic indoor environments with reverberation, using SoundSpaces two point zero and AudioSet clips. So the model is trained on sounds bouncing off walls, coming from different angles and distances — messy, real-world-ish conditions.

Jane: That’s the part that really impressed me. They didn’t just record a few sounds in a studio. They built a virtual world with ninety different buildings, thousands of rooms, and they placed sound sources at random locations, then simulated how the sound would actually reach a listener’s two ears. That’s a lot of careful engineering.

Tom: And the results? The Spatial-AST encoder alone gets a mean average precision of about fifty percent on sound event detection, and it can estimate direction within about eighteen degrees of error on average. Then when you combine it with the LLM, the full BAT model answers spatial reasoning questions with about seventy-seven percent accuracy. That’s way above random guessing, which would be fifty percent.

Jane: Right, and that seventy-seven percent is on questions that require the model to separate two overlapping sounds, figure out where each one is, and then reason about their relationship. That’s not just perception — that’s actual spatial reasoning. I think this could be a big deal for robotics, for augmented reality, even for hearing aid technology.

Tom: I was just thinking about that. Imagine a robot that can hear you call from across a room and actually turn toward you. Or a hearing aid that can tell you “the person speaking is to your right, three meters away.” This paper lays the groundwork for all of that.

Jane: And it’s the first of its kind, Tom. The authors say it’s the first spatial audio-based LLM. So we’re at the very beginning of something new here.

Tom: I love it. So what’s next? We’ve got the title and the big picture — let’s get into the actual method and how they made this work.

Summary: Tom: So we’ve established that “BAT: Learning to Reason about Spatial Sounds with Large Language Models” is a big deal. But let’s get into the nitty-gritty of how they actually built it. Jane, you’re the one who can explain this in plain English.

Jane: Happy to, Tom. So the core idea is pretty simple: they take a two-channel audio clip — left ear and right ear — and they process it in two stages. First, they have this encoder called Spatial-AST. It takes the audio and turns it into a set of tokens, like a compressed summary of what’s happening. But here’s the clever part: they add three special tokens, each one trained to focus on a different thing — one for the sound class, one for direction, one for distance.

Tom: So it’s like having three little specialists inside the encoder, each one looking for a different piece of information.

Jane: Exactly. And they train it on three tasks at once: “what is this sound?”, “which direction is it coming from?”, and “how far away is it?” They use cross-entropy loss for all three, but they train in two stages. First, they just focus on classification, because that’s the hardest. Then they add the direction and distance losses.

Tom: And that two-stage approach really mattered, right? I saw in the paper that training all three at once from the start hurt performance.

Jane: Yeah, the numbers show it clearly. With binaural audio and IPD — that’s interaural phase difference, the tiny timing difference between when a sound hits your left ear versus your right ear — the two-stage training gets the direction error down to about eighteen degrees. Joint training, where they do everything at once, gets about eighteen point six degrees. It’s a small difference, but it’s consistent across all the metrics.

Tom: And what about the distance estimation? That’s the part that surprised me.

Jane: Me too. They measure distance error rate — basically, how often the model is within half a meter of the true distance. With the full binaural setup, they get about thirty-two point five percent error rate. That sounds low, but remember, the distances range from zero to ten meters, and the model has to pick from discrete half-meter bins. Random guessing would be way worse.

Tom: So the encoder itself is already a strong spatial audio model. But then they go further and hook it up to the LLM. How does that work?

Jane: They take the output tokens from Spatial-AST — including those three specialist tokens — and they project them into the embedding space of LLaMA-two. Then they fine-tune the LLM with a parameter-efficient method called LLaMA-Adapter v2. The LLM stays mostly frozen, but they add a few trainable layers so it can learn to interpret the audio tokens.

Tom: And that’s where the reasoning comes in. The model can answer questions like “Is the sound of the dog to the left of the sound of the stereo?” — and it has to figure that out from the raw audio, not from any text description.

Jane: Right. And they trained it with a curriculum. Stage one is single-source questions — just “what do you hear and where is it?” Stage two introduces two sources at once, which forces the model to implicitly separate them. Stage three adds the reasoning questions. And the ablations show that if you skip stage two, the reasoning ability collapses to near random.

Tom: That’s a really important finding. It means the model isn’t just memorizing patterns — it actually needs to learn how to separate overlapping sounds before it can reason about their spatial relationships.

Jane: Exactly. And that’s the kind of insight that makes this paper more than just a demo. It tells us something about how to train these models effectively.

Tom: So we’ve got the encoder, the training curriculum, the dataset. But what does this actually mean for the world? Let’s bring in Lu and Meng for that.

Improvements: Tom: Alright, we’ve covered the basics of “BAT: Learning to Reason about Spatial Sounds with Large Language Models.” Now let’s talk about what this paper improves on and where it could go. Lu, you’ve been thinking about the big picture — what’s the real step forward here?

Lu: Thanks, Tom. I think the biggest improvement is that this paper moves audio LLMs from “what” to “where.” Before BAT, models like LTU or Pengi could tell you what sounds were present, but they were working with monaural audio — they had no spatial information at all. BAT is the first to bring binaural spatial cues into the LLM pipeline. That’s a fundamental capability upgrade.

Jane: And it’s not just about adding a new input channel. The paper shows that the model can actually reason about relationships between multiple sound sources. That’s something that no previous audio LLM could do.

Lu: Exactly. And that reasoning ability is what opens the door to real-world applications. Think about autonomous robots navigating a home — they need to know not just that there’s a person talking, but where that person is relative to the robot. BAT provides a template for that.

Meng: I’m with you on the potential, Lu, but let me play the engineer for a second. How practical is this to actually deploy? The paper uses LLaMA-two 7B, which is a pretty big model. And they trained on eight V100 GPUs. That’s not exactly edge-device friendly.

Tom: That’s a fair point, Meng. But they do use parameter-efficient fine-tuning — LLaMA-Adapter v2 — which means most of the LLM stays frozen. So the actual trainable parameters are small. And the encoder, Spatial-AST, is only about ninety million parameters. That part could run on a modest GPU.

Meng: Okay, that’s more reasonable. But what about the dataset? They synthesized everything with SoundSpaces two point zero. That’s a simulator. How well does that transfer to real-world audio? There’s always a sim-to-real gap.

Lu: That’s the big open question, and the authors acknowledge it. They mention it in the limitations section. But the fact that they used real AudioSet clips as sound sources, and real room geometries from Matterportthree dee, means the audio is at least grounded in reality. The reverberation is simulated, but it’s based on actual room layouts.

Jane: And they also show that the IPD — the phase difference between the two ears — is crucial for direction estimation. That’s a physical cue that exists in real binaural recordings too, so the model should generalize at least partially.

Meng: I’d still want to see a real-world test before I’d trust it in a product. But the architecture is sound, and the training curriculum is smart. The fact that they show you need the multi-source separation stage to get reasoning to work — that’s a practical lesson for anyone building this kind of model.

Tom: So what’s the next big improvement you’d want to see?

Lu: I’d love to see them extend this to more than two sound sources. Right now the reasoning questions are all binary — two sources. Real environments have many overlapping sounds. Also, they only handle static sources — no moving sounds. Tracking a moving sound source would be a natural next step.

Meng: And I’d like to see a smaller, distilled version of the model. If you could get the reasoning ability into a model that runs on a phone or a smart speaker, that would be huge for consumer applications.

Jane: The authors also mention integrating visual information. SoundSpaces two point zero actually supports visual rendering too, so you could train a model that understands both what it sees and what it hears in a three dee environment. That’s a really exciting direction.

Tom: Alright, so we’ve got the improvements and the future directions. Let’s wrap this up with a conclusion and say goodbye to BAT.

Conclusion: Tom: Alright, we’ve spent a lot of time with “BAT: Learning to Reason about Spatial Sounds with Large Language Models,” and I think it’s fair to say this is one of those papers that marks a turning point. Jane, how would you sum it up for someone who just tuned in?

Jane: I’d say BAT is the first large language model that can actually hear in three dee. It takes binaural audio — two channels, like two ears — and it can tell you what sounds are present, where they’re coming from, how far away they are, and even reason about the spatial relationships between multiple sounds. It does this using a new encoder called Spatial-AST, which is trained on a massive simulated dataset called Spatial Sound QA.

Tom: And the key numbers — the encoder gets about fifty percent mean average precision on sound event detection, direction error around eighteen degrees, and the full BAT model answers spatial reasoning questions with about seventy-seven percent accuracy. That’s a huge step up from random guessing.

Lu: And the implications go beyond just audio. This shows that LLMs can be extended to handle spatial reasoning in a modality that’s inherently three-dimensional. That’s relevant for robotics, augmented reality, assistive technology, and even gaming.

Meng: From a practical standpoint, the training curriculum is the most valuable takeaway. The paper proves that you can’t just throw all the tasks at the model at once — you need to build up from single-source perception to multi-source separation before reasoning works. That’s a lesson that applies to any multimodal LLM.

Lalam: I’d add that this work has cultural implications too. Spatial audio is how humans experience the world — we don’t just hear sounds, we place them in space. By giving LLMs that ability, we’re making them more aligned with how people actually perceive their environment. That could make human-AI interaction more natural, whether it’s a robot guiding you through a building or an AR assistant telling you where to look.

Tom: That’s a beautiful way to put it, Lalam. And with that, we’re going to say goodbye to BAT. It’s been a fantastic paper — the first of its kind, with a solid dataset, a clever architecture, and a training recipe that actually teaches reasoning. We’ll be watching to see where this line of research goes.

Jane: Absolutely. And if you want to check it out yourself, the authors have released the demo, dataset, code, and model weights. So you can actually play with it. Thanks for listening, everyone — we’ll see you on the next one.

Tom: Take care, and keep listening.

More episodes

← Home