BAT: Learning to Reason about Spatial Sounds with Large Language Models

arXiv:2402.01591 · eess.AS, cs.AI, cs.CL, cs.SD · Submitted 2026-08-13 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "BAT: Learning to Reason about Spatial Sounds with Large Language Models".

Jane: The paper was written by Zhisheng Zheng, Puyuan Peng, Ziyang Ma, Xie Chen, Eunsol Choi et al. from University of Texas at Austin and Shanghai Jiao Tong University.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Title: Tom: Welcome back to the show, everyone. Today we’re digging into a paper that’s got me genuinely excited — it’s called “BAT: Learning to Reason about Spatial Sounds with Large Language Models.” Jane, you’ve been grinning since we started reading this one.

Jane: I really have, Tom. This paper is about teaching a large language model to actually understand where sounds are coming from in a three dee space. You know, not just “I hear a dog barking,” but “the dog is barking to my left, about three meters away, and there’s a stereo playing behind me at the same time.” That’s a huge leap from what most audio models can do.

Tom: And that’s the kicker, right? Most of the audio LLMs out there — they work on monaural audio, which is just a single channel, like one ear. They can tell you what’s making a sound, but they have no idea where it’s coming from. This paper, BAT, uses binaural audio — two channels, like having two ears — and it actually reasons about space.

Jane: Exactly. And the way they did it is pretty clever. They built a new audio encoder called Spatial-AST, which takes in the two-channel audio and extracts three things at once: what the sound is, which direction it’s coming from, and how far away it is. Then they hook that up to a large language model — LLaMA-two specifically — so the model can answer questions like “Is the dog to the left of the stereo?”

Tom: And it’s not just a toy demo. They created a whole new dataset called Spatial Sound QA, with over eight hundred thousand question-answer pairs. They simulated realistic indoor environments with reverberation, using SoundSpaces two point zero and AudioSet clips. So the model is trained on sounds bouncing off walls, coming from different angles and distances — messy, real-world-ish conditions.

Jane: That’s the part that really impressed me. They didn’t just record a few sounds in a studio. They built a virtual world with ninety different buildings, thousands of rooms, and they placed sound sources at random locations, then simulated how the sound would actually reach a listener’s two ears. That’s a lot of careful engineering.

Tom: And the results? The Spatial-AST encoder alone gets a mean average precision of about fifty percent on sound event detection, and it can estimate direction within about eighteen degrees of error on average. Then when you combine it with the LLM, the full BAT model answers spatial reasoning questions with about seventy-seven percent accuracy. That’s way above random guessing, which would be fifty percent.

Jane: Right, and that seventy-seven percent is on questions that require the model to separate two overlapping sounds, figure out where each one is, and then reason about their relationship. That’s not just perception — that’s actual spatial reasoning. I think this could be a big deal for robotics, for augmented reality, even for hearing aid technology.

Tom: I was just thinking about that. Imagine a robot that can hear you call from across a room and actually turn toward you. Or a hearing aid that can tell you “the person speaking is to your right, three meters away.” This paper lays the groundwork for all of that.

Jane: And it’s the first of its kind, Tom. The authors say it’s the first spatial audio-based LLM. So we’re at the very beginning of something new here.

Tom: I love it. So what’s next? We’ve got the title and the big picture — let’s get into the actual method and how they made this work.

Summary: Tom: So we’ve established that “BAT: Learning to Reason about Spatial Sounds with Large Language Models” is a big deal. But let’s get into the nitty-gritty of how they actually built it. Jane, you’re the one who can explain this in plain English.

Jane: Happy to, Tom. So the core idea is pretty simple: they take a two-channel audio clip — left ear and right ear — and they process it in two stages. First, they have this encoder called Spatial-AST. It takes the audio and turns it into a set of tokens, like a compressed summary of what’s happening. But here’s the clever part: they add three special tokens, each one trained to focus on a different thing — one for the sound class, one for direction, one for distance.

Tom: So it’s like having three little specialists inside the encoder, each one looking for a different piece of information.

Jane: Exactly. And they train it on three tasks at once: “what is this sound?”, “which direction is it coming from?”, and “how far away is it?” They use cross-entropy loss for all three, but they train in two stages. First, they just focus on classification, because that’s the hardest. Then they add the direction and distance losses.

Tom: And that two-stage approach really mattered, right? I saw in the paper that training all three at once from the start hurt performance.

Jane: Yeah, the numbers show it clearly. With binaural audio and IPD — that’s interaural phase difference, the tiny timing difference between when a sound hits your left ear versus your right ear — the two-stage training gets the direction error down to about eighteen degrees. Joint training, where they do everything at once, gets about eighteen point six degrees. It’s a small difference, but it’s consistent across all the metrics.

Tom: And what about the distance estimation? That’s the part that surprised me.

Jane: Me too. They measure distance error rate — basically, how often the model is within half a meter of the true distance. With the full binaural setup, they get about thirty-two point five percent error rate. That sounds low, but remember, the distances range from zero to ten meters, and the model has to pick from discrete half-meter bins. Random guessing would be way worse.

Tom: So the encoder itself is already a strong spatial audio model. But then they go further and hook it up to the LLM. How does that work?

Jane: They take the output tokens from Spatial-AST — including those three specialist tokens — and they project them into the embedding space of LLaMA-two. Then they fine-tune the LLM with a parameter-efficient method called LLaMA-Adapter v2. The LLM stays mostly frozen, but they add a few trainable layers so it can learn to interpret the audio tokens.

Tom: And that’s where the reasoning comes in. The model can answer questions like “Is the sound of the dog to the left of the sound of the stereo?” — and it has to figure that out from the raw audio, not from any text description.

Jane: Right. And they trained it with a curriculum. Stage one is single-source questions — just “what do you hear and where is it?” Stage two introduces two sources at once, which forces the model to implicitly separate them. Stage three adds the reasoning questions. And the ablations show that if you skip stage two, the reasoning ability collapses to near random.

Tom: That’s a really important finding. It means the model isn’t just memorizing patterns — it actually needs to learn how to separate overlapping sounds before it can reason about their spatial relationships.

Jane: Exactly. And that’s the kind of insight that makes this paper more than just a demo. It tells us something about how to train these models effectively.

Tom: So we’ve got the encoder, the training curriculum, the dataset. But what does this actually mean for the world? Let’s bring in Lu and Meng for that.

Improvements: Tom: Alright, we’ve covered the basics of “BAT: Learning to Reason about Spatial Sounds with Large Language Models.” Now let’s talk about what this paper improves on and where it could go. Lu, you’ve been thinking about the big picture — what’s the real step forward here?

Lu: Thanks, Tom. I think the biggest improvement is that this paper moves audio LLMs from “what” to “where.” Before BAT, models like LTU or Pengi could tell you what sounds were present, but they were working with monaural audio — they had no spatial information at all. BAT is the first to bring binaural spatial cues into the LLM pipeline. That’s a fundamental capability upgrade.

Jane: And it’s not just about adding a new input channel. The paper shows that the model can actually reason about relationships between multiple sound sources. That’s something that no previous audio LLM could do.

Lu: Exactly. And that reasoning ability is what opens the door to real-world applications. Think about autonomous robots navigating a home — they need to know not just that there’s a person talking, but where that person is relative to the robot. BAT provides a template for that.

Meng: I’m with you on the potential, Lu, but let me play the engineer for a second. How practical is this to actually deploy? The paper uses LLaMA-two 7B, which is a pretty big model. And they trained on eight V100 GPUs. That’s not exactly edge-device friendly.

Tom: That’s a fair point, Meng. But they do use parameter-efficient fine-tuning — LLaMA-Adapter v2 — which means most of the LLM stays frozen. So the actual trainable parameters are small. And the encoder, Spatial-AST, is only about ninety million parameters. That part could run on a modest GPU.

Meng: Okay, that’s more reasonable. But what about the dataset? They synthesized everything with SoundSpaces two point zero. That’s a simulator. How well does that transfer to real-world audio? There’s always a sim-to-real gap.

Lu: That’s the big open question, and the authors acknowledge it. They mention it in the limitations section. But the fact that they used real AudioSet clips as sound sources, and real room geometries from Matterportthree dee, means the audio is at least grounded in reality. The reverberation is simulated, but it’s based on actual room layouts.

Jane: And they also show that the IPD — the phase difference between the two ears — is crucial for direction estimation. That’s a physical cue that exists in real binaural recordings too, so the model should generalize at least partially.

Meng: I’d still want to see a real-world test before I’d trust it in a product. But the architecture is sound, and the training curriculum is smart. The fact that they show you need the multi-source separation stage to get reasoning to work — that’s a practical lesson for anyone building this kind of model.

Tom: So what’s the next big improvement you’d want to see?

Lu: I’d love to see them extend this to more than two sound sources. Right now the reasoning questions are all binary — two sources. Real environments have many overlapping sounds. Also, they only handle static sources — no moving sounds. Tracking a moving sound source would be a natural next step.

Meng: And I’d like to see a smaller, distilled version of the model. If you could get the reasoning ability into a model that runs on a phone or a smart speaker, that would be huge for consumer applications.

Jane: The authors also mention integrating visual information. SoundSpaces two point zero actually supports visual rendering too, so you could train a model that understands both what it sees and what it hears in a three dee environment. That’s a really exciting direction.

Tom: Alright, so we’ve got the improvements and the future directions. Let’s wrap this up with a conclusion and say goodbye to BAT.

Conclusion: Tom: Alright, we’ve spent a lot of time with “BAT: Learning to Reason about Spatial Sounds with Large Language Models,” and I think it’s fair to say this is one of those papers that marks a turning point. Jane, how would you sum it up for someone who just tuned in?

Jane: I’d say BAT is the first large language model that can actually hear in three dee. It takes binaural audio — two channels, like two ears — and it can tell you what sounds are present, where they’re coming from, how far away they are, and even reason about the spatial relationships between multiple sounds. It does this using a new encoder called Spatial-AST, which is trained on a massive simulated dataset called Spatial Sound QA.

Tom: And the key numbers — the encoder gets about fifty percent mean average precision on sound event detection, direction error around eighteen degrees, and the full BAT model answers spatial reasoning questions with about seventy-seven percent accuracy. That’s a huge step up from random guessing.

Lu: And the implications go beyond just audio. This shows that LLMs can be extended to handle spatial reasoning in a modality that’s inherently three-dimensional. That’s relevant for robotics, augmented reality, assistive technology, and even gaming.

Meng: From a practical standpoint, the training curriculum is the most valuable takeaway. The paper proves that you can’t just throw all the tasks at the model at once — you need to build up from single-source perception to multi-source separation before reasoning works. That’s a lesson that applies to any multimodal LLM.

Lalam: I’d add that this work has cultural implications too. Spatial audio is how humans experience the world — we don’t just hear sounds, we place them in space. By giving LLMs that ability, we’re making them more aligned with how people actually perceive their environment. That could make human-AI interaction more natural, whether it’s a robot guiding you through a building or an AR assistant telling you where to look.

Tom: That’s a beautiful way to put it, Lalam. And with that, we’re going to say goodbye to BAT. It’s been a fantastic paper — the first of its kind, with a solid dataset, a clever architecture, and a training recipe that actually teaches reasoning. We’ll be watching to see where this line of research goes.

Jane: Absolutely. And if you want to check it out yourself, the authors have released the demo, dataset, code, and model weights. So you can actually play with it. Thanks for listening, everyone — we’ll see you on the next one.

Tom: Take care, and keep listening.

Zhisheng Zheng, Puyuan Peng, Ziyang Ma, Xie Chen, Eunsol Choi, David Harwath

University of Texas at Austin · Shanghai Jiao Tong University

eess.AS, cs.AI, cs.CL, cs.SD

Submitted: 2026-08-13

Updated: 2026-08-17

Comments: Accepted to ICML 2024. Our demo, dataset, code and model weights are available at: https://zhishengzheng.com/bat

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 68/100

The gist: BAT: Learning to Reason about Spatial Sounds with Large Language Models Abstract: "Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based

Terminology

Summary

BAT: Learning to Reason about Spatial Sounds with Large Language Models

Abstract: "Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed S PATIAL S OUND QA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or S PATIAL-AST, which by itself achieves strong performance across sound event detection, spatial localization, and distance estimation. By integrating S PATIAL-AST with LLaMA-2 7B model, BAT transcends standard Sound Event Localization and Detection (SELD) tasks, enabling the model to reason about the relationships between the sounds in its environment. Our experiments demonstrate BAT’s superior performance on both spatial sound perception and reasoning, showcasing the immense potential of LLMs in navigating and interpreting complex spatial audio environments."

Introduction: "To bridge this gap, we introduce BAT, the first spatial audio-based LLM designed to reason about sounds in a 3-D environment. Recognizing the current lack of large-scale datasets for in-the-wild spatial audio and spatial audio-based question answering, we synthesized a large-scale binaural audio dataset using Audioset clips as sound sources and Soundspaces 2.0 for simulating diverse, 3-D, reverberant acoustic environments. In conjunction with this dataset, we developed S PATIAL S OUND QA, a diverse collection of question answering tasks designed to train and evaluate spatial sound understanding across varying levels of complexity."

Key contributions: "We present the first spatial audio-based question answering dataset S PATIAL S OUND QA, offering a range of 3-D audio understanding tasks from perception to reasoning... We propose S PATIAL-AST, a binaural spatial audio encoder architecture that jointly performs sound event detection, spatial localization, and distance estimation that achieves strong performance across all three tasks. We introduce BAT, which integrates S PATIAL-AST with the LLaMA-2 LLM, resulting in a model capable of answering complex reasoning questions about multiple sound sources situated within a 3-D environment."

Spatial Audio Generation: "We use the state-of-the-art audio simulator, SoundSpaces 2.0... This platform performs on-the-fly geometry-based sound rendering, enabling realistic acoustic reverberation with arbitrary source-receiver locations... we leverage Matterport3D for our environmental meshes, which includes highly detailed mesh renderings of 90 complete buildings... For a given arbitrary source location s, a monaural sound source As, receiver location r, and the receiver’s heading direction θ in a specific mesh environment, the audio signal Ar received by a microphone is given by the convolution of the room impulse response with the sound source... To minimize perceptual dissonance in auditory localization when visual cues are absent, we ensure that both the sound source and the receiver are located within the same room... In total, 21,131 reverberations are generated, averaging approximately 9.58 reverberations per room."

Sound Sources: "We sample from AudioSet to specify our monaural sound source As. AudioSet consists of roughly 2 million 10-second in-the-wild YouTube clips used for audio classification... we exclude labels that require visual information by manual inspection... we further exclude most noise-related labels... This curation results in a set of 355 audio event labels that are appropriate for our setting, identifiable solely through audio cues. After filtering by these categories, we are left with 1,861,750 clips in the AudioSet-2M split, while our filtered AudioSet-20K split contains 18,373 clips, and our evaluation set consists of 17,148 clips."

Question-Answer Pair Generation: "We curated S PATIAL S OUND QA to include diverse question-answer pairs, all centered around the challenges of spatial audio perception and reasoning... each sample in the dataset is structured as a tuple of (audio, question, answer)... The questions in these tasks are paraphrased using GPT-4... The answers, on the other hand, are generated uniformly through a systematic rule-based approach. The dataset includes five question types: A: Detection (139K, 15.9%) with 1 source, B: DoA & DP (139K, 15.9%) with 1 source, C: Detection (118K, 13.5%) with 2 sources, D: DoA & DP (118K, 13.5%) with 2 sources, and E: Reasoning (358K, 41.2%) with 2 sources. For spatial reasoning, the three-dimensional space is segmented into eight regions. Each region is defined by a tuple representing the directional axis (left/right, front/behind, above/below) with respect to the receiver. Additionally, distance is quantified in increments of 0.5 meters, spanning a range from 0 to 10 meters."

Spatial Audio Encoder (S PATIAL-AST): We propose a novel architecture S PATIAL-AST to capture spatial audio information. The front-end "integrates both Mel-Spectrogram and Interaural Phase Difference (IPD)... we initially transform the time-domain signal x(n) into the frequency domain X(t, f) using the short-time Fourier transform (STFT)... we compute both the Mel-Spectrogram and the Interaural Phase Difference (IPD)... The final front-end output, Z, is a concatenation of these processed components. It includes both the Mel-Spectrograms of the left and right channels, as well as the cosine and sine transformations of the IPD. The backbone uses a 3x3 2D convolution followed by a batch normalization layer and a GELU layer... we employ a Patch-Embed CNN... we concatenate three [CLS] tokens at the beginning of the audio tokens, each specifically designated for extracting information about the audio’s category, distance, and direction respectively. All tokens are then fed into 12-layer Transformer encoder blocks. The pre-training objective is L = λ1 Lcls + λ2 Ldis + λ3 Ldoa for detection, distance, and direction, with cross-entropy loss for all three tasks and targets discretized into intervals of 0.5 meters and angle into intervals of 1 degree."

BAT Model: "To extend this ability to encompass spatial reasoning about multiple sound sources present within an environment, we fuse S PATIAL-AST to the LLaMA-2 7B LLM. Input spatial audio is first processed by the S PATIAL-AST encoder, and then a projection module is used to map its set of output tokens into LLaMA-2’s input text embedding space. For efficient fine-tuning, we use the LLaMA-adapter v2. The fine-tuning objective is to predict the answer text conditioned on its paired question text and the corresponding audio input... maximizing the probability of predicting the next answer token. The training uses a perception-to-reasoning curriculum with three stages: Stage I (A, B; 278K samples; 31.8%), Stage II (A, B, C, D; 514K samples; 58.8%), Stage III (A, B, C, D, E; 872K samples; 100%)."

Experimental Results for S PATIAL-AST: Table 3 presents the experimental results for our proposed S PATIAL-AST audio encoder. For binaural input with IPD and two-stage training, S PATIAL-AST achieves mAP of 50.03%, ER20° of 23.89, MAE of 17.94°, and DER of 32.54%. For monaural input with joint training, it achieves mAP of 51.39, ER20° of 95.85, MAE of 88.52, and DER of 26.87. The paper notes: Binaural data significantly improves performance in tasks involving direction perception... The performance of monaural data suggests it is sufficiently effective for SED and DP tasks. Also, "two-stage training consistently demonstrates improvements in its second stage, regardless of whether IPD is used. The enhancement is particularly significant with the usage of IPD, achieving the lowest MAE of 23.89 and DER of 17.94."

Experimental Results for BAT: Table 4 shows the performance of various configurations of our BAT model on S PATIAL S OUND QA. The full three-stage BAT with binaural input achieves: Detection (mAP) of 0.69 0.65 for types A and C, DoA (Acc) of 75.54 37.65 for types B and D, DP (DER) of 29.16 47.90 for types B and D, and Reasoning (BA) of 69.77 (direction), 84.01 (distance), and 76.89 (average). The one-stage BAT (only stage III) achieves lower reasoning performance: 65.12 (direction), 81.98 (distance), 73.55 (average). The monaural BAT achieves Binary Accuracy (BA) of only 54.53% in reasoning, showing monaural audio input alone cannot provide sufficient spatial information for complex reasoning challenges. The paper highlights: "We observed a significant drop in the model’s reasoning ability when question types C and D, related to sound source separation, were removed from the training set. This is evident from the Binary Accuracy (BA) results, which approximate random chance in the reasoning tasks."

Limitations and Future Work: "One of the primary challenges we encounter is extracting as much information as possible from spatial audio, with a specific focus on how to better utilize phase information. Currently, our model handles a maximum of two sound sources, with an emphasis primarily on audio. Looking ahead, there is potential to expand into multi-source scenarios and to integrate both audio and speech processing for a more holistic approach. Additionally, while our current framework is limited to binaural audio, exploring ambisonics could provide a more immersive and realistic spatial audio experience. Moreover, expanding S PATIAL S OUND QA to be more open-ended would better align with human usage patterns... Another important consideration is the integration of additional modalities, such as visual information... Lastly, the sim2real gap remains an aspect that requires further investigation."

Conclusion: "In this work, we have presented BAT, the first LLM capable of processing spatial audio input. To train and evaluate BAT, we introduced S PATIAL S OUND QA, the first extensive spatial audio-based question answering dataset. We also proposed S PATIAL-AST, a novel spatial audio encoder capable of efficiently handling sound event detection, spatial localization, and distance estimation. BAT and S PATIAL S OUND QA showcase the immense potential of LLMs for reasoning about spatial sound. S PATIAL S OUND QA also enables rich future work, such as reasoning about the environment itself (materials, shape, layout), modeling moving sounds to for tracking problems, or incorporating the visual modality which SoundSpaces2.0 natively supports."

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems, along with what the improved system can do:


Improvement: Integrate a novel binaural audio encoder that fuses Mel-spectrograms with Interaural Phase Difference (IPD) features, processed through a 12-layer transformer with three dedicated [CLS] tokens for classification, direction-of-arrival, and distance estimation. Use a two-stage training curriculum (first detection-only, then multi-task) with loss weighting (λ1=1250, λ2=1, λ3=2).

What the improved system can do:

  • Detect 355 distinct sound event classes from reverberant, in-the-wild audio with 50.03% mAP (vs. 47.18% for AudioMAE).

  • Estimate sound source direction with a mean angular error of 17.94° (vs. 88.52° for monaural input).

  • Predict sound source distance within 0.5 meters with a 32.54% error rate (vs. 67% random baseline).

  • Process both monaural and binaural inputs, with binaural providing significant gains in direction estimation.

Improvement: Fuse the Spatial-AST encoder with LLaMA-2 7B using LLaMA-adapter v2 (trainable zero-init attention, projection, norm, bias, scale). Train with a three-stage curriculum: (I) single-source perception (types A, B), (II) multi-source perception (types C, D), (III) spatial reasoning (type E). Use a projection module with 64 learnable query tokens to map audio features into text embedding space.

Improvement: Generate a large-scale, rule-based QA dataset using SoundSpaces 2.0 (Matterport3D environments) and AudioSet clips. Include five question types: (A) single-source detection, (B) single-source direction/distance, (C) multi-source detection with location constraints, (D) multi-source direction/distance, (E) spatial reasoning (binary Yes/No). Use GPT-4 to paraphrase questions for diversity.

Improvement: Implement a progressive training schedule: Stage I (31.8% of data, types A+B), Stage II (58.8%, types A+B+C+D), Stage III (100%, all types). This prevents catastrophic forgetting and builds reasoning on top of solid perception.

Improvement: Use 32kHz binaural audio, 1024-point STFT with hop 320, 128 mel-bins, and compute cosine/sine-transformed IPD to avoid phase wraparound. Apply loudness normalization and two-stage training (first detection-only, then multi-task) to balance task performance.

Improvement: Use a text-embedding-based similarity metric (text-embedding-3-small) for open-ended answers, alongside standard metrics (mAP, ER20°, MAE, DER, binary accuracy). Include random and oracle baselines for calibration.

Improvement: Use SoundSpaces 2.0 with Matterport3D meshes (90 buildings, 24.5 rooms average) and AudioSet's 1.86M clips for diverse, realistic training. Exclude visually-dependent and noise-only labels to ensure audio-only learnability.

The improved system (BAT + Spatial-AST + SpatialSound QA) can:

  • Listen to binaural audio and answer natural language questions about what sounds are present, where they come from, and how far away they are.

  • Reason about spatial relationships between multiple simultaneous sounds (e.g., Is the dog to the left of the music?).

  • Perform sound event detection, direction-of-arrival estimation, and distance prediction in a single unified model.

  • Train efficiently on simulated data and be ready for real-world deployment in robotics, virtual reality, hearing aids, and smart home assistants.

Abstract

Spatial sound reasoning is a fundamental human skill, enabling us to navigate and interpret our surroundings based on sound. In this paper we present BAT, which combines the spatial sound perception ability of a binaural acoustic scene analysis model with the natural language reasoning capabilities of a large language model (LLM) to replicate this innate ability. To address the lack of existing datasets of in-the-wild spatial sounds, we synthesized a binaural audio dataset using AudioSet and SoundSpaces 2.0. Next, we developed SpatialSoundQA, a spatial sound-based question-answering dataset, offering a range of QA tasks that train BAT in various aspects of spatial sound perception and reasoning. The acoustic front end encoder of BAT is a novel spatial audio encoder named Spatial Audio Spectrogram Transformer, or Spatial-AST, which by itself achieves strong performance across sound event detection, spatial localization, and distance estimation. By integrating Spatial-AST with LLaMA-2 7B model, BAT transcends standard Sound Event Localization and Detection (SELD) tasks, enabling the model to reason about the relationships between the sounds in its environment. Our experiments demonstrate BAT's superior performance on both spatial sound perception and reasoning, showcasing the immense potential of LLMs in navigating and interpreting complex spatial audio environments.

Sources

Related papers