Spatial-Omni: Spatial Audio Understanding Integration in Multimodal LLMs via FOA Encoding
eess.AS, cs.AI
Submitted: 2026-06-09
Updated: 2026-09-07
Code: https://github.com/dieKarotte/Spatial-Omni
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial relation reasoning,
Terminology
Abstract
Recent multimodal large language models mainly process audio as monaural signals, thereby discarding the spatial cues contained in spatial audio for sound localization, spatial relation reasoning, and spatial scene understanding. We propose Spatial-Omni, a lightweight method that implements SO-Encoder to inject First-Order Ambisonics (FOA) spatial audio into existing Omni LLMs as an independent modality, without modifying their original audio encoders. SO-Encoder provides spatial tokens with limited additional context cost and improves spatial audio understanding through efficient staged training. To support training and evaluation, we construct SO-Dataset, SO-QA, and SO-Bench from open-source data, real recordings, and simulations, containing 400K FOA spatial audio clips and 2.1M spatial question answering pairs. SO-Bench covers 16 spatial audio understanding subtasks, including basic detection and location estimation, spatial relation understanding, and complex spatial reasoning. Experiments show that Spatial-Omni outperforms existing open-source Large Audio-Language Models (LALMs) and Omni LLM models on spatial audio understanding tasks while retaining a reasonable level of general audio understanding. Code and data are available at https://github.com/dieKarotte/Spatial-Omni.
Sources
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Qwen2-Audio Technical Report
- GPT-4 Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- PhaseCoder: Microphone Geometry-Agnostic Spatial Audio Understanding for Multimodal LLMs
- Kimi-Audio Technical Report
- OWL: Geometry-Aware Spatial Reasoning for Audio Large Language Models
- AST: Audio Spectrogram Transformer
- JAEGER: Joint 3D Audio-Visual Grounding and Reasoning in Simulated Physical Environments
- STAR-Bench: Probing Deep Spatio-Temporal Reasoning as Audio 4D Intelligence
- SALM: Spatial Audio Language Model with Structured Embeddings for Understanding and Editing
- Deep Learning for Personalized Binaural Audio Reproduction
- Audio Flamingo: A Novel Audio Language Model with Few-Shot Learning and Dialogue Abilities
- Spatial Audio Processing with Large Language Model on Wearable Devices
- OmniAudio: Generating Spatial Audio from 360-Degree Video
- STARSS22: A dataset of spatial recordings of real scenes with spatiotemporal annotations of sound events
- Towards Spatial Audio Understanding via Question Answering
- Both Ears Wide Open: Towards Language-Driven Spatial Audio Generation
- Habitat-Matterport 3D Dataset (HM3D): 1000 Large-scale 3D Environments for Embodied AI
- AudioMotionBench: Evaluating Auditory Motion Perception in Audio LLMs
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions