Audio-Maestro: Enhancing Large Audio-Language Models with Tool-Augmented Reasoning
cs.SD, cs.AI
Submitted: 2025-10-13
Updated: 2026-09-06
Comments: Accepted in ISCSLP 2026
Code: https://github.com/gary920209/Audio-Maestro
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Recent advancements in large multimodal models (LMMs) have shown strong capabilities in audio understanding.
Terminology
Abstract
Recent advancements in large multimodal models (LMMs) have shown strong capabilities in audio understanding. However, most systems rely solely on end-to-end reasoning, limiting interpretability and accuracy for tasks that require structured knowledge or specialized signal analysis. In this work, we present Audio-Maestro -- a tool-augmented audio reasoning framework that enables audio-language models to autonomously call external tools and integrate their timestamped outputs into the reasoning process. This design allows the model to analyze, transform, and interpret audio signals through specialized tools rather than relying solely on end-to-end inference. Experiments show that Audio-Maestro consistently improves general audio reasoning performance: Gemini-2.5-flash's average accuracy on MMAU-Test rises from 67.4% to 72.1%, DeSTA-2.5 from 58.3% to 62.8%, and GPT-4o from 60.8% to 63.9%. To our knowledge, Audio-Maestro is the first framework to integrate structured tool output into the large audio language model reasoning process.
Sources
- Mellow: a small audio language model for reasoning
- GPT-4o System Card
- Brouhaha: multi-task training for voice activity detection, speech-to-noise ratio, and C50 room acoustics estimation
- Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
- ART: Automatic multi-step reasoning and tool-use for large language models
- Gemini: A Family of Highly Capable Multimodal Models
- emotion2vec: Self-Supervised Pre-Training for Speech Emotion Representation
- TASTE: Text-Aligned Speech Tokenization and Embedding for Spoken Language Modeling
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- Step-Audio 2 Technical Report
- Audio-Thinker: Guiding Audio Language Model When and How to Think via Reinforcement Learning
- Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
- Omni-R1: Reinforcement Learning for Omnimodal Reasoning via Two-System Collaboration
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment