StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios
cs.SD, cs.AI
Submitted: 2026-09-03
Updated: 2026-09-03
Terminology
Sources
- CMGAN: Conformer-based Metric GAN for Speech Enhancement
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Gen-Searcher: Reinforcing Agentic Search for Image Generation
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- An Analysis of the Variance of Diffusion-based Speech Enhancement
- Diffusion Models for Audio Restoration
- ICASSP 2026 URGENT Speech Enhancement Challenge
- JarvisArt: Liberating Human Artistic Creativity via an Intelligent Photo Retouching Agent
- MP-SENet: A Speech Enhancement Model with Parallel Denoising of Magnitude and Phase Spectra
- NISQA: A Deep CNN-Self-Attention Model for Multidimensional Speech Quality Prediction with Crowdsourced Datasets
- The INTERSPEECH 2020 Deep Noise Suppression Challenge: Datasets, Subjective Testing Framework, and Challenge Results
- UTMOS: UTokyo-SaruLab System for VoiceMOS Challenge 2022
- Interspeech 2025 URGENT Speech Enhancement Challenge
- Proximal Policy Optimization Algorithms
- Gemini: A Family of Highly Capable Multimodal Models
- Audio-Reasoner: Improving Reasoning Capability in Large Audio Language Models
- Qwen3-Omni Technical Report
- TIGER: Time-frequency Interleaved Gain Extraction and Reconstruction for Efficient Speech Separation
- DanceGRPO: Unleashing GRPO on Visual Generation
- EmbodiedBench: Comprehensive Benchmarking Multi-modal Large Language Models for Vision-Driven Embodied Agents
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment