Audible World Models: Spatially Aware Sound Generation for 3D Worlds
cs.CV, cs.SD
Submitted: 2026-09-29
Updated: 2026-09-29
Terminology
Sources
- ZoeDepth: Zero-shot Transfer by Combining Relative and Metric Depth
- LucidDreamer: Domain-free Generation of 3D Gaussian Splatting Scenes
- Audio Flamingo 2: An Audio-Language Model with Long-Audio Understanding and Expert Reasoning Abilities
- TangoFlux: Super Fast and Faithful Text to Audio Generation with Flow Matching and Clap-Ranked Preference Optimization
- ViSAGe: Video-to-Spatial Audio Generation
- AudioGen: Textually Guided Audio Generation
- Neural Acoustic Context Field: Rendering Realistic Room Impulse Response With Neural Fields
- AudioLDM: Text-to-Audio Generation with Latent Diffusion Models
- OmniAudio: Generating Spatial Audio from 360-Degree Video
- DreamFusion: Text-to-3D using 2D Diffusion
- SAM 2: Segment Anything in Images and Videos
- HunyuanVideo-Foley: Multimodal Diffusion with Representation Alignment for High-Fidelity Foley Audio Generation
- Lyra 2.0: Explorable Generative 3D Worlds
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World Modeling
- HunyuanWorld 1.0: Generating Immersive, Explorable, and Interactive 3D Worlds from Words or Pixels
- Audiobox: Unified Audio Generation with Natural Language Prompts
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models