VIBE: Video Instruction-aligned Background music gEneration

arXiv:2608.30125 · cs.SD, cs.AI, cs.CL, cs.CV, cs.LG · Submitted 2026-08-31 · Read on arXiv

cs.SD, cs.AI, cs.CL, cs.CV, cs.LG

Submitted: 2026-08-31

Updated: 2026-08-31

License: http://creativecommons.org/licenses/by/4.0/

The gist: Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representational bottleneck

Terminology

Abstract

Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representational bottleneck of static cross-modal conditioning in Diffusion Autoregressive (DAR) architectures. To resolve this, we introduce VIBE, a novel text-and-video-to-music (T+V2M) generation model that leverages: (1) Conditioning Connection, a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and (2) a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints (e.g., tempo, key) and soft, subjective qualities (e.g., musicality, multimodal alignment) with a structured 5-stage training curriculum. Upon evaluation using audio-visual alignment, instruction following, and audio quality metrics, along with a subjective human evaluation study, we observe that VIBE demonstrates enhanced controllability and instruction adherence while performing comparably to most evaluated baselines on generation fidelity and multimodal alignment.

Sources

Related papers