TimeSteer: Inference-Time Speech Scheduling in Joint Audio-Visual Diffusion Models
cs.CV, cs.AI, cs.MM
Submitted: 2026-09-01
Updated: 2026-09-01
License: http://creativecommons.org/licenses/by/4.0/
The gist: Although pretrained joint audio-visual diffusion models offer rich control over what to generate, they provide no explicit control over when an utterance should occur.
Terminology
Abstract
Although pretrained joint audio-visual diffusion models offer rich control over what to generate, they provide no explicit control over when an utterance should occur. To address this, we study inference-time speech scheduling, a novel task that places coupled speech and visual articulation within user-specified begin--end intervals without finetuning the backbone model. We uncover two intrinsic properties of the denoising process that enable this task. First, a timing-sensitive text-to-audio cross-attention head exposes each utterance's model-implied source span along the latent timeline. Second, the predicted clean latent already organizes coupled speech and visual articulation, allowing their temporal placement to be edited without regenerating the content. Building on these discoveries, we propose TimeSteer, a training-free framework that localizes each utterance's source span through Source Span Localization and transfers the associated audio-visual latent content from the source interval to the specified target interval through Region-Aware Latent Remapping. We further introduce SpeechShift, the first benchmark for interval-level speech scheduling in joint audio-visual generation. Experiments across two representative backbones show that TimeSteer substantially improves interval controllability over training-free baselines while maintaining competitive overall generation quality.
Sources
- Prompt Relay: Inference-Time Temporal Control for Multi-Event Video Generation
- Unison: Harmonizing Motion, Speech, and Sound for Human-Centric Audio-Video Generation
- Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model
- LTX-2: Efficient Joint Audio-Visual Foundation Model
- A geometric proof of the infinite $(p, q)$-theorem for hyperplane piercing
- Make-An-Audio 2: Temporal-Enhanced Text-to-Audio Generation
- TempoControl: Temporal Attention Guidance for Text-to-Video Models
- Syncphony: Synchronized Audio-to-Video Generation with Diffusion Transformers
- PicoAudio2: Temporal Controllable Text-to-Audio Generation with Natural Language Description
- Harmony: Harmonizing Audio and Video Generation through Cross-Task Synergy
- AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation
- SyncDPO: Enhancing Temporal Synchronization in Video-Audio Joint Generation via Preference Learning
- Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
- Inference-Time Scaling for Joint Audio-Video Generation
- Movie Gen: A Cast of Media Foundation Models
- AADiff: Audio-Aligned Video Synthesis with Text-to-Image Diffusion
- Native Audio-Visual Alignment for Generation
- Meta Audiobox Aesthetics: Unified Automatic Quality Assessment for Speech, Music, and Sound
- JavisDiT: Joint Audio-Video Diffusion Transformer with Hierarchical Spatio-Temporal Prior Synchronization
- SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models