OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
eess.AS, cs.AI, eess.IV
Submitted: 2026-09-18
Updated: 2026-09-28
License: http://creativecommons.org/licenses/by/4.0/
The gist: We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model.
Terminology
Abstract
We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user's query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two constraints: data availability and evaluation. Recordings of people using their own devices are scarce. Furthermore, a good reply often needs to account for the user's surroundings, facial expressions, and nearby objects, and such responses can be expressed in many different ways, making keyword matching unreliable for evaluating reply quality. Recent progress in agent systems and video generation makes generation for comprehension viable, which means using synthesized dialogues for training and evaluation. Therefore, we present OmniVChat-Studio, a multi-agent data engine for synthesizing single- and multi-turn audio-visual dialogues. We use synthesized dialogues to build OmniVChat-Bench, an evaluation benchmark that evaluates omni models' basic dialogue abilities across five ability categories. We also present OmniVChat-RL, a reinforcement learning reward design that jointly targets reply correctness, efficiency, and style in OmniVChat. Training Qwen3-Omni-Instruct with OmniVChat-RL on synthesized dialogues improves its performance on both OmniVChat-Bench and the human-recorded OmniVChat-Bench-Human. These gains validate the reward design and show transfer to real-world dialogues in training and evaluation.
Sources
- OmniInteract: Benchmarking Real-World Streaming Interaction for Real-Time Omnimodal Assistants
- JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
- A Survey of AI Agent Protocols
- Controllable Video Generation: A Survey
- Qwen3-Omni Technical Report
- Group Sequence Policy Optimization
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- LoRA: Low-Rank Adaptation of Large Language Models
- Qwen3.5-Omni Technical Report
- Seed2.0 Model Card: Towards Intelligence Frontier for Real-World Complexity
- Qwen2.5-Omni Technical Report
- MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
- Qwen3-ASR Technical Report
- Baichuan-Omni-1.5 Technical Report
- Ola: Pushing the Frontiers of Omni-Modal Language Model
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- InternLM-XComposer2.5-OmniLive: A Comprehensive Multimodal System for Long-term Streaming Video and Audio Interactions
- video-SALMONN 2: Caption-Enhanced Audio-Visual Large Language Models
- AV-Odyssey Bench: Can Your Multimodal LLMs Really Understand Audio-Visual Information?
- Daily-Omni: Towards Audio-Visual Reasoning with Temporal Alignment across Modalities
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions