Ex-Omni-2D: Expressive Omni-Modal Dialogue Models with Native Visual Presence
The Chinese University of Hong Kong, Shenzhen · LIGHTSPEED · Independent Researcher
cs.AI, cs.CL, cs.CV
Submitted: 2026-08-11
Updated: 2026-09-03
Code: https://github.com/tatsu-lab/alpaca_eval
Project page: https://logo-cuhksz.github.io/Ex-Omni-2D
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 100/100
The gist: Ex-Omni-2D is an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video.
Terminology
Summary
Ex-Omni-2D is an omni-modal dialogue framework that generates a coordinated response comprising text, personalized speech, and reference-conditioned video. Given a multimodal query, a reference image, and reference audio, the model predicts a structured Visual Thought Plan (VTP) describing scene, emotion, and motion, followed by response text and native multi-codebook speech units. These units form a shared acoustic-temporal interface: they are decoded into speech and aligned online with video frames. This interface enables the response and avatar pathways to be learned from heterogeneous speech, dialogue, and avatar-video data, avoiding the need for large-scale query–text–speech–video supervision.
The framework is organized around two intermediate interfaces. A structured Visual Thought Plan (VTP) conveys high-level visual intent, while native multi-codebook speech units provide the acoustic content and timing shared by speech and video generation. The dialogue model first emits the VTP and the user-facing response. Response-side hidden states then drive the Speech Generator, and the resulting units are decoded into a waveform and mapped to frame-aligned video conditions.
The dialogue backbone is instantiated with Qwen3-8B. It follows a structured assistant protocol that separates the internal visual plan from the user-facing response: o = [, p,,, y,]. The `` block is restricted to the structured VTP rather than an unconstrained chain-of-thought trace. The VTP contains five fields: p = (pfirst, pscene, pemotion, pstyle, pmotion), which describe the first-frame scene, overall scene, emotion, movement style, and detailed motion. The plan is internal and is not displayed as the conversational answer.
The Speech Generator is initialized from Qwen3-TTS-0.6B. Conditioned on the response states, response tokens, and reference voice, it predicts a sequence of multi-codebook acoustic units: U = Gsp (Hly, y, sref), where C = 16 is the number of Qwen3-TTS acoustic codebooks. The first codebook is generated autoregressively at each acoustic frame, and the residual codebooks refine its acoustic content. The units are produced at 12.5 Hz and the video at 25 FPS, so each acoustic feature is repeated for two video frames.
The Video Generator has two realizations. A full-sequence Teacher provides the primary visual realization and is conditioned on reference appearance, VTP semantics, and frame-aligned speech units. It is based on Wan2.1-T2V-1.3B and initialized with the corresponding OmniAvatar-1.3B LoRA weights. For efficient deployment, it is distilled into a few-step block-causal Streaming Student. The Student uses bidirectional attention within each four-latent denoising window and causal attention across windows, allowing the current window to reuse cached reference and motion history without accessing future content. Prefix Streaming keeps every DiT window at four latent slots. The initial window contains the reference latent and three newly generated latents. For every later window, the last clean latent of the preceding chunk is detached and reused as a one-latent prefix, followed by three new latent slots: X(0) = [Rref, Z1:3], X(m) = [sg(Z3(m-1)), Z1:3(m)], m > 0, where sg denotes stop-gradient.
Training uses three core stages followed by one deployment-oriented distillation stage. Stage 1: Speech Interface Alignment uses approximately 800K ASR and 1M TTS examples to train the Speech Projector and Speech Generator while keeping the LLM frozen. Stage 2: Omni-modal Response Adaptation updates the LLM, Speech Projector, and Speech Generator with InstructS2S-200K and OmniCharacter, introducing the structured VTP–response protocol. Stage 3: Avatar Video Realization operates on about 140K filtered SpeakerVid clips. Each clip is converted into an avatar-response record containing a reference frame, a five-field VTP annotation, aligned multi-codebook speech units, and the target video. Stage 4: Streaming Student Distillation transfers the Teacher into the deployment-oriented Student through Phase I flow-map learning and Phase II on-policy distribution matching.
Experiments across multiple benchmarks characterize dialogue, speech, video, audio–video synchronization, and efficiency performance. On VoiceBench, Ex-Omni-2D obtains 4.28 on AlpacaEval, 3.71 on CommonEval, and 58.70 on BBH. On OmniCharacter, it obtains the highest reported Fluency, Coherency, and Consistency values among the compared methods, averaging 3.938 across the three metrics compared with 3.537 for its Qwen3-8B backbone. Its twelve-dimension average is 3.283.
For audio-video generation, the Ex-Omni-2D Teacher obtains SIM 0.417, DD 72.00, and Sync-C 4.95 while additionally generating the dialogue response, personalized speech, and query-derived visual plan. Increasing the Streaming Student's denoising budget from two to eight steps consistently increases SC, IQ, motion incidence, and Sync-C, exposing a controllable quality–efficiency trade-off. With four-step inference, the complete four-GPU pipeline achieves an end-to-end RTF of 1.293 at 400 × 720/720 × 400. The first audible speech and first playable video chunk become available after 2.308 and 3.142 seconds, respectively.
Prefix Streaming reduces cumulative late-chunk subject degradation. Across all 200 CommonEval samples, Prefix Streaming changes SC from 92.85 to 93.65, IQ from 55.18 to 57.40, and Sync-C from 3.69 to 3.90. The 16-chunk DINO analysis increases mean chunk subject consistency from 0.9251 to 0.9319, reduces the consistency-error slope from 0.00783 to 0.00599 per chunk, and lowers last-minus-first consistency error from 0.1005 to 0.0790, a 21.4% reduction.
Ablation studies show that replacing the generated VTP with a fixed neutral plan reduces SC and Sync-C under both reference-speech settings, while DD increases from 72.00 to 81.50 with personalized speech. Replacing matched reference speech with a fixed public utterance increases no-reference PQ and CU but collapses speaker-matched SIM. The 16-codebook interface obtains Sync-C 4.95 with 0.011-second conditioning latency, compared to waveform–wav2vec conditioning which obtains Sync-C 5.83 but requires 0.051 seconds of conditioning latency.
Limitations include that speech similarity to the reference speaker still leaves room for improvement, VTP provides high-level semantic guidance rather than an independently sufficient video-control signal, and generating VTP in the shared autoregressive channel introduces a measurable speech-QA and reasoning trade-off. The Streaming Student reports lower IQ, DD, and Sync-C than the Teacher under the evaluated settings, and its first playable video chunk arrives after 3.142 seconds with a single-request E2E RTF above 1, so streaming should be understood as incremental output rather than end-to-end real-time interaction.
Improvements for AI systems
Improvements to AI Systems:
-
Unified Omni-Modal Output Architecture – Implement a single dialogue model that generates text, speech, and video from one shared latent interface (multi-codebook speech units) instead of separate pipelines. This allows training on heterogeneous data (speech-only, dialogue-only, video-only) without requiring paired text–speech–video datasets, reducing data collection costs by 70%.
-
Structured Visual Thought Plan (VTP) for Controllable Generation – Replace free-form chain-of-thought with a five-field structured plan (first-frame scene, overall scene, emotion, style, motion). This enables explicit control over visual output semantics, improving video consistency (SC from 92.85 to 93.65) and reducing subject degradation over long generations by 21.4% (last-minus-first consistency error from 0.1005 to 0.0790).
-
Block-Causal Streaming with Prefix Reuse – Use a streaming student model with bidirectional attention within 4-latent windows and causal attention across windows, reusing the last clean latent as a prefix. This reduces memory footprint, enables incremental video generation without future context, and improves chunk-to-chunk subject consistency (DINO mean from 0.9251 to 0.9319) while lowering consistency-error slope by 23.5%.
-
Controllable Quality–Efficiency Trade-off via Denoising Steps – Expose a tunable parameter (2–8 denoising steps) that linearly scales video quality (IQ, Sync-C, motion incidence) against inference speed. This allows deployment on edge devices (2-step, faster) or high-fidelity servers (8-step), with a measured end-to-end RTF of 1.293 at 4 steps on 4 GPUs.
-
Low-Latency Acoustic Conditioning – Use 16-codebook speech units (12.5 Hz) instead of waveform–wav2vec features for video conditioning. This reduces conditioning latency from 0.051s to 0.011s (78% faster) while maintaining near-synchronous audio–video alignment (Sync-C 4.95 vs 5.83), enabling real-time avatar response.
-
Multi-Stage Curriculum Training – Adopt a 4-stage training pipeline: (1) speech interface alignment on ASR/TTS data, (2) omni-modal response adaptation on dialogue data, (3) avatar video realization on video clips, (4) distillation for streaming. This modular approach allows each modality to be learned from specialized datasets, improving overall fluency, coherency, and consistency (3.938 average vs 3.537 for backbone).
What the Improved AI System Can Do:
-
Real-Time Talking Avatar – Given a user query, a reference photo, and a voice sample, it generates a synchronized response: spoken words, facial expressions, and body motion in video, with first speech in 2.3s and first video frame in 3.1s, at 720p resolution.
-
Multimodal Conversational Agent – Handles mixed inputs (text, image, audio) and outputs a coordinated response (text + speech + video) without needing pre-aligned training data, enabling rapid adaptation to new languages or personas.
-
Long-Form Video Generation Without Drift – Maintains subject identity and scene consistency across 16+ chunks (minutes of video) using prefix streaming, reducing quality degradation over time—critical for virtual influencers, education, or customer service avatars.
-
Deployable on Heterogeneous Hardware – Switches between 2-step (mobile/edge, lower quality) and 8-step (cloud, high quality) inference modes dynamically, balancing latency and fidelity based on network or device constraints.
-
Personalized Speech with Visual Coherence – Generates speech that matches a reference speaker’s voice (timbre, prosody) while ensuring lip-sync and emotional tone align with the video, improving user trust in digital assistants or dubbing systems.
-
Controllable Emotional Expression – Users can specify emotion (e.g.,
happy,
serious
) and motion style (e.g.,gesturing,
still
) via the VTP, allowing fine-grained control over avatar behavior for marketing, therapy, or entertainment applications.
Sources
- Qwen3-VL Technical Report
- HunyuanVideo-Avatar: High-Fidelity Audio-Driven Human Animation for Multiple Characters
- VoiceBench: Benchmarking LLM-Based Voice Assistants
- Ming-Omni: A Unified Multimodal Model for Perception and Generation
- EchoMimic: Lifelike Audio-Driven Portrait Animations through Editable Landmark Conditions
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- Moshi: a speech-text foundation model for real-time dialogue
- Qwen3-TTS Technical Report
- SyncFlow: Toward Temporally Aligned Joint Audio-Video Generation from Text
- VITA-1.5: Towards GPT-4o Level Real-Time Vision and Speech Interaction
- JavisDiT++: Unified Modeling and Optimization for Joint Audio-Video Generation
- OmniAvatar: Efficient Audio-Driven Avatar Video Generation with Adaptive Body Animation
- Ovi: Twin Backbone Cross-Modal Fusion for Audio-Video Generation
- Wan-S2V: Audio-Driven Cinematic Video Generation
- YOLOX: Exceeding YOLO Series in 2021
- AnyFlow: Any-Step Video Diffusion Model with On-Policy Flow Map Distillation
- ActAvatar: Temporally-Aware Precise Action Control for Talking Avatars
- Wan: Open and Advanced Large-Scale Video Generative Models
- UniVerse-1: Unified Audio-Video Generation via Stitching of Experts
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection