Qwen3.8-Omni: Towards Native Omni-Modal Agents
cs.CL, cs.CV, cs.MM
Submitted: 2026-09-22
Updated: 2026-09-22
Code: https://github.com/QwenLM/Qwen-MM-Plugins
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce Qwen3.8-Omni-Flash, a natively multimodal agentic model for real-world multimodal productivity.
Terminology
Abstract
We introduce Qwen3.8-Omni-Flash, a natively multimodal agentic model for real-world multimodal productivity. Compared with previous omni models, which primarily emphasized perception and interaction, Qwen3.8-Omni-Flash substantially improves multimodal understanding and reasoning, as well as performance on long-horizon agentic tasks. These capabilities are supported by a native multimodal co-training strategy that preserves strong text-domain capabilities while facilitating the transfer of agentic capabilities from text to audio and video tasks. The model inherits the sparse mixture-of-experts (MoE) architecture of Qwen3.8-Next and extends the context window to one million tokens, supporting long-context multimodal reasoning and long-horizon planning. These advances enable integration into production workflows as a primary agent or a specialized sub-agent, supporting video editing, long-form audio and video translation, music-conditioned music video or movie generation, and video-based note or omni-skill creation. To address the lack of native audio and video support in existing agent harnesses, we release Qwen-MM-Plugins, a lightweight open-source plugin framework for multimodal productivity. We further frame real-time multimodal interaction as a system-level challenge requiring orchestration of context and memory management, tool use, and sub-agent delegation. Accordingly, we release Qwen-Live-Harness, an open-source framework for building responsive, real-time multimodal agents based on Qwen3.8-Omni-Flash. Extensive evaluations demonstrate that Qwen3.8-Omni-Flash achieves strong performance across multimodal understanding, reasoning, long-horizon agentic execution, and video productivity tasks. These results and the accompanying open-source tools support Qwen3.8-Omni-Flash as a practical foundation for deploying natively multimodal agents in research and production.
Sources
- AgenticVBench: Can AI Agents Complete Real-World Post-Production Tasks?
- JointAVBench: A Benchmark for Joint Audio-Visual Reasoning Evaluation
- Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction
- VoiceBench: Benchmarking LLM-Based Voice Assistants
- UniClawBench: A Universal Benchmark for Proactive Agents on Real-World Tasks
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?
- NL2Repo-Bench: Towards Long-Horizon Repository Generation Evaluation of Coding Agents
- WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
- Video-MME-v2: Towards the Next Stage in Benchmarks for Comprehensive Video Understanding
- MMOU: A Massive Multi-Task Omni Understanding and Reasoning Benchmark for Long and Complex Real-World Videos
- OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue
- Vision2Web: A Hierarchical Benchmark for Visual Website Development with Agent Verification
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
- MusTBench: Benchmarking and Advancing Temporal Grounding in Music LLMs
- OmniVideoBench: Towards Audio-Visual Understanding Evaluation for Omni MLLMs
- OmniGAIA: Towards Native Omni-Modal AI Agents
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering