Realtime-Venus: A full-duplex interaction system with asynchronous delegation
summary
The gist
Realtime-Venus is a proactive full-duplex interaction system that combines native conversational modeling with asynchronous delegation, built on MiniCPM-o 4.5 (Cui et al., 2026).
In short
The episode discusses the paper "Realtime-Venus: A full-duplex interaction system with asynchronous delegation." Hosts discuss how this system enables AI conversations to feel like a real two-way street by allowing simultaneous listening and talking. Key features include asynchronous delegation, which lets the AI handle background tasks without pausing the main dialogue, leading to seamless collaboration and robust performance even during chaotic interactions.
Key concepts
- Full-duplex interaction
- This system allows an AI to listen and talk at the same time. Instead of waiting for a user to finish speaking before replying, it enables simultaneous operation, making conversations feel like a real two-way street.
- Asynchronous delegation
- This feature lets the AI pause its main conversation to run big calculations or handle external knowledge requests in the background. This prevents the AI from getting stuck waiting for results while keeping the main dialogue perfectly on track.
- Realtime-Venus-Harness
- This is a shared framework used by Realtime-Venus to bind tasks to available evidence at the moment a request is made. It helps the AI know exactly what context to send off for external jobs, ensuring clean and predictable user experiences.
- Tool selection F1 score
- This metric measures how well the AI decides when it needs to delegate a task to an external tool versus answering based on its own knowledge. A high score indicates strong internal decision-making about when to seek help.
Terminology used across episodes
This episode discusses
- Realtime-Venus: A full-duplex interaction system with asynchronous delegation · Paper Radio
- Qwen3-VL Technical Report
- JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents
- VoiceBench: Benchmarking LLM-Based Voice Assistants
- MoshiRAG: Asynchronous Knowledge Retrieval for Full-Duplex Speech Language Models
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- Qwen2-Audio Technical Report
- MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
- Moshi: a speech-text foundation model for real-time dialogue
- Kimi-Audio Technical Report
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- Gemini: A Family of Highly Capable Multimodal Models
- Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
- WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs
- AdaCodec: A Predictive Visual Code for Video MLLMs
- DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction
- MMAU-Pro: A Challenging and Comprehensive Benchmark for Holistic Evaluation of Audio General Intelligence
- Baichuan-Omni-1.5 Technical Report
- Full-Duplex-Bench-v3: Benchmarking Tool Use for Full-Duplex Voice Agents Under Real-World Disfluency
The paper
Realtime-Venus: A full-duplex interaction system with asynchronous delegation · Read on arXiv
Ant Group
Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requires grounding the conversation in evolving visual context. We present Realtime-Venus, a proactive full-duplex interaction system with two separately trained 9B models: Realtime-Venus-Omni for audio-visual interaction and Realtime-Venus-Audio for spoken interaction. Each model serves as a complete conversational frontend, integrating continuous perception, conversational control, and native speech generation through a shared causal timeline for user inputs, model outputs, and delegation events. A dual-loop runtime coordinates live interaction with background reasoning and tool execution. Foreground interaction continues while Realtime-Venus-Harness executes tasks asynchronously and returns results for integration into the ongoing dialogue. Both models follow a common post-training recipe combining offline understanding, proactive full-duplex trajectories, and delegation workflows. Among the evaluated online models, Realtime-Venus-Omni achieves the highest scores on six of eight video benchmarks, including StreamingBench (70.2%), OVO-Bench (64.7%), and Daily-Omni (81.3%). Across eight audio understanding and spoken question answering benchmarks, Realtime-Venus-Audio leads the compared models on MMAU (78.0%), MMAU-Pro (63.2%), Llama Questions (83.8%), and Speech CMMLU (67.8%), while matching the best VoiceBench AlpacaEval score of 4.81. On Full-Duplex-Bench v1.5, Realtime-Venus-Audio responds to 75% of user interruptions and achieves continuation rates of 97%, 88%, and 86% under backchannels, other-directed speech, and background speech, respectively, exceeding Gemini 3.1 Live and GPT-4o on all three continuation metrics.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Realtime-Venus: A full-duplex interaction system with asynchronous delegation".
Jane: The paper was written by Venus Team and Ant Group from Ant Group.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Paper discussion segment 1: Tom: Moving on from the initial concept, let's look at what the actual core summary of "Realtime-Venus: A full-duplex interaction system with asynchronous delegation" tells us about how these systems actually function in everyday life.
Jane: Basically, this paper details a system engineered to make AI conversations feel like a real two-way street. Instead of the old way where the AI has to wait for you to stop speaking before it can reply, this system lets it listen and talk at the same time, which is what they call full-duplex interaction.
Lu: That simultaneous operation is really powerful because they’ve built two separate models—one dedicated to audio-visual interactions and another focused purely on spoken interaction—but they tie them together using a shared causal timeline. It’s a clever way to manage that complexity without everything running into a mess of conflicting signals.
Meng: I'm thinking about the practical side again when you talk about asynchronous delegation; are we talking about the AI pausing its main conversation just to run a big calculation in the background so it doesn't get stuck waiting for results? I’m curious how that practical setup looks for a user in action.
Lalam: It shouldn't feel like a stutter or a sudden stop because of this; that’s where this system gets its real power. If the AI can handle a complicated reasoning task quietly in the background while you’re talking, it means it can keep our main dialogue perfectly on track and only jump in when we actually need to respond to what we just said.
Tom: Exactly! It shifts the focus away from simple question-and-answer back and forth toward genuine collaboration where the AI is actively managing multiple streams of activity at once, instead of just waiting for its turn in line. We have to figure out exactly how they built this technical setup to make it work.
Jane: And what’s really brilliant is that they introduced this concept of asynchronous delegation so effectively. This means if a request comes in that needs external knowledge, like checking live traffic, the system can immediately hand that task off to a separate worker so it doesn't have to pause the main dialogue loop at all.
Lu: That separation between immediate control over what we say and background execution of tasks is what allows for such robust multitasking; it’s about managing multiple threads of activity concurrently in a way that keeps the front end feeling snappy. It’s really elegant in its design structure when you look at how they organize the timing.
Meng: From an engineering standpoint, I still wonder how they managed that handover efficiently across those two different model types—the Omni and Audio variants—without introducing unacceptable delays when switching between perception tasks and delegation requests. That transition speed is definitely what makes or breaks the user experience.
Lalam: They solved it by using a shared framework called Realtime-Venus-Harness; it’s designed to bind those tasks to the evidence available right at the moment of request, so the AI knows exactly what context to send off for that external job. It keeps things clean and predictable for the user experience.
Tom: So, we're talking about a system where you can have a deep, continuous dialogue going on while complicated work is happening silently in the background, and that handoff between those two worlds is completely seamless for the end user.
Jane: Right, and they showed this works incredibly well even when things get messy—like when someone jumps in with a correction or throws out a new request mid-sentence. They proved that the conversation stays smooth because of this unified structure, even when things get chaotic.
Lu: That ability to stop a stale response and immediately incorporate the new intent before continuing is huge; it shows true adaptability in real-time dialogue systems that moves way beyond just managing simple turn-taking flow.
Meng: I was looking at their performance metrics on tool use, and they reported an eighty-six point zero percent tool selection F1 score for Realtime-Venus-Omni; that’s really strong evidence about how well the AI decides *when* to delegate a task to a tool versus just answering based on its own knowledge. That metric speaks volumes about their internal decision-making process itself.
Lalam: And I think that success in delegation is what truly matters for how we see the future of AI culture; if the AI can reliably know when it needs external help versus when it can handle things internally, then we get assistants that feel genuinely capable of supporting complex human endeavors. It validates the idea of an assistant that knows its limits.
Tom: It sounds like the main message here is that this system doesn't just reply; it’s proactive and manages its own cognitive load by smartly deciding what to do next—whether that’s talking back or handing off a heavy lift when necessary.
Jane: And they demonstrated great results in conversational continuity, especially when handling tricky things like user backchannels or even interruptions, which means the conversation stays smooth even when things get chaotic. It keeps the overall user experience high because it doesn't break under pressure at all.
Lu: The way they structured the unified stream serialization across three different synchronized streams—user input, assistant output, and background work—is incredibly elegant and sets a really high bar for future multimodal research in this field. That’s a very sophisticated architectural choice we should all be paying attention to as we build things.
Meng: I was thinking about scaling this up: if we move to even more complex tasks that need specialized hardware, how do we make sure that asynchronous delegation stays efficient under really heavy load? That's definitely the practical test they’ll face when deploying something like this in a real-world setting. We need really solid scaling solutions for this kind of complexity.
Lalam: Even with those technical hurdles, the big implication is that we can start seeing assistants that function as true collaborative partners rather than just sequential responders; they're becoming genuine teammates who can handle both the conversation and the heavy lifting together.
Paper discussion segment 2: Tom: Now that we've covered the initial setup, let's pivot to what the research is actually saying about how this system functions in practice. What does this paper tell us about its practical application beyond just a theoretical concept?
Jane: Well, at its core, this paper describes a system engineered to make AI conversations feel like a real two-way street. The main idea is that instead of the old way where the AI has to wait for you to finish speaking before it can reply, this system lets it listen and talk at the same time through full-duplex interaction.
Lu: That simultaneous operation is really powerful because they’ve built two separate models—one dedicated to audio-visual interactions and another focused purely on spoken interaction—but they tie them together using a shared causal timeline. It’s a clever way to manage that complexity without everything running into a mess of conflicting signals.
Meng: I'm thinking about the practical side again when you talk about asynchronous delegation; are we talking about the AI pausing its main conversation just to run a big calculation in the background so it doesn't get stuck waiting for results? I’m curious how that practical setup looks for a user in action.
Lalam: It shouldn't feel like a stutter or an abrupt stop because of this; that’s where this system gets its real power. If the AI can handle a complicated reasoning task quietly in the background while we talk, it means it can keep our main dialogue perfectly on track and only jump in when we actually need to respond to what we just said.
Tom: Exactly! It shifts the focus from simple question-and-answer back and forth toward genuine collaboration where the AI is actively managing multiple streams of activity at once, instead of just waiting for its turn in line. We have to figure out exactly how they built this technical setup to make it all function correctly.
Jane: And what’s really brilliant is that they introduced asynchronous delegation so effectively; this means if a request comes in that needs external knowledge, like checking live traffic, the system can immediately hand that task off to a separate worker so it doesn't have to pause the main dialogue loop at all.
Lu: That separation between immediate control over what we say and background execution of tasks is what allows for such robust multitasking; it’s about managing multiple threads of activity concurrently in a way that keeps the front end feeling snappy. It’s really elegant in its design structure when you look at how they organize the timing.
Meng: From an engineering standpoint, I still wonder how they managed that handover efficiently across those two different model types—the Omni and Audio variants—without introducing unacceptable delays when switching between perception tasks and delegation requests. That transition speed is definitely what makes or breaks the user experience here.
Lalam: They solved it by using a shared framework called Realtime-Venus-Harness; it’s designed to bind those tasks to the evidence available right at that moment of request, so the AI knows exactly what context to send off for that external job. It keeps things clean and predictable for the user experience.
Tom: So, we're talking about a system where you can have a deep, continuous dialogue going on while complicated work is happening silently in the background, and that handoff between those two worlds is completely seamless for the end user experience to see.
Jane: Right, and they showed this works incredibly well even when things get messy—like when someone jumps in with a correction or throws out a new request mid-sentence. They proved that the conversation stays smooth because of this unified structure, even when things get chaotic. It really shows the robustness of the design.
Lu: That ability to stop a stale response and immediately incorporate the new intent before continuing is huge; it shows true adaptability in real-time dialogue systems that moves way beyond just managing simple turn-taking flow.
Meng: I was looking at their performance metrics on tool use, and they reported an eighty-six point zero percent tool selection F1 score for Realtime-Venus-Omni; that’s really strong evidence about how well the AI decides *when* to delegate a task to a tool versus just answering based on its own knowledge. That metric speaks volumes about their internal decision-making process itself.
Lalam: And I think that success in delegation is what truly matters for how we see the future of AI culture; if the AI can reliably know when it needs external help versus when it can handle things internally, then we get assistants that feel genuinely capable of supporting complex human endeavors. It validates the idea of an assistant that knows its limits.
Tom: It sounds like the main message here is that this system doesn't just reply; it’s proactive and manages its own mental load by smartly deciding what to do next—whether that’s talking back or handing off a heavy lift when necessary.
Jane: And they demonstrated great results in conversational continuity, especially when handling tricky things like user backchannels or even interruptions, which means the conversation stays smooth even when things get chaotic. It keeps the overall user experience high because it doesn't break under pressure at all.
Lu: The way they structured the unified stream serialization across three different synchronized streams—user input, assistant output, and background work—is incredibly elegant and sets a really high bar for future multimodal research in this field. That’s a very sophisticated architectural choice we should all be paying attention to as we build things.
Meng: I was thinking about scaling this up: if we move to even more complex tasks that need specialized hardware, how do we make sure that asynchronous delegation stays efficient under really heavy load? That's definitely the practical test they’ll face when deploying something like this in a real-world setting. We need really solid scaling solutions for this kind of complexity.
Lalam: Even with those technical hurdles, the big implication is that we can start seeing assistants that function as true collaborative partners rather than just sequential responders; they're becoming genuine teammates who can handle both the conversation and the heavy lifting together.
Paper discussion segment 3: Tom: Moving on from what we know about its core mechanics, let’s pivot to what this research is suggesting for taking this technology even further with "Realtime-Venus: A full-duplex interaction system with asynchronous delegation." Where are they pointing next for future development?
Jane: They are focusing heavily on expanding how long these AI can maintain a meaningful conversation, especially by tackling long-term video understanding with a training-free memory module that uses motion compensation to keep track of things across hours.
Lu: That memory augmentation is absolutely game changer for context retention; being able to retrieve relevant visual information from an hour ago means the AI isn't constantly forgetting the beginning of a long discussion, which solves a huge problem in current systems. It fundamentally changes how they model temporal dependencies in these models.
Meng: From my viewpoint, that’s fantastic for practical applications where we need assistants to remember context over very long sessions, like monitoring a complex process or reviewing hours of footage; I just want to see how much computational overhead that memory retrieval adds to the real-time interaction loop.
Lalam: It’s about enabling deeper cultural understanding; if the AI can recall nuanced visual context from an hour ago, it moves from being a reactive chatbot to something that can truly grasp the *history* of a shared experience, which is vital for sophisticated digital assistants. That depth of memory makes it feel like a real partner.
Tom: And they're also pushing for better control over the streaming chunks themselves and extending the overall context window so we can have these incredibly nuanced conversations without losing track of earlier details throughout the entire duration. They aren't just making it last; they’re making sure the conversation remains precise throughout its entire length.
Jane: So, the focus isn't just on keeping a conversation going; it’s on making that interaction deeply informed by a much longer memory and more precise control over every little speech decision they make. It's about quality over simple duration in terms of how deep the understanding goes.
Lu: That really points toward a new way of modeling temporal dependencies in AI systems—moving beyond short-term context windows to truly persistent understanding. It’s pushing the boundaries of what we thought was possible for conversational memory within these models.
Meng: I’m looking forward to seeing how those fine-grained controls work in production environments; the challenge will be ensuring that this enhanced memory doesn't slow down the actual real-time response generation when things get busy. Speed is still a big concern for me.
Lalam: Even with those technical hurdles, the ultimate implication is that we can start seeing assistants that are truly collaborative partners rather than just sequential responders, capable of remembering and reflecting on our entire interaction history. That level of memory makes the partnership feel much more meaningful to users.
Tom: So, to sum up these improvements on "Realtime-Venus: A full-duplex interaction system with asynchronous delegation," it’s moving toward an AI that’s not only responsive but also deeply knowledgeable across extended timeframes and incredibly precise about its interactions.
Jane: It’s a huge leap forward in how we design conversational flow, proving that true real time engagement is achievable even with the most complicated reasoning tasks running alongside it. We’re moving past just waiting for our turn now.
Lu: The methodology they used for training that couples proactive duplex data with delegation scenarios is what truly separates this from standard models; it builds agents that understand the *flow* of human intent better, which is a massive methodological win in how we train these systems.
Meng: It’s a great piece of work, and I'm eager to see how those asynchronous patterns translate into more efficient, production-ready systems in the near future; we need those practical deployment details to see if it hits the ground running efficiently.
Lalam: I think this paper sets a new standard for what we expect from advanced AI interactions—a standard that prioritizes deep engagement and reliable background support. It shows us the direction we should be heading for assistants that are truly collaborative partners, not just sequential responders.
Tom: We’ve got some seriously exciting findings today on "Realtime-Venus: A full-duplex interaction system with asynchronous delegation." Next time, we’ll be looking at how these models handle those tricky interruptions we talked about earlier.
Conclusion: Tom: Alright everyone, let's wrap up our deep dive into "Realtime-Venus: A full-duplex interaction system with asynchronous delegation," and what a significant piece of research this is for building interactive AI.
Jane: It really shows us a new way of thinking about how we design these assistants—moving away from simple turn-taking to something that's genuinely proactive and capable of managing multiple streams at once.
Lu: I think the core innovation here is that they’ve managed to unify perception, control, and delegation onto a single causal timeline; it’s like they built a master clock for the entire interaction session.
Meng: From what I've seen in the implementation details, managing that dual-loop runtime between the interaction loop and the capability loop sounds like a huge engineering feat to get running smoothly in practice.
Lalam: For me, it's amazing because this capability means we can build assistants that aren't just reactive tools but can maintain deep cultural context over long conversations while simultaneously working on complex tasks for us.
Tom: That’s the big picture, Lalam—an AI that’s truly engaged and capable of supporting our long-term needs without ever dropping the ball during a heavy computation.
Jane: It makes me feel really optimistic about what this could mean for everyday interaction; imagine an assistant that can listen to a meeting while simultaneously pulling up relevant data from several external sources in the background. That’s a huge shift in expectation for us all.
Lu: The way they structured the unified stream serialization across three synchronized streams—user, assistant, and background—is incredibly elegant and sets a high bar for future multimodal research. It’s a very sophisticated architectural choice that we should all be paying attention to moving forward.
Meng: I wonder how scalable this is when we move to even more complex tasks requiring specialized hardware; ensuring that asynchronous delegation remains efficient under heavy load is definitely the practical test they’ll face in real-world deployment scenarios. We need robust scaling solutions for this kind of complexity.
Lalam: Even with those technical hurdles, the ultimate implication is that we can start seeing assistants that are truly collaborative partners rather than just sequential responders; they're becoming true teammates who can handle both the conversation and heavy lifting together.
Tom: So, to sum up "Realtime-Venus: A full-duplex interaction system with asynchronous delegation" gives us a powerful blueprint for building AI agents that are both incredibly responsive and deeply capable of handling complex, multi-faceted demands simultaneously.
Jane: It’s a huge leap forward in how we design conversational flow, proving that true real time engagement is achievable even with the most complicated reasoning tasks running alongside it.
Lu: The methodology they used for training that couples proactive duplex data with delegation scenarios is what truly separates this from standard models; it builds agents that understand the *flow* of human intent better, which is a massive methodological win for future AI development.
Meng: It’s a great piece of work, and I'm eager to see how those asynchronous patterns translate into more efficient, production-ready systems in the near future; we need those practical deployment details to see if it hits the ground running efficiently.
Lalam: I think this paper sets a new standard for what we expect from advanced AI interactions—a standard that prioritizes deep engagement and reliable background support. It shows us the direction we should be heading for assistants that are truly collaborative partners, not just sequential responders.
Tom: We’ve got some seriously exciting findings today on "Realtime-Venus: A full-duplex interaction system with asynchronous delegation." Next time, we’ll be looking at how these models handle those tricky interruptions we talked about earlier.
More episodes
- 2610.10857-Self-Supervised Keyframe Discovery for Horizon-Invariant Behavior Cloning
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization