FlowAct-R2: Beyond Talking Avatar via Streaming Multimodal References and Proactive Agent Planning
cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Project page: https://bone-11.github.io/Flowact-R2
Terminology
Sources
- Wan-Streamer v0.1: End-to-end Real-time Interactive Foundation Models
- DuplexOmni: Real-Time Listening, Seeing, Thinking, and Speaking for Full-Duplex Interaction
- Live Avatar: Streaming Real-time Audio-Driven Avatar Generation with Infinite Length
- StreamHOI: Interaction-aware Temporal Memory Adaptation for Streaming HOI Video Generation
- FlowAct-R1: Towards Interactive Humanoid Video Generation
- X-Streamer: Unified Human World Modeling with Audiovisual Interaction
- Vidu S2: Real-Time Interactive, Editable, and Spatial Video Generation
- Vidu S1: A Real-Time Interactive Video Generation Model
- LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time
- INFP: Audio-Driven Interactive Head Generation in Dyadic Conversations
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models