JoyAI-VL-Interaction: Real-Time Vision-Language Interaction Intelligence
cs.CV, cs.AI
Submitted: 2026-06-10
Updated: 2026-09-24
Code: https://github.com/jd-opensource/JoyAI-VL-Interaction
Project page: https://joyai-vl-video-future-academy-jd.github.io/JoyAI-VL-Interaction
Terminology
Sources
- Qwen3-VL Technical Report
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- Video Streaming Thinking: VideoLLMs Can Watch and Think Simultaneously
- Efficient Memory Management for Large Language Model Serving with PagedAttention
- OVO-Bench: How Far is Your Video-LLMs from Real-World Online Video Understanding?
- StreamingBench: Assessing the Gap for MLLMs to Achieve Streaming Video Understanding
- Dispider: Enabling Video LLMs with Active Real-Time Interaction via Disentangled Perception, Decision, and Reaction
- EasyVideoR1: Easier RL for Video Understanding
- Qwen3.5-Omni Technical Report
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- VideoLLM Knows When to Speak: Enhancing Time-Sensitive Video Comprehension with Video-Text Duet Interaction Format
- Streaming Video Instruction Tuning
- StreamingVLM: Real-Time Understanding for Infinite Video Streams
- Proact-VL: A Proactive VideoLLM for Real-Time AI Companions
- Qwen3 Technical Report
- HERMES: KV Cache as Hierarchical Memory for Efficient Streaming Video Understanding
- SGLang: Efficient Execution of Structured Language Model Programs
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models