EgoVoice: Proactive Spoken Assistance from Egocentric Multimodal Streams
cs.CL, cs.CV, cs.SD
Submitted: 2026-10-08
Updated: 2026-10-08
Code: https://github.com/snakers4/silero-vad
Project page: https://egocentricvoice.github.io
Terminology
Sources
- MiniCPM-o 4.5: Towards Real-Time Full-Duplex Omni-Modal Interaction
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- Moshi: a speech-text foundation model for real-time dialogue
- AURA: Always-On Understanding and Real-Time Assistance via Video Streams
- GPT-4o System Card
- OpenAI GPT-5 System Card
- PersonaPlex: Voice and Role Control for Full Duplex Conversational Speech Models
- SAM Audio: Segment Anything in Audio
- Qwen3-ASR Technical Report
- Mini-Omni: Language Models Can Hear, Talk While Thinking in Streaming
- Mini-Omni2: Towards Open-source GPT-4o with Vision, Speech and Duplex Capabilities
- Qwen3-Omni Technical Report
- Proact-VL: A Proactive VideoLLM for Real-Time AI Companions
- GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
- Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering