OmniFusion: Simultaneous Multilingual Multimodal Translations via Modular Fusion
cs.CL, cs.AI
Submitted: 2025-11-28
Updated: 2026-08-28
Terminology
Sources
- SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- No Language Left Behind: Scaling Human-Centered Machine Translation
- PaddleOCR 3.0 Technical Report
- Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture
- StreamUni: Achieving Streaming Speech Translation with a Unified Large Speech-Language Model
- Discrete Tokenization for Multimodal LLMs: A Comprehensive Survey
- MCIF: Multimodal Crosslingual Instruction-Following Benchmark from Scientific Talks
- Gemma 3 Technical Report
- How to Connect Speech Foundation Models and Large Language Models? What Matters and What Does Not
- CoVoST 2 and Massively Multilingual Speech-to-Text Translation
- OmniVinci: Enhancing Architecture and Data for Omni-Modal Understanding LLM
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering