Diffusion Large Language Models for Visual Speech Recognition
cs.AI, cs.CV, eess.AS
Submitted: 2026-05-27
Updated: 2026-09-01
Comments: Accepted to EMNLP 2026. Code: https://github.com/JeongHun0716/dllm-vsr
Code: https://github.com/JeongHun0716/dllm-vsr
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- LRS3-TED: a large-scale dataset for visual speech recognition
- LipNet: End-to-End Sentence-level Lipreading
- Improving Variable-Length Generation in Diffusion Language Models via Length Regularization
- The Llama 3 Herd of Models
- Transformer-Based Video Front-Ends for Audio-Visual Speech Recognition for Single and Multi-Person Video
- Learning Audio-Visual Speech Representation by Masked Multimodal Cluster Prediction
- Dream-Coder 7B: An Open Diffusion Language Model for Code
- Qwen2.5 Technical Report
- Diffusion LLM with Native Variable Generation Lengths: Let [EOS] Lead the Way
- Dream 7B: Diffusion Large Language Models
- LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
- Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection