Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions
cs.CL, cs.AI
Submitted: 2026-08-31
Updated: 2026-08-31
Terminology
Sources
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- Reasoning Beyond Language: A Comprehensive Survey on Latent Chain-of-Thought Reasoning
- Earnings-22: A Practical Benchmark for Accents in the Wild
- How Contrastive Decoding Enhances Large Audio Language Models
- Granite-speech: open-source speech-aware LLMs with strong English ASR capabilities
- Qwen3-ASR Technical Report
- Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism
- Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters
- Open ASR Leaderboard: Towards Reproducible and Transparent Multilingual and Long-Form Speech Recognition Evaluation
- Speech-Hands: A Self-Reflection Voice Agentic Approach to Speech Recognition and Audio Reasoning with Omni Perception
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering