AURAL: Adaptive Latent Reasoning with Joint Chunk for Speech Language Models
cs.CL, cs.LG, cs.SD
Submitted: 2026-10-01
Updated: 2026-10-01
Terminology
Sources
- Playing a Part: Speaker Verification at the Movies
- The MSP-Podcast Corpus
- GigaSpeech: An Evolving, Multi-domain ASR Corpus with 10,000 Hours of Transcribed Audio
- Compressed Chain of Thought: Efficient Reasoning Through Dense Representations
- Training Verifiers to Solve Math Word Problems
- LatentOmni: Rethinking Omni-Modal Understanding via Unified Audio-Visual Latent Reasoning
- LLM Latent Reasoning as Chain of Superposition
- Latent-GRPO: Group Relative Policy Optimization for Latent Reasoning
- Continuous Audio Thinking for Large Audio Language Models
- Don't Overthink it. Preferring Shorter Thinking Chains for Improved LLM Reasoning
- Measuring Mathematical Problem Solving With the MATH Dataset
- Qwen3-TTS Technical Report
- Kimi-Audio Technical Report
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- EXPRESSO: A Benchmark and Analysis of Discrete Expressive Speech Resynthesis
- EARS: An Anechoic Fullband Speech Dataset Benchmarked for Speech Enhancement and Dereverberation
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- EmotionTalk: An Interactive Chinese Multimodal Emotion Dataset With Rich Annotations
- Step-Audio-R1 Technical Report
- Marco-Voice Technical Report
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering