Spoken Language Models that Think Aloud
cs.CL, cs.SD, eess.AS
Submitted: 2026-09-22
Updated: 2026-09-22
Comments: Accepted at SLT 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- MinMo: A Multimodal Large Language Model for Seamless Voice Interaction
- STITCH: Simultaneous Thinking and Talking with Chunked Reasoning for Spoken Language Models
- Qwen2-Audio Technical Report
- Moshi: a speech-text foundation model for real-time dialogue
- Kimi-Audio Technical Report
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- LLaMA-Omni: Seamless Speech Interaction with Large Language Models
- LLaMA-Omni2: LLM-based Real-time Spoken Chatbot with Autoregressive Streaming Speech Synthesis
- LUCY: Linguistic Understanding and Control Yielding Early Stage of Her
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- GPT-4o System Card
- OpenAI o1 System Card
- TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension
- Reinforcement Learning Outperforms Supervised Fine-Tuning: A Case Study on Audio Question Answering
- Baichuan-Audio: A Unified Framework for End-to-End Speech Interaction
- VITA-Audio: Fast Interleaved Cross-Modal Token Generation for Efficient Large Speech-Language Model
- Can Speech LLMs Think while Listening?
- Freeze-Omni: A Smart and Low Latency Speech-to-speech Dialogue Model with Frozen LLM
- Towards Spoken Mathematical Reasoning: Benchmarking Speech-based Models over Multi-faceted Math Problems
- SARI: Structured Audio Reasoning via Curriculum-Guided Reinforcement Learning
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering