Controlling Speaking Rate in Autoregressive TTS via Activation Steering
cs.SD, cs.CL, cs.LG, eess.AS
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/SYSTRAN/faster-whisper
Terminology
Sources
- Seed-TTS: A Family of High-Quality Versatile Speech Generation Models
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- Qwen3-TTS Technical Report
- Representation Engineering: A Top-Down Approach to AI Transparency
- Steering Language Models With Activation Engineering
- TADA! Tuning Audio Diffusion Models through Activation Steering
- Overcoming State Inertia in Full-Duplex Spoken Language Models via Activation Steering
- Steering Where to Listen: Instruction-Based Activation Steering Redirects Temporal Attention in Large Audio-Language Models
- SALSA: Speech Aware LLM Adaptation via Learned Steering Activation Vectors
- Understanding Unreliability of Steering Vectors in Language Models: Geometric Predictors and the Limits of Linear Approximations
- Refusal in LLMs is an Affine Function
- Improving Steering Vectors by Targeting Sparse Autoencoder Features
- Sparse Autoencoders Make Audio Foundation Models more Explainable
- Sparse Autoencoders for Interpretable Emotion Control in Text-to-Speech
- Interpreting and Steering a Text-to-Speech Language Model with Sparse Autoencoders
- CoCoEmo: Composable and Controllable Human-Like Emotional TTS via Activation Steering
- EmoShift: Lightweight Activation Steering for Enhanced Emotion-Aware Speech Synthesis
- EmoSteer-TTS: Fine-Grained and Training-Free Emotion-Controllable Text-to-Speech via Activation Steering
- Activation Steering for Accent-Neutralized Zero-Shot Text-To-Speech
- Activation Steering for Accent Adaptation in Large Audio Language Models
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment