When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models
cs.CL, cs.SD, eess.AS
Submitted: 2026-08-31
Updated: 2026-08-31
Code: https://github.com/sizigi/animeGRPO.1
Project page: https://sizigi.github.io
Terminology
Sources
- DLPO: Diffusion Model Loss-Guided Reinforcement Learning for Fine-Tuning Text-to-Speech Diffusion Models
- High Fidelity Neural Audio Compression
- CosyVoice: A Scalable Multilingual Zero-shot Text-to-speech Synthesizer based on Supervised Semantic Tokens
- Differentiable Reward Optimization for LLM based TTS system
- Emilia: An Extensive, Multilingual, and Diverse Speech Dataset for Large-Scale Speech Generation
- Who Finds This Voice Attractive? A Large-Scale Experiment Using In-the-Wild Data
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- Group Relative Policy Optimization for Text-to-Speech with Large Language Models
- Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
- AnimeScore: A Preference-Based Dataset and Framework for Evaluating Anime-Like Speech Style
- MLS: A Large-Scale Multilingual Dataset for Speech Research
- Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale
- Fine-Tuning Language Models from Human Preferences
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering