Tokens Change, Structure Endures: Spectral Watermarking for Generated Speech
cs.SD, cs.CL, cs.CR, cs.LG, eess.AS
Submitted: 2026-09-27
Updated: 2026-09-27
Code: https://github.com/sarulab-speech/UTMOSv2
Project page: https://hwiora.github.io/redwing-demo
Terminology
Sources
- AudioLM: a Language Modeling Approach to Audio Generation
- WavMark: Watermarking for Audio Generation
- VoiceBench: Benchmarking LLM-Based Voice Assistants
- CRAW: Codec Robust Audio Watermarking
- Undetectable Watermarks for Language Models
- Simple and Controllable Music Generation
- High Fidelity Neural Audio Compression
- Moshi: a speech-text foundation model for real-time dialogue
- CosyVoice 3: Towards In-the-wild Speech Generation via Scaling-up and Post-training
- MOSS-TTS Technical Report
- Clarabel: An interior-point solver for conic programs with quadratic objectives
- Unbiased Watermark for Large Language Models
- Watermarking Autoregressive Image Generation
- On the Reliability of Watermarks for Large Language Models
- Robust Distortion-free Watermarks for Language Models
- High-Fidelity Audio Compression with Improved RVQGAN
- A Survey of Text Watermarking in the Era of Large Language Models
- AudioMarkBench: Benchmarking Robustness of Audio Watermarking
- Proactive Detection of Voice Cloning with Localized Watermarking
- Latent Watermarking of Audio Generative Models
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment