Repetition, Not Length: Isolating the Counting Failure in Neural Text-to-Speech
cs.CL, cs.SD
Submitted: 2026-09-29
Updated: 2026-09-29
Code: https://github.com/lab260ru/tts-counting-failure
Terminology
Sources
- Neural Codec Language Models are Zero-Shot Text to Speech Synthesizers
- RALL-E: Robust Codec Language Modeling with Chain-of-Thought Prompting for Text-to-Speech Synthesis
- CosyVoice 2: Scalable Streaming Speech Synthesis with Large Language Models
- VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
- From Dispersion to Attraction: Spectral Dynamics of Hallucination Across Whisper Model Scales
- Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
- Qwen3-TTS Technical Report
- Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs
- Transformers need glasses! Information over-squashing in language tasks
- When Can Transformers Count to n?
- The Counting Power of Transformers
- Counting Ability of Large Language Models and Impact of Tokenization
- Theoretical Limitations of Self-Attention in Neural Sequence Models
- What Formal Languages Can Transformers Express? A Survey
- (How) Do Language Models Track State?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering