Reconstruct! Don't Encode: Self-Supervised Representation Reconstruction Loss for High-Intelligibility and Low-Latency Streaming Neural Audio Codec
eess.AS, cs.AI
Submitted: 2026-03-06
Updated: 2026-09-01
Code: https://github.com/jhcodec843/jhcodec
Terminology
Sources
- Moshi: a speech-text foundation model for real-time dialogue
- High Fidelity Neural Audio Compression
- WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling
- MagiCodec: Simple Masked Gaussian-Injected Codec for High-Fidelity Reconstruction and Generation
- Scaling Transformers for Low-Bitrate High-Quality Speech Coding
- FocalCodec-Stream: Streaming Low-Bitrate Speech Coding via Causal Distillation
- SpeechTokenizer: Unified Speech Tokenizer for Speech Large Language Models
- SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
- BigCodec: Pushing the Limits of Low-Bitrate Neural Speech Codec
- XY-Tokenizer: Mitigating the Semantic-Acoustic Conflict in Low-Bitrate Speech Codecs
- TS3-Codec: Transformer-Based Simple Streaming Single Codec
- High-Fidelity Simultaneous Speech-To-Speech Translation
- When Worse is Better: Navigating the compression-generation tradeoff in visual tokenization
- Language Model Beats Diffusion -- Tokenizer is Key to Visual Generation
- Heptapod: Language Modeling on Visual Signals
- GLU Variants Improve Transformer
- Layer Normalization
- Effective Context in Neural Speech Models
- Llasa: Scaling Train-Time and Inference-Time Compute for Llama-based Speech Synthesis
- Vector-quantized Image Modeling with Improved VQGAN
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions