The Semantic Bottleneck: Leveraging Semantic Representations for Non-Invasive Speech Decoding
cs.CL, cs.LG
Submitted: 2026-09-09
Updated: 2026-10-07
Comments: 12 pages, 8 figures
License: http://creativecommons.org/licenses/by/4.0/
The gist: Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult.
Terminology
Abstract
Non-invasive speech decoding remains constrained by the low signal-to-noise ratio of neural recordings, which makes fine-grained reconstruction of phonemes or individual words difficult. Motivated by neuroscientific evidence that high-level semantic representations are distributed across cortical regions and evolve over slower temporal scales, we hypothesize that semantic content may provide a more suitable target for non-invasive decoding than low-level acoustic or lexical features. We introduce Brain2Semantics2Text, a method that reconstructs text through an intermediate semantic embedding space. Our model maps sentence-level MEG responses into a semantic manifold and then inverts the predicted embeddings into natural language. This semantic bottleneck enables recovery of high-level meaning without word-level alignment. We describe the core principles of the approach, its implementation, and the strategies used to mitigate the challenges of learning a reliable neural-to-semantic mapping. Finally, we compare against prior non-invasive Brain2Text methods and show improved sentence-level results.
Sources
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech Representations
- VICReg: Variance-Invariance-Covariance Regularization for Self-Supervised Learning
- Decoding individual words from non-invasive brain recordings across 723 participants
- SONAR: Sentence-Level Multimodal and Language-Agnostic Representations
- Harnessing the Universal Geometry of Embeddings
- Are EEG-to-Text Models Working?
- BrainECHO: Semantic Brain Signal Decoding through Vector-Quantized Spectrogram Reconstruction for Whisper-Enhanced Text Generation
- Manifold learning: what, how, and why
- Language Model Inversion
- MTEB: Massive Text Embedding Benchmark
- Learning Transferable Visual Models From Natural Language Supervision
- Semantic reconstruction of continuous language from MEG signals
- Brain-to-Text Benchmark '24: Lessons Learned
- NeuGPT: Unified multi-modal Neural GPT
- NeuSpeech: Decode Neural signal as Speech
- MAD: Multi-Alignment MEG-to-Text Decoding
- Sigmoid Loss for Language Image Pre-Training
- LibriBrain: Over 50 Hours of Within-Subject MEG to Improve Speech Decoding Methods at Scale
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering