Using Prosody to Predict Syntactic Structure
cs.CL, cs.AI, cs.LG
Submitted: 2026-08-31
Updated: 2026-08-31
Comments: 15 pages, 4 figures. EMNLP 2026 camera-ready
Code: https://github.com/aatlantise/prosody-syntax-interface
License: http://creativecommons.org/licenses/by/4.0/
The gist: While it is well-established that prosody carries crucial cues for syntactic structure, the degree and nature of correspondence between these two domains remains contested.
Terminology
Abstract
While it is well-established that prosody carries crucial cues for syntactic structure, the degree and nature of correspondence between these two domains remains contested. We investigate the syntax-prosody interface through an information-theoretic lens, quantifying the interaction between prosodic features and syntactic representations as their mutual information. We provide a general-purpose framework for estimating this quantity over large speech-text corpora using multimodal language models. Our framework is structure-agnostic and modular, insofar as it can be used to measure the contributions of individual prosodic features or components of structure. We evaluate the syntax-prosody relationship for two features (word duration and inter-word pauses) across two domains--read audiobooks and spontaneous conversations--both in English. Our results demonstrate that prosody contains measurable syntactic information, with prosodic features reducing syntactic uncertainty in spontaneous conversations by up to 10.2%. Our findings offer new empirical support for several theoretical accounts of the syntax-prosody interface.
Sources
- Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
- PaliGemma: A versatile 3B VLM for transfer
- On entropy for mixtures of discrete and continuous variables
- What Do Prosody and Text Convey? Characterizing How Meaningful Information is Distributed Across Multiple Channels
- LibriTTS: A Corpus Derived from LibriSpeech for Text-to-Speech
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering