The Public Discourse Corpus (PDC): A Speaker-Attributed Dataset for Valence and Epistemic Modality with Target Speaker Participation
cs.CL
Submitted: 2026-07-27
Updated: 2026-07-27
Comments: 16 pages, 1 figure
License: http://creativecommons.org/licenses/by/4.0/
The gist: We introduce the Public Discourse Corpus (PDC), the first dataset of public-figure interview speech jointly annotated for affective valence and epistemic modality.
Terminology
Abstract
We introduce the Public Discourse Corpus (PDC), the first dataset of public-figure interview speech jointly annotated for affective valence and epistemic modality. The corpus contains 998 videos from 100 speakers across seven professional domains, yielding 186,642 sentences (3.1 million words) after sentence segmentation and filtering. To ensure that all retained videos contain analyzable speech from the intended speaker, we introduce Target Speaker Participation (TSP)---a five-category annotation taxonomy with documented inter-annotator reliability (κ= 0.616)---as a key methodological contribution that any corpus construction project can adopt. Target-speaker turns are separated from interviewer and third-party speech through an audio-first diarization pipeline combining local Whisper ASR with pyannote speaker separation, released as an open-source implementation. We release the annotated corpus, the annotation tools, the cross-provider validation sample, and the complete processing pipeline. The dataset is available at https://huggingface.co/datasets/ictchenbo/public-discourse-corpus.
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering