VoxPrivacy: A Benchmark for Evaluating Interactional Privacy of Speech Language Models
eess.AS, cs.AI, cs.SD
Submitted: 2026-01-27
Updated: 2026-09-03
Project page: https://interactionalprivacy.github.io
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- GPT-4 Technical Report
- Common Voice: A Massively-Multilingual Speech Corpus
- SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak Attacks
- Qwen2-Audio Technical Report
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Kimi-Audio Technical Report
- AISHELL-2: Transforming Mandarin ASR Research Into Industrial Scale
- MTalk-Bench: Evaluating Speech-to-Speech Models in Multi-Turn Dialogues via Arena-style and Rubrics Protocols
- LLaMA-Omni: Seamless Speech Interaction with Large Language Models
- Emilia: A Large-Scale, Extensive, Multilingual, and Diverse Dataset for Speech Generation
- SOVA-Bench: Benchmarking the Speech Conversation Ability for LLM-based Voice Assistant
- AudioTrust: Benchmarking the Multifaceted Trustworthiness of Audio Large Language Models
- Baichuan-Omni-1.5 Technical Report
- DeepSeek-V3 Technical Report
- Voxtral
- Large Language Model Can Transcribe Speech in Multi-Talker Scenarios with Versatile Instructions
- Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the Wild
- Can LLMs Keep a Secret? Testing Privacy Implications of Language Models via Contextual Integrity Theory
- MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations
- DNSMOS: A Non-Intrusive Perceptual Objective Speech Quality metric to evaluate Noise Suppressors
Related papers
- X-VC: Zero-shot Streaming Voice Conversion in Codec Space
- Autoregressive Guidance of Deep Spatially Selective Filters using Bayesian Tracking for Efficient Extraction of Moving Speakers
- Anonymization, Not Elimination: Utility-Preserved Speech Anonymization
- Towards Audio Token Compression in Large Audio Language Models
- WaveScat: Wavelet Scattering Front-Ends with Self-Supervised Features for Speech Deepfake Detection
- ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions