VoiceLongMemEval: Do Assistants Remember How You Sounded?
cs.AI, cs.LG
Submitted: 2026-09-01
Updated: 2026-09-02
Terminology
Sources
- LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory
- PersonaMem-v2: Towards Personalized Intelligence via Learning Implicit User Personas and Agentic Memory
- LongMemEval-V2: Evaluating Long-Term Agent Memory Toward Experienced Colleagues
- DialSim: A Dialogue Simulator for Evaluating Long-Term Multi-Party Dialogue Understanding of Conversational Agents
- Do LLMs Recognize Your Preferences? Evaluating Personalized Preference Following in LLMs
- A-MBER: Affective Memory Benchmark for Emotion Recognition
- Audio2Tool: Speak, Call, Act -- A Dataset for Benchmarking Speech Tool Use
- MMSU: A Massive Multi-task Spoken Language Understanding and Reasoning Benchmark
- S2S-Arena: Evaluating Paralinguistic Instruction Following in Speech-to-Speech Models
- ParaS2S: Benchmarking and Aligning Spoken Language Models for Paralinguistic-aware Speech-to-Speech Interaction
- GPT-4o System Card
- Qwen2-Audio Technical Report
- Qwen2.5-Omni Technical Report
- Moshi: a speech-text foundation model for real-time dialogue
- GLM-4-Voice: Towards Intelligent and Human-Like End-to-End Spoken Chatbot
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection