WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
summary
The gist
WASIL introduces a novel, in-the-wild dataset capturing real Arabic spoken interactions with Large Language Models (LLMs).
In short
WASIL is a new dataset of real Arabic spoken interactions with LLMs, featuring user feedback and gold transcripts. It helps researchers study how errors from speech-to-text systems affect LLM performance. The data allows isolating whether poor results come from bad transcription or poor response content.
Key concepts
- Cascaded ASR→LLM Pipeline
- This refers to a system where speech is first converted to text by an Automatic Speech Recognition (ASR) system, and then that text is fed into a Large Language Model (LLM). Errors can happen at both steps, making it hard to tell if the final poor answer was due to bad transcription or the LLM's inability to understand the text.
- Intrinsic Answerability
- This label classifies user turns based on whether they are inherently answerable. Labels include 'Answerable/Clear,' 'Ambiguous/Needs-Clarification,' and 'Out-of-Domain/Unsupported.' This helps researchers distinguish failures caused by the user's question itself from failures caused by errors in the transcription or the LLM's response.
- Multi-judge LLM Scoring
- This is an evaluation method where several Large Language Models are used to score responses based on a detailed rubric. Judges assess quality across criteria like intent precision, context awareness, and coherence. This provides a more robust and scalable way to measure response quality than relying on just one model.
- Like/Dislike Signals
- These are explicit user feedback signals where users can indicate if they liked or disliked an assistant's response. This feedback is crucial because it helps researchers link specific types of errors—like factual inaccuracy or failure to follow instructions—to the underlying cause, whether it stems from transcription issues or the LLM's output.
Terminology used across episodes
This episode discusses
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs · Paper Radio
- VoiceAssistant-Eval: Benchmarking AI Assistants across Listening, Speaking, and Viewing
- The Cascade Equivalence Hypothesis: When Do Speech LLMs Behave Like ASR to LLM Pipelines?
- An Analysis of Dialogue Repair in Voice Assistants
- Reject or Not?: A Benchmark for Voice Assistant Query Rejection in Smart Home Scenario and an Improved Method Based on LLMs
- Fanar: An Arabic-Centric Multimodal Generative AI Platform
- Gemini: A Family of Highly Capable Multimodal Models
- Instruction-Following Evaluation for Large Language Models
- Constitutional AI: Harmlessness from AI Feedback
- OpenAI GPT-5 System Card
- Beyond Seeing: Evaluating Multimodal LLMs on Tool-Enabled Image Perception, Transformation, and Reasoning
The paper
WASIL: In-the-Wild Arabic Spoken Interactions with LLMs · Read on arXiv
Zien Sheikh Ali ID, Hamdy Mubarak ID, Soon-Gyo Jung, Hunzalah Hassan Bhatti ID, Firoj Alam ID, Shammur Absar Chowdhury ID
Qatar Computing Research Institute, Qatar
Large Language Models (LLMs) voice assistants are commonly built as cascaded Automatic Speech recognition (ASR) to LLM systems, where recognition errors can distort user intent. Dislikes may also arise from ambiguous, out-of-domain, or non-request turns, making it hard to isolate ASR effects. We release WASIL (it denotes connection or linking in Arabic): in-the-wild Arabic spoken interaction prompts with audio, ASR hypotheses, assistant responses, and explicit like/dislike feedback (8,529 turns; 14.2% dislikes), plus a 2,000-turn test set covering Modern Standard Arabic (MSA) and four major dialects with their labels. We provide low-cost gold transcripts via multi-ASR agreement-guided post-editing and annotate answerability (answerable, ambiguous/needs-clarification, unsupported, not-a-request/noise) to separate intrinsic unanswerability from ASR-induced degradation. Finally, we describe scalable reference-free evaluation of responses from ASR vs. gold transcripts using multi-judge LLM scoring.
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "WASIL: In-the-Wild Arabic Spoken Interactions with LLMs".
Jane: WASIL introduces a novel, in-the-wild dataset capturing real Arabic spoken interactions with Large Language Models (LLMs).
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Let's start by looking at the title of this paper, "WASIL: In-the-Wild Arabic Spoken Interactions with LLMs." It tells you immediately that they are focusing on interactions happening naturally, not just scripted questions.
Jane: And the authors—Zien Sheikh Ali, Hamdy Mubarak, Soon-Gyo Jung, Hunzalah Hassan Bhatti, Firoj Alam, and Shammur Absar Chowdhury—they are clearly bringing together expertise from different linguistic and research backgrounds to tackle this complex area.
Lu: The combination of researchers from Qatar Computing Research Institute alongside others suggests a deep dive into the nuances of Arabic spoken language across various contexts. It sets the stage for a very thorough investigation.
Meng: I'm thinking about what those authors are trying to prove with that specific title; they are trying to bridge the gap between spoken input and LLM output quality in Arabic. Does this mean their work is specifically focused on dialectal robustness?
Lalam: Exactly, Meng; because the paper mentions testing across four major dialects, it seems like a big part of their mission is understanding how variations in speech affect AI performance. It’s not just about getting the words right; it’s about keeping the conversation going.
The paper's summary: Tom: Moving on to what the WASIL paper actually does, they introduce a dataset that includes audio, ASR hypotheses, assistant responses, and explicit like or dislike feedback. It's a big collection of real interactions from users across Algeria, Egypt, Sudan, and Syria.
Jane: That feedback mechanism is key because it lets them gather direct user opinions on the assistant's performance rather than relying on just automated metrics. They even provide a multi-label taxonomy for dislikes like "failed to follow instructions" or issues with "style and format."
Lu: The core contribution seems to be the explicit labeling of answerability—categorizing turns as answerable, ambiguous, unsupported, or not-a-request or noise. That helps separate if the user was genuinely asking something the AI couldn't handle versus if the ASR just misheard them.
Meng: I see that separation as crucial for practical deployment; knowing whether a failure is because of bad speech recognition versus a lack of knowledge in the LLM helps us decide where to invest our engineering efforts. That distinction makes sense for real-world application.
Lalam: For me, the summary highlights how they are using this feedback set to analyze answerability, which helps isolate ASR effects from response content issues. It’s a very structured way of dissecting why users might be unhappy with the final result.
The paper's improvements: Tom: The paper suggests several ways to improve the overall system by focusing on separating intrinsic unanswerability from ASR-induced degradation, which they tackle through their annotation schema. They are using these labels to directly correlate with user dislikes.
Jane: They also propose a method for creating low-cost gold transcripts by using multi-ASR agreement as a proxy for transcription reliability, accepting high-agreement utterances with minimal edits and prioritizing post-editing for the ones that don't agree well.
Lu: That strategy of using multiple ASR systems like Fanar and Gemini to create a reliable reference set is clever because it aims to reduce the manual effort required for transcription significantly, potentially cutting human effort by sixty-eight percent.
Meng: I'm interested in the practical implication of that low-cost gold creation; if we can generate high-quality training data efficiently without massive manual annotation, that makes scaling up our testing frameworks much more feasible. That’s a solid engineering idea.
Lalam: I think the improvement they suggest regarding answerability labeling is very powerful because it gives us a clear roadmap for tuning the system to handle ambiguity better, which directly addresses how users feel about the AI's helpfulness.
Conclusion: Tom: So, to wrap up on WASIL, the paper shows that performance under audio and ASR inputs can really diverge from what you’d see with just gold transcripts. They found that providing those transcripts substantially reduces these degradations in depth and specificity, even when dealing with complex input.
Jane: The authors conclude that most user dissatisfaction comes from a small set of failure modes, primarily related to helpfulness and correctness, which gives us a focus for improving the LLM's core reasoning capabilities. They also point out that dialect coverage is important because it affects how transcription errors translate into task failures.
Lu: I think the overall implication is that we need more principled ways to study this error propagation from ASR to LLM reasoning, and WASIL provides a foundation for that using their structured approach. It opens up new avenues for testing robustness across different Arabic varieties.
Meng: From a practical standpoint, the paper gives us a way to use like or dislike signals alongside audio judgments to benchmark models in ways that reflect actual user preference, not just internal scores. That’s useful for setting real-world performance targets.
Lalam: Ultimately, WASIL gives us the tools to build cultural and dialect-aware Arabic assistants that are more reliable because we understand exactly where the breakdowns happen in the pipeline. It's a really valuable resource for advancing how we develop these kinds of systems.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck