WASIL: In-the-Wild Arabic Spoken Interactions with LLMs

arXiv:2605.16364 · cs.SD, cs.AI, cs.CL · Submitted 2026-05-09 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "WASIL: In-the-Wild Arabic Spoken Interactions with LLMs".

Jane: WASIL introduces a novel, in-the-wild dataset capturing real Arabic spoken interactions with Large Language Models (LLMs).

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Let's start by looking at the title of this paper, "WASIL: In-the-Wild Arabic Spoken Interactions with LLMs." It tells you immediately that they are focusing on interactions happening naturally, not just scripted questions.

Jane: And the authors—Zien Sheikh Ali, Hamdy Mubarak, Soon-Gyo Jung, Hunzalah Hassan Bhatti, Firoj Alam, and Shammur Absar Chowdhury—they are clearly bringing together expertise from different linguistic and research backgrounds to tackle this complex area.

Lu: The combination of researchers from Qatar Computing Research Institute alongside others suggests a deep dive into the nuances of Arabic spoken language across various contexts. It sets the stage for a very thorough investigation.

Meng: I'm thinking about what those authors are trying to prove with that specific title; they are trying to bridge the gap between spoken input and LLM output quality in Arabic. Does this mean their work is specifically focused on dialectal robustness?

Lalam: Exactly, Meng; because the paper mentions testing across four major dialects, it seems like a big part of their mission is understanding how variations in speech affect AI performance. It’s not just about getting the words right; it’s about keeping the conversation going.

The paper's summary: Tom: Moving on to what the WASIL paper actually does, they introduce a dataset that includes audio, ASR hypotheses, assistant responses, and explicit like or dislike feedback. It's a big collection of real interactions from users across Algeria, Egypt, Sudan, and Syria.

Jane: That feedback mechanism is key because it lets them gather direct user opinions on the assistant's performance rather than relying on just automated metrics. They even provide a multi-label taxonomy for dislikes like "failed to follow instructions" or issues with "style and format."

Lu: The core contribution seems to be the explicit labeling of answerability—categorizing turns as answerable, ambiguous, unsupported, or not-a-request or noise. That helps separate if the user was genuinely asking something the AI couldn't handle versus if the ASR just misheard them.

Meng: I see that separation as crucial for practical deployment; knowing whether a failure is because of bad speech recognition versus a lack of knowledge in the LLM helps us decide where to invest our engineering efforts. That distinction makes sense for real-world application.

Lalam: For me, the summary highlights how they are using this feedback set to analyze answerability, which helps isolate ASR effects from response content issues. It’s a very structured way of dissecting why users might be unhappy with the final result.

The paper's improvements: Tom: The paper suggests several ways to improve the overall system by focusing on separating intrinsic unanswerability from ASR-induced degradation, which they tackle through their annotation schema. They are using these labels to directly correlate with user dislikes.

Jane: They also propose a method for creating low-cost gold transcripts by using multi-ASR agreement as a proxy for transcription reliability, accepting high-agreement utterances with minimal edits and prioritizing post-editing for the ones that don't agree well.

Lu: That strategy of using multiple ASR systems like Fanar and Gemini to create a reliable reference set is clever because it aims to reduce the manual effort required for transcription significantly, potentially cutting human effort by sixty-eight percent.

Meng: I'm interested in the practical implication of that low-cost gold creation; if we can generate high-quality training data efficiently without massive manual annotation, that makes scaling up our testing frameworks much more feasible. That’s a solid engineering idea.

Lalam: I think the improvement they suggest regarding answerability labeling is very powerful because it gives us a clear roadmap for tuning the system to handle ambiguity better, which directly addresses how users feel about the AI's helpfulness.

Conclusion: Tom: So, to wrap up on WASIL, the paper shows that performance under audio and ASR inputs can really diverge from what you’d see with just gold transcripts. They found that providing those transcripts substantially reduces these degradations in depth and specificity, even when dealing with complex input.

Jane: The authors conclude that most user dissatisfaction comes from a small set of failure modes, primarily related to helpfulness and correctness, which gives us a focus for improving the LLM's core reasoning capabilities. They also point out that dialect coverage is important because it affects how transcription errors translate into task failures.

Lu: I think the overall implication is that we need more principled ways to study this error propagation from ASR to LLM reasoning, and WASIL provides a foundation for that using their structured approach. It opens up new avenues for testing robustness across different Arabic varieties.

Meng: From a practical standpoint, the paper gives us a way to use like or dislike signals alongside audio judgments to benchmark models in ways that reflect actual user preference, not just internal scores. That’s useful for setting real-world performance targets.

Lalam: Ultimately, WASIL gives us the tools to build cultural and dialect-aware Arabic assistants that are more reliable because we understand exactly where the breakdowns happen in the pipeline. It's a really valuable resource for advancing how we develop these kinds of systems.

Zien Sheikh Ali ID, Hamdy Mubarak ID, Soon-Gyo Jung, Hunzalah Hassan Bhatti ID, Firoj Alam ID, Shammur Absar Chowdhury ID

Qatar Computing Research Institute, Qatar

cs.SD, cs.AI, cs.CL

Submitted: 2026-05-09

Updated: 2026-09-30

Comments: Spoken Prompts, Multilingual LLMs, Speech-based Evaluation, Dialectal Speech, Low-resource Languages, Conversational AI, Speech-to-Text QA, Real-world Interaction, Spoken Language Understanding

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 92/100

The gist: WASIL introduces a novel, in-the-wild dataset capturing real Arabic spoken interactions with Large Language Models (LLMs).

Key concepts

Cascaded ASR→LLM Pipeline
This refers to a system where speech is first converted to text by an Automatic Speech Recognition (ASR) system, and then that text is fed into a Large Language Model (LLM). Errors can happen at both steps, making it hard to tell if the final poor answer was due to bad transcription or the LLM's inability to understand the text.
Intrinsic Answerability
This label classifies user turns based on whether they are inherently answerable. Labels include 'Answerable/Clear,' 'Ambiguous/Needs-Clarification,' and 'Out-of-Domain/Unsupported.' This helps researchers distinguish failures caused by the user's question itself from failures caused by errors in the transcription or the LLM's response.
Multi-judge LLM Scoring
This is an evaluation method where several Large Language Models are used to score responses based on a detailed rubric. Judges assess quality across criteria like intent precision, context awareness, and coherence. This provides a more robust and scalable way to measure response quality than relying on just one model.
Like/Dislike Signals
These are explicit user feedback signals where users can indicate if they liked or disliked an assistant's response. This feedback is crucial because it helps researchers link specific types of errors—like factual inaccuracy or failure to follow instructions—to the underlying cause, whether it stems from transcription issues or the LLM's output.

Terminology

Summary

WASIL introduces a novel, in-the-wild dataset capturing real Arabic spoken interactions with Large Language Models (LLMs). This resource is crucial because it addresses the significant evaluation challenge arising from cascaded ASR→LLM pipelines, where user dissatisfaction can stem from multiple confounded sources, including ASR errors that distort intent and intrinsic ambiguity in user turns. By providing explicit like/dislike feedback and gold transcripts with intrinsic answerability labels, WASIL enables researchers to isolate the effects of transcription quality versus response content on downstream LLM performance.

Dataset Composition and Data Collection

WASIL is a dataset of in-the-wild Arabic spoken interactions with an LLM-based assistant, containing approximately 9,304 turns from 93 users across four Arab countries (Algeria, Egypt, Sudan, and Syria). The data spans diverse topics including open discussion and creative writing. A key feature is the inclusion of explicit user feedback on assistant responses, including like or dislike signals and a multi-label taxonomy for dislikes such as failed to follow instructions, lacked factual accuracy, or issues related to style and format.

Gold Transcript Creation Strategy

To ensure high-quality reference data, WASIL employs a low-cost reference creation strategy. This involves using multi-ASR agreement as a proxy for transcription reliability by collecting multiple ASR hypotheses from systems like Fanar and Gemini. The paper notes that high-agreement utterances are accepted with minimal edits, while low-agreement utterances are prioritized for post-editing, which is hypothesized to reduce human effort by 68%. This approach yields gold transcripts that show strong agreement (WER = 0.070) and nearly identical meaning (cosine similarity = 0.98) when compared to fully manual transcriptions.

Annotation Schema for Error Isolation

The dataset includes a detailed annotation schema designed to separate intrinsic unanswerability from ASR-induced degradation. This involves annotating intrinsic answerability using four mutually exclusive labels: Answerable/Clear, Ambiguous/Needs-Clarification, Out-of-Domain/Unsupported, and Not-aRequest/Backchannel/Noise. The paper analyzes the association between these intrinsic labels and user dislikes to help separate failures driven by transcription issues from those driven by response content.

Evaluation Methodology

The authors propose a scalable, reference-free evaluation of responses using multi-judge LLM scoring protocols. This method benchmarks multiple LLMs (open and closed) across different input conditions, including transcript using ASR vs. gold transcripts and raw audio. They utilize a rubric-based LLM-as-a-judge method to assess response quality based on criteria such as Intent Precision, Context Awareness, Specificity, Depth & Thoroughness, Grounding & Honesty, Format & Language Compliance, and Coherence.

Key Findings on Performance

The experiments reveal that performance under audio and ASR inputs can diverge substantially from gold transcription settings. Specifically, while direct audio processing shows significant drops in depth and specificity compared to gold transcripts, providing transcripts substantially mitigates these degradations. Furthermore, the analysis shows that most dissatisfaction is driven by a small set of failure modes, with the dominant marginal dimensions being helpfulness/task success and correctness/truthfulness. The findings suggest that dialect coverage and a model’s ability to handle dialectal lexical and syntactic variation play an important role in how transcription errors translate into task failures.

Conclusion

WASIL provides a foundation for principled studies of error propagation from ASR to LLM reasoning, preference-based alignment using like/dislike signals with audio-grounded judgments, and culturally aware Arabic assistant development. It enables researchers to evaluate exactly how ASR errors degrade LLM responses across various dialects and input modalities.


**(Word Count Check: Approximately 480 words.

Improvements for AI systems

Here are specific improvements to AI systems based on the WASIL paper, focusing on mitigating ASR errors and enhancing LLM reasoning:

  1. Acknowledge and Mitigate ASR Error Propagation:

Identify that user dissatisfaction often stems from ASR errors distorting user intent (Section 1, 2.2).

Improving AI systems can involve implementing a two-stage evaluation pipeline where the output of an ASR system is fed into a diagnostic module that assesses the semantic distance or similarity between the transcribed text and potential ground truths (using metrics like SemDist discussed in Section 1). This allows for pre-emptive flagging of transcriptions likely to cause downstream errors.

  1. Implement Robust, Low-Cost Reference Transcription Strategies:

For systems relying on ASR transcripts, integrate a multi-ASR agreement scoring mechanism (Section 2.4). Instead of immediately accepting the first transcript, compare outputs from multiple ASR models (like Fanar and Gemini) and use the agreement score as a proxy for reliability.

The improved system can then dynamically adjust its confidence: if the inter-ASR agreement is low, it triggers a fallback mechanism—either asking for clarification or generating a response based on the most likely transcript rather than committing to an error. This directly addresses Section 3.6.5's findings that high-agreement utterances require minimal post-editing effort, optimizing human effort while maintaining quality.

  1. Enhance Intent and Context Awareness via Answerability Labeling:

Integrate the intrinsic answerability annotation schema (Section 3.4) into the LLM’s internal reasoning process or as a pre-processing step for query classification.

The improved AI system can explicitly classify an incoming query as:

  • Answerable/Clear (proceed with standard response).

  • Ambiguous/Needs Clarification (trigger a clarifying question to reduce uncertainty).

  • Out-of-Domain/Unsupported (initiate a graceful rejection or request for scope adjustment).

This directly addresses the findings in Section 3.6.7, ensuring that the system doesn't waste resources generating complex answers for prompts that are inherently unanswerable or require different interaction strategies, thereby improving user satisfaction metrics like Intent Precision and Context Awareness (as measured by the LLM-as-a-judge rubric).

  1. Improve Response Quality via Multi-Judge Rubric Scoring:

Move beyond simple correctness checks by adopting the detailed rubric-based LLM-as-a-judge protocol (Section 4.2).

The improved system should be trained or prompted to evaluate its own outputs against a comprehensive rubric that includes: Intent Precision, Context Awareness, Specificity, Depth/Thoroughness, Grounding/Honesty (avoiding hallucination), Format Compliance, and Coherence. This forces the model to reason beyond surface-form matching and focus on the nuanced dimensions of conversational success that users actually value.

  1. Develop Dialect-Aware Robustness:

In systems handling spoken Arabic, implement dialect detection as an early feature (Section 3.3).

The improved AI system can dynamically load dialect-specific models or parameters based on the detected dialect (MSA vs. Egyptian vs. Sudanese, etc.). This directly addresses the findings in Section 5.2 that performance degradation is most pronounced in raw audio, especially for dialects like Algerian Arabic, and shows how high-quality transcription (via ASR) can mitigate this gap significantly.

  1. Refine Safety and Cultural Alignment Mechanisms:

Utilize the fine-grained meta-feedback categories (Section 3.6.7) to create a more sophisticated safety layer than simple keyword filtering.

The improved system can be fine-tuned to recognize patterns associated with specific meta-clusters, such as the co-occurrence of Helpfulness/Task Success and Cultural Alignment. This allows the AI to provide responses that are not only factually correct but also contextually appropriate for Arabic/Islamic cultural norms, reducing instances of undesirable cultural misalignments (Meta-cluster E).

  1. Optimize Cost-Effective Gold Data Creation:

For continuous improvement loops, use the multi-ASR agreement score (Section 3.6.5) to create a smart gold set.

Instead of manually transcribing everything, prioritize high-agreement utterances for human post-editing and only apply full manual transcription effort to low-agreement samples. This creates a highly accurate reference dataset efficiently, allowing for rapid iteration on ASR and LLM performance without the prohibitive cost of transcribing 100% of the data from scratch.

Abstract

Large Language Models (LLMs) voice assistants are commonly built as cascaded Automatic Speech recognition (ASR) to LLM systems, where recognition errors can distort user intent. Dislikes may also arise from ambiguous, out-of-domain, or non-request turns, making it hard to isolate ASR effects. We release WASIL (it denotes connection or linking in Arabic): in-the-wild Arabic spoken interaction prompts with audio, ASR hypotheses, assistant responses, and explicit like/dislike feedback (8,529 turns; 14.2% dislikes), plus a 2,000-turn test set covering Modern Standard Arabic (MSA) and four major dialects with their labels. We provide low-cost gold transcripts via multi-ASR agreement-guided post-editing and annotate answerability (answerable, ambiguous/needs-clarification, unsupported, not-a-request/noise) to separate intrinsic unanswerability from ASR-induced degradation. Finally, we describe scalable reference-free evaluation of responses from ASR vs. gold transcripts using multi-judge LLM scoring.

Sources

Related papers