CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

arXiv:2608.13101 · cs.CL, eess.AS · Submitted 2026-08-13 · Read on arXiv

Nhan Phan, Ilona Lähteenmäki, Anna von Zansen, Olli-Pekka Pauna, Yaroslav Getman, Tamás Grósz, Mikko Kurimo

Aalto University · University of Helsinki · Walton Institute

cs.CL, eess.AS

Submitted: 2026-08-13

Updated: 2026-08-14

Comments: To be submitted to ICASSP 2027. Code is available at https://github.com/aalto-speech/casa

Code: https://github.com/aalto-speech/casa

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 50/100

The gist: CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model Abstract Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large

Terminology

Summary

CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model

Abstract

Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners’ speaking performance. However, existing studies provide limited analysis of how acoustic and content information contribute to predictions and how stable the resulting performance is. We propose CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content. On the Speak & Improve Corpus 2025, CASA achieves a root mean square error (RMSE) of 0.358, improving on the previous best RMSE while using approximately half the estimated inference parameters. The general-purpose architecture is designed for adaptation to other ASA corpora without structural changes and relies on three handcrafted fluency features. Through ablations and repeated runs, we analyze the individual and complementary contributions of acoustic and content information, examine performance variability, and demonstrate the potential of large language model reasoning for training-free content validation.

Introduction

Automatic speaking assessment (ASA) aims to automatically estimate second-language (L2) learners’ oral proficiency from their spoken responses. Compared with assessments conducted exclusively by human raters, ASA can reduce scoring costs, improve scoring consistency, and enable scalable and timely feedback. Nevertheless, holistic speaking assessment remains challenging because speaking proficiency is reflected in multiple interacting sources of evidence, including speech delivery (how something is said) and speech content (what is said).

With recent advances in large language models (LLMs), speech LLMs have increasingly been explored for ASA. Ma et al. used Qwen2-Audio for L2 oral proficiency assessment and achieved state-of-the-art (SOTA) performance on the S&I test set, obtaining a Pearson correlation coefficient (PCC) of 0.833; RMSE was not reported. More recently, Lin et al. combined Phi-4-Multimodal with a Whisper-large encoder, achieving an RMSE of 0.360 on the same test set. In contrast, Cai et al. achieved a comparable RMSE of 0.364 without using a speech LLM by combining a Whisper-medium encoder and BERT-based graders with handcrafted acoustic and linguistic features.

Despite their strong performance, speech-LLM-based graders rely on large multimodal backbones, which increase computational requirements during inference and may limit practical deployment. Although Cai et al. avoid using a speech LLM, their system combines multiple graders and handcrafted features. It also relies on a separate ASR model to produce more verbatim transcripts, increasing pipeline complexity and potentially limiting portability. Moreover, previous studies provide limited analysis of the respective contributions of acoustic and textual information. As their implementations have not been publicly released, reproducibility and further investigation are also limited.

In this work, we propose CASA (content-acoustic speaking assessment), a compact, general-purpose two-branch ASA architecture that explicitly separates acoustic information from linguistic information represented in ASR transcripts. The model employs a Whisper-medium encoder for the acoustic branch, and Qwen3.5-2B for the content branch, making it substantially smaller than existing speech-LLM-based systems. Through detailed ablation and component-level analyses, we investigate the individual and complementary contributions of acoustic and textual information. Our model achieves an RMSE of 0.358 on the S&I test set, slightly outperforming the SOTA result of 0.360 while requiring substantially fewer parameters during inference. The architecture is designed as a general-purpose ASA framework that is not tailored to any dataset and uses only three simple fluency features. Our code is publicly available.

Method

Dataset: The spoken language assessment (SLA) task of the S&I Corpus consists of four parts: Parts 1, 3, 4, and 5. Part 1 contains six short questions, with responses typically lasting 10–20 seconds. Part 3 consists of one-minute open-ended questions in which speakers express their opinions, whereas Part 4 contains one-minute tasks requiring speakers to describe a process illustrated in a graphical prompt. Finally, Part 5 contains five opinion-based questions, each with a maximum response time of 20 seconds. In total, the corpus contains approximately 315 hours of speech and is divided into training, development, and test sets, with the score distributions kept as balanced as possible across the splits. Each part was rated by human assessors and assigned a holistic score based on the CEFR scale. The CEFR levels map to a 2.0–5.5 scale in 0.5 steps (A2=2.0, A2+=2.5,..., C1+=5.5); the final score is the arithmetic mean of the four part scores.

Model: To support adaptation to other ASA settings, a single shared model scores each part independently. The overall score is then calculated as the arithmetic mean of the per-part predictions. The design builds on Phan et al., who showed that a single Whisper-small encoder achieves strong performance but still lags behind SOTA. Their model uses global average pooling, which may discard fine-grained acoustic variation. CASA instead adopts a Whisper-medium encoder, kept frozen: the encoder–decoder still produces the ASR transcript, while LoRA adapters on the encoder produce the representation used by the acoustic branch.

The acoustic branch, originating from the Whisper encoder, assesses the speaker’s delivery. As the S&I corpus provides no separate delivery score, it is supervised using the holistic per-part CEFR label—the main loss target y—which the branch must predict from acoustics alone. Because Whisper processes at most 30 s at a time, each answer is split into 30 s chunks up to a 2 min cap, with each chunk yielding 1,500 frame-level vectors. Rather than global average pooling, which discards acoustic detail, the chunk encodings are concatenated and adjacent frame pairs are averaged, reducing the temporal resolution from 20 ms to 40 ms and halving the sequence length while preserving most acoustic information.

Parts P1 and P5 comprise several short answers, so before the Transformer two zero-initialized, learned embeddings are added to every frame: a task embedding (which part, P1–P5) and a segment embedding (which answer within the part). This is intended to let the aggregator distinguish frames from different answers. The tagged frames pass through a two-layer Transformer encoder with rotary position embeddings (RoPE), which help preserve positional information over long concatenated sequences. Following the [CLS] pooling of BERT, shown to transfer well to audio representation learning, the sequence is summarized with a [CLS] token.

The acoustic summary feeds two heads. An MLP projector maps it to four acoustic soft tokens for the LLM, allowing the acoustic information to be represented across multiple embeddings while remaining compact. A linear auxiliary head maps it to a scalar CEFR estimate that is detached and rendered as text (e.g., acoustic cefr estimate: 4.5) in the LLM prompt, and also drives the auxiliary loss. Both objectives backpropagate into the shared branch, so the main loss (through the soft tokens) and the auxiliary loss (through the auxiliary head) jointly shape the acoustic representation. Since acoustics alone cannot determine CEFR, the auxiliary term is kept from dominating by assigning it a small weight (0.1) and, more importantly, using a tolerance of ±1 score point. Predictions within this margin incur no auxiliary loss, and only deviations beyond the margin are penalized. Among the tested auxiliary-loss weights and tolerance levels, this combination achieved the best performance.

On the content branch, the frozen Whisper encoder–decoder transcribes each answer. The transcripts are generated offline before training for computational efficiency. For each part, the transcript is paired with its task prompt using followed by the corresponding Q1/A1 pairs. Qwen3.5-2B receives the fused input sequence in the following order: the four acoustic soft tokens, the detached acoustic CEFR estimate represented as text, the scoring rubric, the question–answer pairs, and three simple fluency statistics derived from the audio and ASR output: duration, silence ratio, and speech rate (words/s). The base Qwen parameters remain frozen and are adapted with another LoRA adapter. The model performs a single forward pass without text generation, and a linear regression head applied to the final-token representation predicts the part score ŷ.

The training objective combines the main loss from the fused scorer with the auxiliary loss from the acoustic branch:

L = MSE(ŷ, y) + 0.1 MSEτ(ŷaux, y), where MSEτ(a, b) = max(a − b − τ, 0)2, τ = 1, where ŷ is the final part prediction, ŷaux is the acoustic-only prediction, and y is the ground-truth per-part CEFR score.

Results

Training was conducted on a single NVIDIA H100 80 GB GPU with a batch size of 16 and gradient accumulation over two steps. Learning rates were 2×10−4 (acoustic LoRA), 1×10−4 (LLM LoRA), 5×10−5 (other modules). Training took approximately two hours. The model configuration and checkpoint were selected solely according to overall RMSE on the development set. CASA achieves an RMSE of 0.358, marginally lower than the previous SOTA result of 0.360. As this difference is small, the authors do not claim a meaningful improvement in accuracy. Instead, CASA provides a better accuracy–model-size tradeoff, achieving comparable SOTA performance with approximately half the estimated inference parameters of the NTNU speech-LLM system. Team Perezoso can likewise be considered SOTA-level despite not using an LLM, although its pipeline combines multiple graders, handcrafted features, and a separate fine-tuned Parakeet-TDT-1.1B ASR model. One Whisper remains the most compact system, using only a Whisper-small encoder, although with a higher RMSE of 0.372.

In terms of architecture, CASA is closer to Cai et al. than to speech-LLM approaches. Cai et al. probe the final six layers of a frozen Whisper-medium encoder, while Lin et al. derive an acoustic proficiency prior from the final representation of a frozen Whisper-large encoder. CASA instead adapts the Whisper-medium encoder across its layers using LoRA. Since acoustic and semantic information is distributed unevenly across Whisper layers, this motivates CASA’s task-specific adaptation rather than reliance on a fixed final-layer representation. Cai et al. further combine their Whisper grader with several BERT-based graders and handcrafted features, whereas CASA integrates the acoustic representation with an LLM-based content branch and a tolerance-based auxiliary objective. Apart from three fluency statistics, CASA requires no handcrafted assessment features.

Analysis

The modular acoustic encoder allows swapping in CrisperWhisper, a verbatim-transcription Whisper-large that preserves the disfluencies that standard Whisper smooths away. Using the same hyperparameters as for Whisper-medium, the overall RMSE may be worse than what proper tuning would yield. CASA-Crisper improves on A2, indicating that verbatim transcripts help the system catch the content errors of A2 and B1 speakers. This gain matters because A2 is an under-represented class. B2 and C1, however, get worse: for otherwise fluent speakers, verbatim transcription may introduce inauthentic errors that hide the range marking the top band.

Two self-supervised learning (SSL) models were also explored: WavLM and wav2vec2 XLS-R 300M, swapping only the acoustic encoder while keeping CASA’s Whisper ASR transcript, so the acoustic representation is the sole difference from CASA. Both substantially underperform CASA, and an English-finetuned variant also provides no improvement. This gap is attributed partly to architectural mismatch: CASA’s aggregator and training recipe were developed for Whisper representations and may not transfer optimally to SSL encoders. SSL encoders are effective for ASA in other architectures; adapting CASA to exploit them (e.g., a dedicated pronunciation head) is left to future work.

Neither a larger acoustic encoder (CrisperWhisper) nor a larger LLM (Qwen3.5-4B, RMSE 0.364) improves RMSE; even pairing Whisper-large-v3 with the 4B LLM yields only 0.362. Performance thus appears not to be limited by model capacity, though larger models may require different hyperparameters.

CASA achieves per-part RMSEs of 0.476, 0.454, 0.490, and 0.444 for P1, P3, P4, and P5, respectively. Performance is weaker on P1 and P4, consistent with Lin et al. A qualitative analysis of the ASR transcripts, conducted with a co-author in applied linguistics, suggests that task format may partly explain this pattern: short personal questions in P1 and process descriptions in P4 provide less freedom to demonstrate broad linguistic and content proficiency than the opinion-based tasks in P3 and P5. Analyzing the same task formats, Banno et al. similarly found that P1 and P4 place greater emphasis on fluency, whereas P3 and P5 place greater weight on thematic development and other content-related dimensions. P1 and P4 may therefore provide fewer discriminative textual cues. In CASA, the task embeddings in the Transformer block proved ineffective: they remained close to their zero initialization, contributing little task information to the acoustic representation. Future work could investigate separate prediction heads or graders for more constrained and open-ended tasks.

To assess the stability of the SOTA-level result, 9 additional runs were performed for each configuration: 4 with seed 1011 and 5 with distinct seeds. Here, 4e-4 doubles the acoustic-branch LoRA learning rate, while aux-0 sets the auxiliary-loss weight to 0. The auxiliary head provides a small but consistent improvement across repeated runs, reducing mean RMSE by 0.004 and helping CASA reach SOTA-level performance. In two additional single-run ablations, blocking either the auxiliary-loss gradient or the soft-token-path (main-loss) gradient from reaching the acoustic encoder degraded performance, suggesting the two paths provide complementary supervision to the shared acoustic encoder. Doubling the learning rate to 4e-4 unexpectedly yielded an RMSE of 0.350 on the first run, but subsequent runs did not reproduce it, indicating an unstable configuration rather than a better one.

For CASA, even runs using the same seed produced test RMSEs ranging from 0.357 to 0.365, indicating nondeterminism in the training pipeline. Most distinct seeds produced comparable results, except for the poorer run with seed 6066. The 0.358 result in Table 1 was obtained from the first run of the configuration and was selected on the development set, rather than from the best repeated run. Full run-level results are available in the repository.

Although efficiency remains a primary goal, LLMs offer substantial value for ASA beyond their parameter scale. In particular, they can adapt to different assessment tasks and provide speech content assessment and feedback without task-specific training. As an illustration, CASA’s Qwen3.5-2B is used for few-shot content validation, prompting it at inference time to judge whether an answer addresses the given question. Replacing the original questions with an unrelated question about nuclear reactors, the LLM flags 99.9% of responses as not addressing the question; pairing responses with real questions from different parts yields 97.3%. This is also practical: judgments took just under 0.1 s per response. Such validation requires reasoning over the semantic relationship between question and answer, which is not available to an acoustic-only grader. The same judge could also be prompted to flag A2-level content even from ASR transcripts, though inconsistently; this is left as a direction for future work.

Conclusion

CASA is a model architecture that explicitly separates speech delivery from speech content in ASA. CASA achieves SOTA-level performance on the S&I test set while using approximately half the inference parameters of comparable LLM-based systems. The analyses demonstrate the complementary roles of the acoustic and content branches, identify task-specific modeling and verbatim transcription as promising directions for further improvement, and show that the LLM branch can also support training-free content validation. By releasing the efficient, general-purpose architecture and experimental results, the authors aim to support reproducible research and further investigation of acoustic and linguistic information in ASA.

Improvements for AI systems

Based on the paper, here are the specific improvements I can make to AI systems and what the improved system can do:

  1. Dual-branch architecture with explicit acoustic/content separation: I can implement a two-branch system where one branch processes raw audio through a frozen Whisper encoder with LoRA adapters, and the other processes ASR transcripts through a frozen LLM with LoRA. This allows independent optimization and analysis of each modality.

  2. Tolerance-based auxiliary loss for acoustic branch: I can add an auxiliary prediction head on the acoustic branch that predicts the score, but only penalize errors beyond ±1 score point with a small weight (0.1). This prevents the acoustic branch from dominating while still providing useful gradient signal.

  3. Soft token integration into LLM: Instead of concatenating acoustic features directly, I can map the acoustic summary to 4 compact soft tokens that are prepended to the LLM input sequence, allowing the LLM to interpret acoustic information flexibly.

  4. Detached text-based auxiliary prediction: I can render the acoustic branch's prediction as text (e.g., acoustic cefr estimate: 4.5) and include it in the LLM prompt, providing interpretable context without backpropagating through the text.

  5. Chunked processing with frame-level tagging: For long responses, I can split audio into 30-second chunks, concatenate the frame vectors, and add task-specific and segment-specific learned embeddings to each frame before passing through a Transformer aggregator with RoPE positional encoding.

  6. Training-free content validation via LLM reasoning: I can prompt the LLM at inference time to judge whether an answer addresses the given question, using few-shot examples. This works without any fine-tuning for the validation task.

  7. Three simple fluency features: I can compute duration, silence ratio, and speech rate from audio and ASR output, and include these as text in the LLM prompt to ground the assessment in basic delivery metrics.

  • Achieve state-of-the-art speaking assessment performance (RMSE 0.358) with approximately half the inference parameters of comparable LLM-based systems, making it more deployable in resource-constrained environments.

  • Provide interpretable separation between speech delivery and content: The system can independently analyze how something is said (acoustic branch) versus what is said (content branch), enabling targeted feedback for learners.

  • Flag off-topic responses in real-time: The LLM can detect when a learner's answer doesn't address the given question (99.9% detection rate for unrelated questions, 97.3% for cross-part mismatches) in under 0.1 seconds per response, without any additional training.

  • Adapt to new assessment corpora without structural changes: The general-purpose architecture uses only three handcrafted features and can be applied to different speaking assessment tasks by simply changing the task prompt and rubric.

  • Provide robust performance across repeated runs: The auxiliary loss provides consistent small improvements (mean RMSE reduction of 0.004), and the system maintains SOTA-level performance across different random seeds.

  • Handle varying task formats: The system can score short personal questions, open-ended opinion tasks, process descriptions, and other formats by using task embeddings and part-specific prompts, though performance is better on open-ended tasks.

  • Support training-free content quality assessment: Beyond topic relevance, the LLM can be prompted to flag low-level content (e.g., A2-level responses) from ASR transcripts, though this capability is currently inconsistent and requires further development.

Abstract

Research on automatic speaking assessment (ASA) has increasingly adopted multimodal speech large language models to assess learners' speaking performance. However, existing studies provide limited analysis of how acoustic and content information contribute to predictions and how stable the resulting performance is. We propose CASA, a simpler architecture combining Whisper-medium and Qwen3.5-2B that achieves state-of-the-art performance while providing a more interpretable separation between speech delivery and content. On the Speak & Improve Corpus 2025, CASA achieves a root mean square error (RMSE) of 0.358, improving on the previous best RMSE while using approximately half the estimated inference parameters. The general-purpose architecture is designed for adaptation to other ASA corpora without structural changes and relies on three handcrafted fluency features. Through ablations and repeated runs, we analyze the individual and complementary contributions of acoustic and content information, examine performance variability, and demonstrate the potential of large language model reasoning for training-free content validation.

Related papers