When Demonstrations Fail: Diagnosing the Limits of In-Context Learning in Large Audio-Language Models with Progressive Cue Removal
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "When Demonstrations Fail".
Jane: While Large Audio-Language Models (LALMs) have shown degraded instruction-following capabilities, their ability to infer task patterns from in-context examples under audio conditioning remains unstudied.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So, we're talking about this paper called "When Demonstrations Fail: Diagnosing the Limits of In-Context Learning in Large Audio-Language Models with Progressive Cue Removal." It’s really digging into why these models might be getting good at mimicking instructions but missing the mark on actual audio tasks when they see examples.
Jane: Exactly, Tom; it’s about testing how much we can trust those examples when we take away the explicit guidance step by step. Basically, the authors are showing that simply giving a model more examples isn't always a good thing if the underlying instructions aren't strong enough.
Lu: The title itself suggests they are systematically peeling back layers of instruction to see where the models break down, which is fascinating because it moves past just seeing if an example works or not.
Meng: From an engineering standpoint, that progressive removal sounds like a really rigorous way to isolate whether the model is learning formatting rules or actually understanding the physics of the audio task.
Lalam: I think this research points toward a needed evolution in how we train and prompt these models, because if they can’t handle this kind of structured evaluation, it suggests our current alignment methods are missing something fundamental about cross-modal learning.
Tom: Right, that's the core idea—we're testing the limits of in-context learning by systematically reducing textual help to see what happens when that help gets weaker.
The paper's summary: Jane: What the paper summarizes is a three-stage framework called ALICE, which they use to evaluate Large Audio-Language Models’ ability to learn from examples under audio conditioning, and they find a consistent pattern across all models.
Lu: They set up this evaluation by starting with Stage one where everything is fully explained, then moving to Stage two where the explicit format constraints are removed from the examples, and finally Stage three where only audio inputs and formatted outputs are left.
Tom: And what they found is that in-context demonstrations consistently improve how well a model follows the output format, but they don't help it get better at doing the actual speech processing task itself, and sometimes even make things worse.
Meng: That’s interesting because it means we can't just rely on seeing good examples to fix performance issues; the model isn't truly internalizing the core concepts from those demonstrations yet.
Lalam: It really highlights an asymmetry where surface-level pattern matching gets a boost, but the deep semantic understanding needed for complex audio tasks stays underdeveloped when relying only on in-context examples without full instruction.
Jane: Precisely; they’re showing that models can pick up on how to structure an output based on demonstrations, but they haven't managed to grasp the actual meaning of what those audio inputs are supposed to achieve.
The paper's improvements: Tom: The paper suggests a few ways we can use this finding, primarily focusing on building better evaluation methods like ALICE itself and perhaps developing new training objectives that force the models to focus on the task objective rather than just surface formatting.
Lu: They propose implementing a multi-stage reasoning architecture, which would mirror their evaluation setup, allowing the model to process complex instructions in stages instead of all at once.
Meng: From a practical standpoint, that means we could design training regimens where the AI learns to decouple what constitutes a correct format from what constitutes a successful task execution.
Lalam: I think integrating explicit cross-modal alignment training objectives would be helpful because it would directly address that gap they found by forcing the model to connect audio features with desired output structures during its learning process.
Jane: That makes sense; if we can train them to prioritize the task logic over just matching a template, then we might see better results in real-world applications like those audio understanding tasks they tested, such as ASR and SER.
Conclusion: Tom: So to wrap up, the main point of "When Demonstrations Fail: Diagnosing the Limits of In-Context Learning in Large Audio-Language Models with Progressive Cue Removal" is that while demonstrations help format compliance, they don't reliably boost the core task performance under audio conditioning.
Jane: That means we have to be careful not to assume that a good set of examples will automatically make a model better at the actual work, especially when dealing with complex audio inputs.
Lu: The implication is that we need methods that probe deeper than just observing what the paper suggests—we need architectures that can handle these multi-stage reasoning challenges we saw in ALICE.
Meng: For us engineers, this means our next step should be to focus on building systems where the model learns to distinguish between necessary output structure and the actual audio feature extraction required for success.
Lalam: I think this paper gives us a clear direction on how to design better training signals that specifically target that cross-modal grounding they discussed in the paper.
National Taiwan University
cs.SD, cs.AI, cs.CL, eess.AS
Submitted: 2026-03-20
Updated: 2026-09-28
Comments: Accepted to IEEE SLT 2026
Code: https://github.com/yenting-biao/ALICE
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 78/100
The gist: While Large Audio-Language Models (LALMs) have shown degraded instruction-following capabilities, their ability to infer task patterns from in-context examples under audio conditioning remains
Key concepts
- ALICE Framework
- A three-stage testing method used to systematically reduce textual guidance in examples. It starts with full instructions, removes format rules, and finally tests models on audio alone to see what they can infer from demonstrations.
- Format Compliance Rate (FCR)
- Measures how well a model's output follows the specific structural rules or constraints given in the demonstration examples. Models consistently improve FCR with demonstrations, even when core task performance stays flat or drops.
- Cross-modal ICL
- The ability of models to learn a task objective by jointly reasoning over both audio inputs and corresponding textual outputs across multiple examples. This is tested as the most difficult stage of the evaluation.
Terminology
Summary
While Large Audio-Language Models (LALMs) have shown degraded instruction-following capabilities, their ability to infer task patterns from in-context examples under audio conditioning remains unstudied. This research introduces ALICE, a three-stage framework that systematically reduces textual guidance to evaluate LALMs’ in-context learning ability under audio conditioning, revealing a consistent asymmetry where demonstrations improve format compliance but fail to improve or degrade core task performance.
The gist
In-context demonstrations reliably improve format compliance but fail to improve, and often degrade, the core task performance when evaluating LALMs under audio conditioning.
ALICE: The Three-Stage Evaluation Framework
The ALICE framework is a systematic three-stage process designed to progressively reduce textual guidance in in-context demonstrations to isolate the contribution of textual cues versus cross-modal semantic grounding. This setup keeps the audio inputs and demonstration outputs unchanged while systematically withholding explicit instructions.
-
In Stage 1 (Explicit Constraint), each demonstration example and test sample contains both a task description and an explicit format constraint instruction alongside the audio input. This stage serves as a
strong upper-bound condition,
where textual instructions fully specify both the task objective and how the output should be formatted, meaning any observed ICL effect reflects learning on top of already complete textual guidance. -
In Stage 2 (Implicit Constraint), the explicit format constraint instructions are removed from both the demonstration examples and test samples while other parts remain intact. This tests if models can
infer formatting conventions from example outputs alone, without being explicitly told the rules.
-
In Stage 3 (Audio Only), task descriptions are further removed, providing only audio inputs paired with correctly formatted outputs as demonstrations. This is the most challenging stage, requiring models to
jointly infer task objectives and formatting patterns from the relationship between audio inputs and textual outputs across demonstrations,
constituting astrict test of cross-modal ICL.
Experimental Setup and Models
The evaluation adopts a subset of Speech-IFEval as the testbed, covering four audio understanding tasks: Automatic Speech Recognition (ASR), Speech Emotion Recognition (SER), Gender Recognition (GR), and Massive Multi-Task Audio Understanding (MMAU). Two constraint categories are used: Closed-Ended Questions (CEQ) and Chain-of-Thought (CoT).
The researchers evaluated five open-source LALMs, including Qwen2-Audio, DeSTA2.5-Audio, BLSP-Emo, Qwen2.5-Omni, and Phi-4 Multimodal. A proprietary model, Gemini 2.5 Flash, was also included under two configurations: dynamic thinking enabled (“On”) and thinking entirely disabled (“Off”). The number of in-context examples (k) was varied from 1 to 8 to observe convergence trends.
Evaluation Metrics
The ICL ability is assessed along two complementary dimensions: Format Compliance Rate (FCR) and Task Performance.
(i) Format Compliance Rate (FCR):
This measures whether a model’s response adheres to the prescribed output constraints. The results show that in-context demonstrations reliably improve format compliance but fail to improve, and often degrade, the core task performance.
For example, FCR is near-saturated for instruction-strong models in Stage 1 zero-shot (e.g., Gemini 2.5 Flash (On): Task Group A: 98.68%; Task Group B: 99.03%), but one-shot examples still yield large compliance jumps for others, suggesting demonstrations help instantiate the intended output schema.
(ii) Task Performance:
This involves standard metrics for each audio understanding task, such as Word Error Rate (WER) for ASR and accuracy for SER, GR, and MMAU. The paper finds that High format compliance is independent of high task performance,
indicating that LALMs can infer output format from demonstrations but can not leverage such examples to internalize the core concepts of accomplishing speech-processing tasks.
Key Findings on Asymmetry
The study uncovers a consistent asymmetry across all stages and LALMs: in-context demonstrations only improve format compliance but do not improve, and often degrade, core speech task performance, even when reasoning traces are provided.
-
In Stage 1 (Explicit Constraint),
textual instructions fully specify both the task objective and how the output should be formatted.
-
In Stage 2 (Implicit Constraint),
Format inference from demonstrations alone is unreliable for LALMs,
as FCR consistently decreases across all models when explicit constraint instructions are removed. -
In Stage 3 (Audio Only), relying solely on audio inputs and formatted outputs to infer the task objective is challenging, with CEQ task performance remaining "substantially below Stage 2, even with multiple shots.
Improvements for AI systems
As a fastidious and diligent AI researcher, I have thoroughly analyzed the findings of ALICE. The core insight is that LALMs excel at pattern matching (format compliance) but fail at true cross-modal semantic grounding required for complex task inference under audio conditioning.
Here are the specific improvements I would implement in AI systems based on these findings:
-
Enhance Cross-Modal Semantic Grounding Modules:
-
Implement a Multi-Stage Reasoning Architecture (Inspired by ALICE):
-
Develop Constraint-Aware Demonstration Selection and Ordering Strategies:
-
Integrate Explicit Cross-Modal Alignment Training Objectives:
Specific Capabilities of the Improved AI System:
-
The improved system will be capable of performing complex audio reasoning tasks (like specialized ASR or SER) not just by matching surface patterns, but by correctly mapping auditory features (e.g., specific acoustic patterns indicating a speaker's emotion or a particular sound event) to high-level semantic concepts and required output structures.
-
The system will exhibit robust performance when the task requires inferring objectives from purely audio-conditioned examples (Stage 3 in ALICE), moving beyond reliance on textual cues embedded in reasoning traces, which are shown to be insufficient for true cross-modal integration.
-
The system will possess superior
instruction following
reliability under complex constraints because it learns to distinguish between necessary format patterns and task objectives, rather than conflating the two (addressing the finding that good instruction-following does not guarantee robust ICL deduction). -
The system will be better equipped to handle novel audio inputs or slightly varied contexts because its inference mechanism is less brittle; it relies on integrated semantic understanding rather than fragile surface-level pattern matching derived from demonstrations.
Sources
- Qwen2-Audio Technical Report
- DeSTA2.5-Audio: Toward General-Purpose Large Audio Language Model with Self-Generated Cross-Modal Alignment
- Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
- ISA-Bench: Benchmarking Instruction Sensitivity for Large Audio Language Models
- SAKE: Towards Editing Auditory Attribute Knowledge of Large Audio-Language Models
- Instruction-Following Evaluation for Large Language Models
- Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities
- Hallucinations in Neural Automatic Speech Recognition: Identifying Errors and Hallucinatory Models
Related papers
- Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
- Spectral Masking and Interpolation Attack (SMIA): A Black-box Adversarial Attack against Voice Authentication and Anti-Spoofing Systems
- AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
- SoundWeaver: Compositional Warm-Starting for Text-to-Audio Diffusion Serving
- WASIL: In-the-Wild Arabic Spoken Interactions with LLMs
- Efficient Audiovisual Speech Processing via MUTUD: Multimodal Training and Unimodal Deployment