Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts

arXiv:2604.03127 · cs.CL, cs.AI · Submitted 2026-04-03 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts".

Jane: The gist The authors present a domain-adapted Retrieval-Augmented Generation (RAG) pipeline for annotating pedagogical dialogue acts,

Tom: First, who's behind it and why it matters.

Paper summary: Tom: Wrapping up this discussion on "Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts," it seems the authors have successfully shown that adapting the retrieval system is a very effective way to boost annotation quality without having to fine-tune the main generative model.

Jane: They demonstrated that by focusing on domain adaptation in the retriever, they can achieve Cohen’s kappa scores up to zero point seven four three on Eedi, which is substantially better than what was possible with no retrieval at all <ref:2604.03127#pg1>.

Lu: The key finding they emphasized is that utterance-level indexing coupled with parent chunk retrieval proves to be superior because it preserves the specific label signal while still providing that rich conversational context needed for classification.

Meng: So, in simple terms, this paper suggests that for complex tasks involving human interaction analysis, you should build a smart search mechanism tailored precisely to your domain’s nuances.

Lalam: It means the system doesn't need a massive overhaul of the core model; it just needs better access to the right labeled examples when it's making a decision.

Tom: The authors are pointing toward future work that includes extending this idea to other tutoring domains and using active learning to let the index improve iteratively as it learns more.

Jane: Ultimately, they’re suggesting that for AI systems dealing with subtle pedagogical moves, the most reliable way forward is integrating domain-specific retrieval into the workflow.

Conclusion: Tom: So, we've been talking about how you can use retrieval to help label tutoring conversations, and now we’re looking at the end of this paper, "Domain-Adapted Retrieval for In-Context Annotation of Pedagogical Dialogue Acts."

Jane: Yeah, it wraps up by showing that by making the retrieval system specific to the domain—the tutoring stuff—you get much better results in terms of how accurately you label those teaching moves.

Lu: The authors are really pushing the idea that adapting the retriever, not just tweaking the main language model itself, is where most of this improvement comes from.

Meng: From a practical standpoint, they show that this method works across different AI backbones, which means it’s more flexible for us when we start working with new models.

Lalam: For me, seeing the results on those dialogue datasets confirms that giving the right context to an AI really helps it understand the specific function of what's happening in a tutoring session.

Tom: It seems like their main conclusion is that utterance-level indexing, where you look at each individual line of dialogue and pull in some surrounding context, beats just chunk-level indexing.

Jane: Exactly. They show that even with different AI models, like those GPT ones they tested, the utterance-level approach gives a bigger jump in accuracy on Eedi data compared to just looking at the whole chunk.

Lu: It opens up possibilities for tailoring these retrieval systems much more precisely to different subjects or tutoring styles down the road.

Meng: That means we could potentially build these specialized search tools for very niche educational areas without needing to retrain a massive new language model every time.

Lalam: I see it as making the AI's understanding of "teaching" much more nuanced and context-aware, which is a big step for how we design these systems.

Tom: It’s really about moving away from one-size-fits-all methods and toward a retrieval strategy that actually understands the specific patterns of expert tutoring.

Jane: And while they show some really high numbers on accuracy, they also point out that this system still relies on those external LLMs to do the final classification step.

Lu: That’s a fair caveat; it shows the power of the retrieval component, but we still have to manage where we get our generative engine from.

Tom: So, for listeners who just want to know what this means, it's that you can make AI annotation much more reliable by making its search engine smarter and more domain-specific.

Cornell University

cs.CL, cs.AI

Submitted: 2026-04-03

Updated: 2026-04-03

Comments: 20 pages, 20 tables, 4 figures

Code: https://github.com/SumnerLab/TalkMoves

License: http://creativecommons.org/licenses/by-nc-sa/4.0/

Importance score: 92/100

The gist: The gist The authors present a domain-adapted Retrieval-Augmented Generation (RAG) pipeline for annotating pedagogical dialogue acts, achieving high Cohen’s κ scores by adapting the retrieval

Key concepts

Domain-Adapted Embeddings
The researchers fine-tuned a sentence embedding model using Multiple Negatives Ranking Loss on tutoring dialogue data. This process adjusts the semantic space so that utterances with similar teaching functions cluster together, regardless of how they are phrased in text. This adaptation makes the retrieval system highly effective for specialized tutoring tasks.
Utterance-Level Indexing
Instead of indexing entire document chunks, the method indexes each individual utterance separately. When a query is made, it retrieves the specific parent chunk containing that utterance. This granular approach helps preserve label-specific signals and provides richer conversational context for the final classification.
In-Context Learning (ICL)
The pipeline uses a frozen, general-purpose Large Language Model to perform classification through in-context learning. It is provided with retrieved examples (labeled demonstrations) and the target utterance's context directly in the prompt. This allows the LLM to classify the action based on these provided examples without requiring model fine-tuning.
Semantic Chunking
The corpus is divided into semantically coherent chunks that maintain label consistency while respecting session boundaries. Boundaries are identified by measuring similarity between overlapping context windows, ensuring that retrieved segments are meaningful and relevant to the specific pedagogical function being annotated.

Terminology

Summary

The gist The authors present a domain-adapted Retrieval-Augmented Generation (RAG) pipeline for annotating pedagogical dialogue acts, achieving high Cohen’s κ scores by adapting the retrieval component rather than fine-tuning the generative model.

How it works

The proposed method employs a three-stage retrieval-augmented annotation pipeline that combines domain adapted embedding, utterance level indexing with parent chunk context, and codebook grounded in-context learning with a frozen LLM The task is to predict the label for each tutor utterance by constructing a retrieval index from the corpus, dynamically building context windows around target utterances, and passing both local context and retrieved examples to a frozen LLM for in-context classification All domain adaptation resides in the retriever that makes the pipeline portable across LLM backbones and reusable as new models become available

The pipeline involves several key steps, including:

  1. Fine-tuning a lightweight sentence embedding model on tutoring dialogue corpora with Multiple Negatives Ranking Loss (MNRL), adapting the representation space so that utterances serving similar pedagogical functions cluster together regardless of surface lexical variation

  2. Indexing dialogues at the utterance level to retrieve labeled few-shot demonstrations

  3. Presenting retrieved examples as few-shot demonstrations alongside the annotation codebook to a general purpose LLM, which performs classification through in context learning (ICL)

Key Components and Strategies

The methodology is built upon several specific technical choices designed to address the challenges of abstract and contextual tutoring moves

Semantic Chunking & Indexing:

The corpus is partitioned into semantically coherent chunks that preserve label homogeneity while respecting session boundaries Semantic boundaries are detected by computing the smoothed cosine similarity between overlapping context windows The boundary threshold τ is determined by leveraging sparse ground truth annotations and sweeping τ over [0.3, 1.0) to maximize F1 on this boundary classification problem

Domain-Adapted Embeddings:

The sentence transformer model is fine-tuned on labeled utterances from the TalkMoves and Eedi training sets combined using MNRL This fine-tuning improves semantic separation of tutoring dialogue labels in the embedding space, and it is trained on both classroom and dyadic chat data to capture pedagogical function across interaction formats

RAG Index Construction:

Two types of FAISS indexes are constructed: a chunk-level index using Cohere embeddings and two BGE-based indexes The utterance-level index embeds each labeled utterance individually, retrieving the parent chunk at query time

Experimental Results and Findings

The experiments were conducted across two real tutoring dialogue datasets (TalkMoves and Eedi) and three LLM backbones (GPT-5.2, Claude Sonnet 4.6, Qwen3-32b)

Performance Gains:

The best configuration achieved Cohen’s κ of 0.526-0.580 on TalkMoves and 0.659-0.743 on Eedi, substantially outperforming no retrieval baselines (κ = 0.275-0.413 and 0.160-0.410)

Retrieval Precision:

The top-1 label match rate under RAG FINETUNED UTT reaches 62.0% on TalkMoves and 73.1% on Eedi, compared to 39.7% and 52.9% for RAG NO FINETUNE This confirms that retrieval precision drives annotation quality

Driver of Improvement:

An ablation study reveals that utterance-level indexing, rather than embedding quality alone, is the primary driver of these gains, with top-1 label match rates improving from 39.7% to 62.0% on TalkMoves and 52.9% to 73.1% on Eedi under domain adapted retrieval

Confound Isolation:

The comparison between RAG NO FINETUNE and RAG FINETUNED UTT shows that the large jump to RAG FINETUNED UTT (+0.085 on TalkMoves, +0.129 on Eedi for GPT-5.2) comes predominantly from switching the index granularity from chunks to individual utterances

Conclusion

Utterance-level indexing with parent chunk retrieval consistently outperforms chunk-level indexing by preserving label specific signal while providing rich conversational context Domain adaptation on the retriever side accounts for the majority of this gain, and the improvement pattern holds across models of varying capability Utterance-level indexing with parent chunk retrieval consistently outperforms chunk-level indexing by preserving label specific signal while providing rich conversational context Future directions include extending to nonmathematical tutoring domains, integrating active learning for iterative index improvement, multi turn annotation that captures dialogue level pedagogical strategies, and alternative training objectives that enforce inter label separation without post hoc corrections

Ethics Statement

We analyze de identified tutoring dialogue datasets in accordance with our Institutional Review Board (IRB)-approved protocol [BLIND FOR REVIEW] We follow IRB approved procedures for data storage and access and use these data solely to understand and improve dialogue annotation workflows Automated annotation of tutoring moves is intended to augment rather than replace expert human judgment We caution against deploying automated labels as ground truth without such verification, particularly for high stakes decisions such as teacher evaluation or student assessment The confidence scores produced by our pipeline support a human in the loop workflow in which uncertain predictions are flagged for manual review Our system relies on commercial and open weight LLMs accessed via API, which introduces dependence on external services and associated cost, latency, and data handling considerations Finally the label taxonomy was developed for mathematics tutoring in English and may not transfer to other languages, subject areas, or cultural contexts without careful adaptation and validation

References

Rishabh Agarwal et al. Many-shot in-context learning Bakhtawar Ahtisham et al. Ai annotation orchestration: Evaluating llm verifiers to improve the quality of llm annotations, 2025 Aliki Anagnostopoulou et al. Human and llm-based assessment of teaching acts in expert-led explanatory dialogues, In Proceedings of the 6th Workshop on Computational Approaches to Discourse, pp. 166–181, 2025 Parishad BehnamGhader et al. Llm2vec: Large language models are secretly powerful text encoders, In Proceedings of the 1st Conference on Language Modeling, 2024 Sinchana Ramakanth Bhat et al. Rethinking chunk size for long document retrieval: A multi-dataset analysis, arXiv preprint arXiv:2505.21700, 2025 Huiyao Chen et al. Retrieval-style in-context learning for few-shot hierarchical text classification, 2024 Cheng-Han Chiang and Hung-yi Lee. Can large language models be an alternative to human evaluations?, In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 15607–15631, Toronto, Canada, July 2023 Association for Computational Linguistics. doi: 10.18653/v1/2023.acl-long.870 Qingxiu Dong et al.

Improvements for AI systems

  1. textbfDomain-Adapted Retrieval for Annotation Quality Improvement (RAG Pipeline): The core system can be improved by implementing a three-stage retrieval-augmented annotation pipeline that combines domainadapted embedding, utterance-level indexing with parent chunk context, and codebook-grounded in-context learning with a frozen LLM. This system achieves Cohen’s κ of 0.526-0.580 on TalkMoves and 0.659-0.743 on Eedi, substantially outperforming no-retrieval baselines (κ = 0.275-0.413 and 0.160-41).

  2. textbfUtterance-Level Indexing Strategy: Precise Signal Preservation: The system can be improved by utilizing an utterance-level indexing strategy that preserves label-specific signal during search while returning parent chunks as few-shot demonstrations. This strategy is shown to drive the largest gains, with top-1 label match rates improving from 39.7% to 62.0% on TalkMoves and 52.9% to 73.1% on Eedi under domain-adapted retrieval.

  3. textbfContextual Prompt Construction: Dynamic Context Window Generation: The system should dynamically construct a context window for each target utterance by expanding backward and forward from the target, stopping at positions where "sim(i) < τ or at session boundaries, up to a maximum of 10 utterances in each direction. This ensures the LLM receives full surrounding dialogue context" tailored to the specific pedagogical function of the target utterance.

  4. textbfEmbedding Adaptation: Fine-Tuning for Domain Specificity: The system can be enhanced by fine-tuning a lightweight embedding model, specifically BGE-large-en-v1.5 on labeled utterances from the TalkMoves and Eedi training sets combined, using Multiple Negatives Ranking Loss (MNRL) to adapt the representation space so that utterances serving similar pedagogical functions cluster together regardless of surface lexical variation.

  5. textbfConfidence Score Triage: Reliable Human-in-the-Loop Workflow: The system can be improved by leveraging confidence scores, as they are shown to shift the distribution rightward, allowing for a triage mechanism where RAG FINETUNED UTT substantially improves the reliability of high-confidence predictions across all models, with Sonnet 4.6 achieving 0.896 Cohen’s κ on its confident subset.

Abstract

Automated annotation of pedagogical dialogue is a high-stakes task where LLMs often fail without sufficient domain grounding. We present a domain-adapted RAG pipeline for tutoring move annotation. Rather than fine-tuning the generative model, we adapt retrieval by fine-tuning a lightweight embedding model on tutoring corpora and indexing dialogues at the utterance level to retrieve labeled few-shot demonstrations. Evaluated across two real tutoring dialogue datasets (TalkMoves and Eedi) and three LLM backbones (GPT-5.2, Claude Sonnet 4.6, Qwen3-32b), our best configuration achieves Cohen's κ of 0.526-0.580 on TalkMoves and 0.659-0.743 on Eedi, substantially outperforming no-retrieval baselines (κ= 0.275 - 0.413 and 0.160 - 0.410). An ablation study reveals that utterance-level indexing, rather than embedding quality alone, is the primary driver of these gains, with top-1 label match rates improving from 39.7% to 62.0% on TalkMoves and 52.9% to 73.1% on Eedi under domain-adapted retrieval. Retrieval also corrects systematic label biases present in zero-shot prompting and yields the largest improvements for rare and context-dependent labels. These findings suggest that adapting the retrieval component alone is a practical and effective path toward expert-level pedagogical dialogue annotation while keeping the generative model frozen.

Sources

Related papers