Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos

summary

Video file (mp4)

The gist

This paper details an evaluation of a course-grounded chatbot designed for STEM lecture videos, emphasizing strict adherence to provided source material.

In short

The episode explores 'Cite or Decline,' a chatbot designed to answer questions based strictly on STEM lecture videos. The system utilizes LLM-generated summaries and traditional search methods to achieve high citation coverage (75%). While students found the timestamped citations highly useful, the system was noted to lack features like practice-question generation.

Key concepts

Soft Prior
The chatbot uses LLM-generated chapter summaries as a 'soft prior' to guide retrieval. This method boosts the ranking of relevant video segments from chapters that relate to a specific query, allowing for efficient searching across the entire corpus.
Strict Course Grounding
This feature enforces course isolation, preventing systems from answering questions based on general knowledge. It ensures the chatbot only provides answers strictly within the context of the provided lecture materials, addressing gaps in previous systems.
Citation Support
The system generates timestamped citations, which is critical because it tells students exactly where in the video they need to look. This allows users to verify if a citation actually supports the answer provided by the chatbot.

Terminology used across episodes

This episode discusses

The paper

Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos · Read on arXiv

S M Masrur Ahmed, Jaspal Subhlok

University of Houston

Recorded lecture videos, often enhanced with search and summarization features, are a standard study resource. However, students cannot easily ask course specific questions or verify answers against an instructor's lecture. We report a semester-long deployment of VideoPoints platform with a retrieval-augmented chatbot that answers from course lecture materials and returns timestamped citations. The chatbot retrieves only from the active course, uses chapter summaries to guide transcript ranking, and returns clickable timestamped citations. Students used it for quick lookups and exam review. Across 833 messages, 70.5% included citations, none crossed a course boundary, and when no lecture evidence matched, the chatbot usually declined rather than answering. Among the users, citations were the most consistently useful feature, while practice-question generation was the strongest unmet request. We also evaluated the design on the real-world test split of EduVidQA, a public multimodal benchmark for lecture-video question answering. Our design improved correct-lecture retrieval by 6.3 percentage points over dense-only retrieval. Together, the results show that effective deployment depends on course isolation, supported citations, and alignment with students' study practices.

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos".

Jane: The paper was written by S M Masrur Ahmed and Jaspal Subhlok from University of Houston.

Tom: Stay tuned as we take you through the paper and discuss its implications.

The Core Mechanism: Tom: So, Jane, let’s talk about the core mechanism of this chatbot. It's not just a standard RAG system; they added some clever layers to make it more reliable.

Jane: The key is how they use chapter summaries as an anchor, which helps guide the retrieval process when we ask a question about a specific topic in the lecture video.

Lu: They’re using these LLM-generated summaries—which were created using GPT-four point one Mini—to act as a "soft prior" to boost the relevance of chunks from chapters that seem most related to our query.

Meng: From an implementation perspective, this soft prior is much more efficient than forcing a hard filter because it allows the entire corpus to be checked while simply boosting the ranking of relevant segments.

Lalam: I see how this significantly improves the user experience, Lalam feels; instead of dumping a huge pile of irrelevant video chunks on us, we get a cleaner list that matches what we are looking for in our notes.

Tom: And while they use these summaries to guide the search, the system still pulls chunks from all-MiniLM-L6-v2 embeddings and applies BM25 to make sure it' gets a comprehensive view of the information.

Jane: It’s a brilliant blend of semantic understanding and traditional lexical search, making sure that "Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos" can handle both conceptual questions and direct keyword searches.

Lu: This whole process allows the system to generate timestamped citations, which is critical because it tells the student exactly where in the video they need to look.

Meng: The efficiency of having both a summary-guided approach and a traditional retrieval method means that "Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos" is very much designed to operate at scale for real classroom use.

Improvements over Existing Systems: Tom: Now, Jane, the paper points out some real gaps in current systems, and this chatbot seems to address them specifically by enforcing course isolation.

Jane: That's a huge improvement; previous systems often lacked this strict grounding, meaning they could answer questions based on general knowledge instead of the specific context of "Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos."

Lu: They didn're not just assuming the system works; they are rigorously testing it against a public benchmark called EduVidQA, which is a very smart way to prove performance.

Meng: The results from that benchmark are quite strong, showing that their design improved correct-lecture retrieval by six point three percentage points over dense-only retrieval methods.

Lalam: I think this shows the practical impact of structure; when we allow the system to look outside of its assigned course content, it degrades quickly, so enforcing those boundaries is key to creating a reliable educational tool.

Tom: It’s not just about performance on the real-world data either; they are also doing a diagnostic audit on two hundred twenty-four unique messages to check if the citations actually support the answer.

Jane: That's important because, as they found, having a citation present doesn't guarantee that citation is relevant or that it supports what’s written in the response from "Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos."

Lu: The fact that they are measuring this "citation support" separately from the overall usage rate gives us a much clearer picture of where the system is actually succeeding and where it’s struggling.

Meng: From an operational standpoint, this means the instructor can trust that "Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos" is delivering answers that are traceable to the materials provided.

Student Perceptions and Needs: Tom: We have a lot of data on how students used this system, which is really interesting, especially seeing how they interacted with it.

Jane: The survey results show that students perceived the timestamped citations as highly useful, rating them around four point five five out of five for usefulness and helpful for exams in "Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos."

Lu: It’s fascinating that they observed that students often interacted with the chatbot using keyword fragments rather than full questions, which is a behavioral pattern we should pay attention to.

Meng: The fact that students clicked on those citations three hundred thirteen times suggests they are actively engaging with the video content, even if the logs don't show how much of the video they actually watched.

Lalam: But one thing that really stands out is what students felt was missing—they rated practice-question generation as a major unmet need, which "Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos" didn's designed to do.

Tom: It seems like the system works great for specific knowledge lookups and exam review, but it falls short when students want the AI to take on an agentic role like grading a quiz.

Jane: The data also shows that "Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos" is performing best when the student asks a direct question, rather than issuing an imperative command.

Lu: This contrast really suggests that our current retrieval-based RAG approach is excellent for recall but may not be powerful enough to handle complex task management.

Meng: We need to make sure that we’re clear about what this AI can do and what it cannot do, and this finding is a very important reminder for developers implementing these systems.

Wrap-up: Tom: Well, Jane, as we wrap up our discussion of "Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos," it’s clear that this was a major step forward in making AI a reliable partner for academic study.

Jane: The system successfully enforced strict course grounding while providing high citation coverage—seventy point five percent of the messages had citations—which builds huge trust with users.

Lu: I think the implication is that we have moved past just asking "can an AI answer this?" to "can an AI answer this *correctly and within these specific constraints*?"

Meng: And from a practical standpoint, it’s showing us exactly where our current RAG architectures need refinement, especially regarding task-oriented requests.

Lalam: Lalam feels that the future needs to see how we bridge the gap between grounded Q andA and the student's desire for instructor-controlled generative activities.

Tom: Thank you all for sharing your insights on "Cite or Decline: A Strict Course-Grounded Chatbot for STEM Lecture Videos."

Jane: It’s been a great discussion, everyone. We hope this provides some practical inspiration for future AI tools to build better education.

More episodes

← Home