EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning".
Tom: Detailed Research Summary: EduGage:
Jane: First, who's behind it and why it matters.
Paper summary: Tom: Hey everyone, welcome back to the show! We're diving into some really interesting new research today with a paper called "EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning."
Jane: I'm so excited to talk about this because it tackles how we can actually measure when people are truly paying attention while they're learning through videos.
Lu: What really catches my eye about this work is the sheer breadth of data they are integrating, covering everything from physiological signals like ECG and EEG to motion tracking via IMU and even eye-tracking data <ref:2605.01238#pg0>. It suggests a very holistic view of what engagement actually looks like during learning activities.
Meng: From an engineering standpoint, I’m curious how they managed to synchronize all those different types of continuous sensor streams together reliably for this study <ref:2605.01238#pg0>? We need to know the practical feasibility of collecting that much data in a real-world setting.
Lalam: I think the potential here is huge because if we can accurately measure these momentary states, it could inform how adaptive learning systems adjust content in real time <ref:2605.01238#pg2>. It moves us closer to truly personal education experiences where the material changes based on the learner's actual state.
Tom: Exactly, Lalam! And this paper sets up a dataset called EduGage, which is a major piece of stuff because it lets other researchers build on their findings without having to collect all that complex hardware data from scratch <ref:2605.01238#pg0>. It's like giving everyone the raw material.
Jane: That dataset includes not just the sensor signals but also probe-aligned engagement labels gathered from participants who did self-reports, which gives them a solid ground truth to train models against <ref:2605.01238#pg0>. It’s a very comprehensive package for anyone interested in this area.
Lu: The way they set up the user study with sixteen participants and then evaluate their multimodal modeling approach across different sensing modalities is quite rigorous, especially when comparing it against sensor-free methods <ref:2605.01238#pg0>. It shows a careful comparison of what works best in this context.
Meng: I'm interested in the performance numbers they reported; they mentioned achieving an MAE of zero point eight one and an eighty-three point seven five percent within-one accuracy, which sounds like a pretty strong result when you’re dealing with complex physiological signals <ref:2605.01238#pg0>. Does this suggest the model is robust enough for actual deployment in an educational setting?
Lalam: Those accuracy metrics are impressive because they show that their integrated multimodal model beats out established baselines, including deep temporal models and traditional LLM-based approaches for characterization <ref:2605.01238#pg0>. That comparison really validates the approach of using multiple data streams together.
Paper summary: Tom: So, to recap for our listeners, this paper introduces EduGage as a multimodal dataset and benchmark designed to estimate learner engagement through physiological and motion signals collected during self-guided video learning <ref:2605.01238#pg0>. It claims their integrated model performs well against many existing methods <ref:2605.01238#pg0>.
Jane: And they highlight that this system is valuable because wearable sensing can provide a continuous, low-burden basis for estimating engagement at the segment level in video-based learning <ref:2605.01238#pg2>. It’s about making physiological perspectives practical for fine-grained measurement.
Lu: The authors emphasize that engagement involves cognitive, emotional, and behavioral dimensions, and this sensing helps reflect all of those aspects simultaneously <ref:2605.01238#pg1>. This holistic view is what makes the data so rich for modeling engagement.
Meng: What I’m wondering is how this translates practically; if we deploy a system based on these findings, would the computational overhead be manageable, or would we need specialized hardware for everyone to use it? <ref:2605.01238#pg0>
Lalam: The implication for culture is that education can become far more responsive; instead of a one-size-fits-all pace, we could adapt the video content moment by moment based on what the learner's physical signals are indicating <ref:2605.01238#pg2>.
Tom: That’s a big picture idea, Lalam! And this paper really gives us that foundation to start building these adaptive systems, which is where the real impact could happen <ref:2605.01238#pg1>.
Jane: So, looking at the overall results from "EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning," we see that their multimodal modeling approach is competitive with other sophisticated techniques <ref:2605.01238#pg0>. It successfully integrates various sensor inputs to produce a reliable assessment of engagement.
Lu: The authors also pointed out the importance of self-regulation, linking deep engagement with stronger learning outcomes, which is a key part of cognitive engagement they are trying to capture <ref:2605.01238#pg1>. This connects the measurement directly to educational effectiveness.
Meng: If we think about the practical impact on startups or educational platforms, having a benchmark like this means we don't have to reinvent the wheel when trying to measure learner attention using wearable tech <ref:2605.01238#pg0>. It provides a clear path forward for development.
Lalam: It gives us a standardized way to measure engagement that goes beyond simple clicks or quiz scores; it measures the actual human experience of learning unfolding in real-time <ref:2605.01238#pg2>. This level of detail is what makes the potential for improving educational culture so significant.
Paper summary: Tom: So, to wrap up this segment, we've looked at how "EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning" uses diverse sensor data to create a powerful tool for understanding learner engagement <ref:2605.01238#pg0>. It really sets the stage for future work in adaptive education.
Jane: And the title itself tells us that this work is providing both a dataset and a benchmark, which is crucial for anyone trying to advance this field <ref:2605.01238#pg0>. The implications are that we can start moving toward systems that react to the learner's physiological state during video learning sessions.
Lu: The core contribution here seems to be validating the feasibility of using wearable sensing in educational settings for fine-grained engagement measurement <ref:2605.01238#pg2>. It’s a practical application of complex signal processing we see in other scientific domains.
Meng: I still want to stress the engineering reality; while the results are strong, integrating PPG, ECG, EEG, and IMU data requires careful handling of noise and latency to ensure those momentary assessments are actually useful in a learning environment <ref:2605.01238#pg0>. That technical challenge remains important.
Lalam: I think the biggest impact is shifting the focus from just *what* learners know to *how* they are processing that knowledge internally during the video consumption process <ref:2605.01238#pg2>. That kind of insight can genuinely reshape how we design educational content for maximum retention.
Tom: That’s a fantastic point, Lalam; it moves us past surface-level metrics and into the actual learning mechanism <ref:2605.01238#pg1>. This paper provides the technical blueprint for getting there using multimodal data.
Jane: So, to bring this segment to a close, we've discussed how the "EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning" paper uses extensive sensor data to create a benchmark that helps researchers understand learner engagement across different modalities <ref:2605.01238#pg0>. It opens the door for more responsive learning systems.
Lu: The potential for future work seems vast because they have established this robust dataset, giving subsequent researchers a solid starting point to explore even deeper correlations between these physiological states and specific learning outcomes <ref:2605.01238#pg2>.
Meng: We'll keep an eye on how the community responds to this benchmark, because having clear metrics helps us decide where the next practical engineering efforts should be focused for real-world deployment <ref:2605.01238#pg0>.
Lalam: It’s exciting because it gives us a way to build systems that truly understand learner states, which is a major step toward making education more personalized and effective <ref:2605.01238#pg2>.
Conclusion: Tom: So, we've been looking at how this paper introduces EduGage as both a dataset and a benchmark for measuring engagement in video learning sessions.
Jane: Right, and the authors are really showing how they integrated various sensor data streams to get this assessment done.
Lu: It’s fascinating to see the breadth of signals they're using, spanning from physiological responses like EEG to motion tracking via IMUs.
Meng: I mean, it’s interesting how they managed to synchronize all that continuous data together so the model could analyze it moment by moment.
Lalam: This work really gets at what engagement means in a deeper way than just watching a video or taking a quiz score.
Tom: Exactly, and that's why this paper is important because it gives us a better tool to look at how people are actually learning in those self-guided video settings.
Jane: It opens up the door for figuring out not just what learners know, but how they are processing that knowledge internally while they're watching.
Lu: That’s where the real potential lies; understanding those internal states could lead to much smarter ways of designing educational content.
Meng: From an engineering standpoint, having this benchmark means we don't have to start from zero when trying to build systems that react in real time.
Lalam: And I think that ability to adapt based on a learner’s moment-to-moment state is what could truly improve how we approach education and learning culture.
Georgia Institute of Technology
cs.HC, cs.CV
Submitted: 2026-05-02
Updated: 2026-10-06
Comments: Accepted at IMWUT
DOI: 10.1145/3857993
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 92/100
The gist: This research introduces a comprehensive system and dataset designed for the fine-grained, multimodal estimation of learner engagement within video-based learning scenarios.
Key concepts
- Multimodal Sensing
- This involves collecting data from several different types of sensors simultaneously. For this study, it included physiological signals (like heart rate) and motion signals (like body movement). Combining these different data sources provides a richer picture of how a learner is engaging with the learning material.
- Physiological Signals
- These are biological measurements taken from the body to gauge internal states. The study used signals such as Photoplethysmography (PPG) and Electroencephalography (EEG). These help researchers understand the learner's emotional or cognitive state during video lessons.
- EduGage Dataset
- This is a comprehensive collection of data used for testing and future research. It includes synchronized sensor readings, engagement labels from participant self-reports, and video context. It serves as a benchmark resource for anyone wanting to study how sensor data can measure learner engagement.
- Momentary Assessment
- This refers to assessing a learner's engagement at specific points during a learning activity. Instead of measuring overall performance, the system aims to capture brief, real-time indicators of how engaged the learner is at any given moment within the video content.
Terminology
Summary
This research introduces a comprehensive system and dataset designed for the fine-grained, multimodal estimation of learner engagement within video-based learning scenarios. The core contribution is the development and rigorous evaluation of a model capable of characterizing learner engagement by integrating physiological and motion signals collected from wearable and camera-based sensing devices.
Methodology and Data Collection:
The study employed a sophisticated multimodal sensing approach to capture rich behavioral data. Researchers utilized a diverse array of sensors to collect simultaneous physiological and motion signals, specifically including:
-
Physiological Signals: Photoplethysmography (PPG), Electrocardiogram (ECG), Electrodermal Activity (EDA), and Electroencephalography (EEG).
-
Motion/Contextual Signals: Inertial Measurement Units (IMU) for motion tracking, heart rate monitoring, and temperature sensing.
-
Visual Data: Eye-tracking data was also incorporated to capture visual attention patterns.
These continuous sensor streams were collected during a user study involving 16 participants who engaged in specific learning tasks within a video-based learning environment. Crucially, the study employed a mixed-methods approach, combining objective sensor data with subjective self-reports: participants provided repeated, in-situ self-reports of their engagement through brief probes throughout the learning process.
Model Development and Performance:
The primary objective of the research was to develop and evaluate a robust system for engagement estimation by comparing various sensing modalities. The resulting multimodal modeling approach demonstrated significant efficacy across participant-based cross-validation. The performance metrics achieved were highly competitive, yielding an MAE (Mean Absolute Error) of 0.81, an 83.75% within-1 accuracy, a 73.93% binary accuracy, and a 68.45% binary Macro-F1 score. These results indicate that the integrated multimodal model substantially outperforms established baselines, including sensor-free methods, statistical models, deep temporal models, foundation models (e.g., LLMs), and traditional LLM-based approaches for engagement characterization.
The EduGage Dataset:
To ensure reproducibility and facilitate further research in this domain, the study released the EduGage dataset. This dataset is a cornerstone of the work, providing synchronized multimodal sensor signals alongside critical contextual metadata. The EduGage dataset includes:
-
Synchronized multimodal sensor signals (PPG, ECG, EDA, EEG, IMU data).
-
Probe-aligned momentary engagement labels derived from participant self-reports.
-
Video metadata pertaining to the learning scenario.
-
Comprehension quizzes relevant to the study topics.
-
Study materials themselves.
Conclusion and Significance:
In essence, this paper presents both a novel, high-performing method for estimating learner engagement using fine-grained sensor data and a valuable benchmark resource (EduGage) that enables future researchers to replicate and build upon these findings in self-guided learning contexts. The work successfully validates the feasibility and effectiveness of multimodal modeling as a powerful tool for characterizing learner engagement in digital education.
Improvements for AI systems
Here are specific improvements to AI systems based on the findings of the EduGage study, categorized by application:
)1. Improved Real-Time Personalized Adaptive Learning Systems (The Proactive Tutor
)
An improved AI system can move beyond simple content delivery or session-level feedback to provide moment-to-moment cognitive scaffolding.
Improvement Specific AI Capability
:---:---
Continuous, Fine-Grained Engagement Monitoring The system can continuously monitor physiological and behavioral signals (PPG, EDA, IMU) to detect subtle shifts in attention difficulty (e.g., sustained increase in arousal or decrease in gaze stability). Instead of waiting for a quiz or session end, the AI identifies the exact time window where engagement drops below a learned threshold.
Context-Aware Intervention Triggering When engagement drops (as measured by the model's prediction), the system uses contextual metadata (video progress) to determine if it should prompt a specific action: 1) Suggest a short break, 2) Prompt reflection on the preceding segment, or 3) Offer adaptive support such as re-explaining a confusing concept.
Modality-Aware Reliability Weighting The AI's fusion module learns which sensor is currently most reliable for that specific learner and context (e.g., in low-light conditions, it might weigh gaze tracking higher; if the learner is highly agitated, it might prioritize EDA). This prevents the system from making a hard
decision based on noisy data.
)2. Enhanced Content Design and Instructional Material Analysis (The Curriculum Optimizer
)
An improved AI system can be used by content creators to proactively design videos that maximize engagement and minimize cognitive load for diverse learners.
Improvement Specific AI Capability
:---:---
Predictive Video Difficulty Mapping An AI analyzes video segments (using the EduGage dataset structure) and predicts which segments are most likely to cause a drop in attention difficulty for different types of learners (e.g., those with high baseline arousal). This allows creators to flag high-difficulty
moments proactively.
Cross-Modal Instructional Refinement The system correlates self-reported attention difficulty with learning outcomes (using the correlation found in Section 3.1). If a segment is predicted to cause high difficulty but the AI knows it historically leads to poor learning gains, it flags that specific part of the video for mandatory revision (e.g., adding an example or simplifying the explanation).
Optimal Modality Requirements Analysis The system performs the ablation study (Section 5.5) in reverse: it identifies which combination of sensors is most predictive for a given learning domain. It can then advise instructors on the minimum necessary sensor suite required to reliably detect disengagement in specific video types, optimizing resource allocation.
)3. Robust and Deployable Sensor Fusion Architectures (The Resilient Sensor Hub
)
An improved AI system can be more resilient to real-world deployment challenges (noise, missing data, heterogeneity).
Improvement Specific AI Capability
:---:---
Dynamic Gating for Heterogeneous Data Streams The core fusion mechanism learns modality-specific contribution weights. This allows the system to gracefully degrade: if a wearable device fails or is removed (e.g., the user takes off the ring), the model automatically re-weight its remaining sensors (like EEG or IMU) to maintain a functional, albeit less accurate, prediction rather than crashing.
Foundation Model Adaptation for Domain Transfer By leveraging frozen foundation models (PulsePPG, NeuroLM) and training only the final fusion layer on the specific engagement task, the system can be rapidly adapted to new learning domains (e.g., switching from engineering videos to medical lectures) with minimal retraining data, drastically reducing deployment time compared to training a model from scratch.
Low-Burden Signal Selection Strategy Based on Section 5.5 analysis, the AI can dynamically select the lightweight
sensor configuration required for the current context (e.g., prioritizing gaze tracking + heart rate over full EEG) to minimize participant burden while maintaining sufficient predictive utility for that specific learning task.
Abstract
Engagement, which links to attentional, emotional, and cognitive dimensions, plays an important role in learning. In online and video-based learning environments, learners often need to regulate their own interactions with instructional materials. Measuring and reflecting on engagement can therefore support both learners and adaptive learning systems. In this study, we use wearable and camera-based sensing devices to collect physiological and motion signals, including PPG, ECG, EDA, EEG, IMU, heart rate, temperature, and eye-tracking data, to estimate learner engagement. We conducted a user study with 16 participants in a video-based learning scenario, where participants completed learning tasks and provided repeated in-situ engagement-probe ratings on a 1-5 scale, reporting how difficult it was to pay attention during the preceding minute. We establish a benchmark for engagement estimation, compare different sensing modalities, and further analyze the feasibility and effectiveness of multimodal modeling for characterizing learner engagement. Across participant-based cross-validation, the modality-aware reference model achieves an MAE of 0.80, 83.18% within-1 accuracy, 70.96% binary accuracy, and 62.90% binary Macro-F1, outperforming sensor-free, statistical, deep temporal, foundation-model, and LLM-based baselines. Our results suggest that fine-grained engagement estimation is feasible but inherently noisy, and that the preferred sensing configuration depends on the target metric and practical sensing burden. We release the EduGage dataset as the primary contribution of this work, including synchronized multimodal sensor signals, probe-aligned engagement-probe ratings, video metadata, quizzes, and study materials, to support reproducible research on fine-grained sensor-based engagement modeling in self-guided learning.
Sources
- Gated Multimodal Units for Information Fusion
- Foundation Models Defining A New Era In Sensor-based Human Activity Recognition: A Survey And Outlook
- NeuroLM: A Universal Multi-task Foundation Model for Bridging the Gap between Language and EEG Signals
- An Electrocardiogram Foundation Model Built on over 10 Million Recordings with External Evaluation across Multiple Domains
- ConSensus: Multi-Agent Collaboration for Multimodal Sensing
Related papers
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support
- An LLM-Native Psychometric Instrument Reveals a Self-Report--Behavior Gap Across 25 Models