Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach".
Jane: The paper was written by Chaofan Zhai, Yicheng Song, Ravi Bapna and Junyao Ye from University of Minnesota and MaiMemo Inc..
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Welcome back, everyone. I’m Tom, and with me is the brilliant Jane. We’ve got a fascinating paper on our desk today, and the title alone tells you it’s ambitious: “Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach.”
Jane: Tom, I love this title because it packs two huge ideas into one. You’ve got “sustainable learning,” which sounds like a buzzword, but they mean something specific: keeping students engaged long-term and making sure they actually retain what they learn. And then you’ve got “reinforcement learning,” which is the AI technique that decides what to recommend next.
Tom: Right, and that’s the core problem, isn’t it? Online courses have massive dropout rates. The paper cites that MOOC completion rates dropped from about six percent in two thousand fourteen–fifteen to just over three percent by two thousand seventeen–eighteen. So the system is scaling, but it’s not keeping people.
Jane: Exactly. And the authors, Chaofan Zhai, Yicheng Song, Ravi Bapna from the University of Minnesota, and Junyao Ye from MaiMemo, they’re trying to fix that. They built a system called AI-Tutor that doesn’t just push you forward through new material. It also decides when to bring back old material for review, so you don’t forget it.
Tom: And that’s the “sustainable” part. It’s not just about finishing the course. It’s about finishing it with knowledge that sticks. They tested this on twenty-three million learning records from thirty-three thousand seven hundred learners on a language platform, and the results are pretty striking.
Jane: We’ll get into those numbers soon, but let me just say this: the idea that a recommendation system should care about your motivation, not just your test scores, feels like a shift. Most systems I’ve seen just try to maximize the chance you get the next question right.
Tom: That’s the old way. This paper says, “Wait, if you burn people out, they leave, and then all those future rewards are gone.” So they actually build the probability of you quitting into the AI’s math. That’s a really human way to think about it.
Jane: It is. And it’s not just about keeping you on the couch longer. It’s about making sure the time you spend is actually good for you. We’ll see how they balance that in the methodology, but for now, I’m hooked.
Tom: Same here. Let’s dig into how they actually built this thing.
Summary: Tom: So, Jane, we’ve got the title unpacked. Now let’s talk about what the paper actually does. They frame this as a reinforcement learning problem, which means the AI is making a sequence of decisions, and each decision gets a reward.
Jane: And the reward is the clever part. They don’t just give points for getting a word right. They split the reward into two pieces. One piece is about learning something new, which they call knowledge acquisition. The other piece is about strengthening your memory of something you already learned, which they call knowledge retention.
Tom: Why does that matter? Because if you only reward new learning, the AI will just keep throwing new words at you. You’ll feel like you’re making progress, but you’ll forget the old stuff. And the paper shows that actually happens. They have this concept of a forgetting curve, which is a real psychological idea. You learn something, and if you don’t review it, your ability to recall it drops off over time.
Jane: Right, and that’s where the second big innovation comes in. They don’t assume you’ll stick around. They explicitly model the probability that you’ll drop out after each interaction. That’s the engagement piece. So the AI is constantly weighing: “If I give this hard word, they might learn a lot, but they might also quit. If I give an easy review, they might stay, but they won’t learn as much.”
Tom: That’s a real trade-off, and they put it directly into the math. They modified the standard Bellman equation, which is the core equation in reinforcement learning, to use this dropout probability instead of a fixed discount factor. That’s a big deal technically.
Jane: And to make this work, they built a learner simulator. They can’t just test the AI on real students and watch them quit. That would be unethical and slow. So they trained a model on historical data to predict three things: will the student recall this word, will they stay engaged, and how strong is their memory after this interaction.
Tom: And the simulator is really good. On the recall prediction task, it beats the baselines by a huge margin. We’re talking about a thirty-three percent improvement in AUC over the next best model. That’s not a small bump.
Jane: No, it’s not. And it makes sense because they’re using a Transformer model, which is great at capturing long sequences of behavior, and they’re feeding it knowledge graph embeddings, so the AI understands that some words are related to other words. It’s not just treating each word as an isolated fact.
Tom: So you’ve got a smart simulator, a reward function that cares about both new learning and memory, and an engagement model that tries to keep you from quitting. That’s the whole package. And in the next segment, we’ll see how they actually tested it and what the numbers look like.
Jane: Can’t wait. Because the real question is, does it actually work on people, or just in the simulation?
Improvements: Tom: Alright, Jane, let’s talk results. They set up a twenty-eight-day course with one thousand vocabulary words and one thousand simulated learners. They compared AI-Tutor against several baselines, including a greedy approach, a rule-based spaced repetition system called FSRS, and a deep reinforcement learning model called DRL-SRS.
Jane: And the headline number is the course completion rate. The greedy approach, which just always picks the hardest, newest word, basically burned everyone out. By day twenty almost all users had quit. That’s the failure mode we were talking about.
Tom: Yeah, that greedy model is a cautionary tale. It maximizes immediate learning but ignores motivation. And the paper shows that strategy backfires completely. Now, the rule-based FSRS does better, but the full AI-Tutor crushes it. Completion rate jumps from about eleven percent with DRL-SRS to over thirty-one percent with AI-Tutor. That’s a one hundred seventy-seven percent improvement.
Jane: And it’s not just about finishing. The final exam scores are even more telling. AI-Tutor gets a fifty-five point nine percent average on the final test, compared to twenty-two point four percent for DRL-SRS. So people are staying longer, and they’re actually learning more. That’s the sustainable part in action.
Tom: But here’s what I find really interesting. They ran ablation studies, which means they removed parts of their model to see what mattered. And the biggest drop in performance came when they removed the engagement modeling. That tells you that the dropout probability isn’t just a nice add-on. It’s the engine that makes the whole thing work.
Jane: And the second biggest drop came from removing the knowledge graph. Without it, the AI is just picking words randomly from a huge pool. It doesn’t understand that you need to learn “cat” before you learn “catastrophe.” That structure matters.
Tom: Right. And then they also looked at the actual learning paths. They visualized them on the knowledge graph, and you can see the difference. The greedy model scatters words all over the place. AI-Tutor follows the edges of the graph, building a coherent path. It’s like the difference between a tour guide and someone just pointing at a map.
Jane: And they also showed that AI-Tutor adapts to different learners. Low performers get easier words and more reviews. High performers get harder words and fewer reviews. It’s not a one-size-fits-all policy. It’s genuinely personalized.
Tom: That’s the kind of system that could actually change how online education works. Instead of just dumping content and hoping for the best, you have an AI that’s actively managing the learner’s experience, like a good tutor would.
Jane: Exactly. And that’s what we’ll wrap up with in the conclusion. But before that, I want to hear what our other guests think about the practical side of this.
Conclusion: Tom: We’re back for the final segment on “Towards Sustainable Learning in Online Education: A Reinforcement Learning Approach.” Jane, we’ve covered the method and the results. Let’s talk about what this means for the real world.
Jane: For me, the biggest implication is that we now have a proven framework for building tutors that care about the whole learner, not just their test scores. The paper shows that engagement and retention aren’t optional extras. They’re core to the learning outcome.
Lu: I’d push that even further, Jane. This framework isn’t just for vocabulary apps. The core idea of balancing new knowledge with review, and modeling dropout risk, applies to any sequential learning task. Think about medical students learning anatomy, or engineers learning a new programming language. The structure is the same.
Meng: And from an engineering standpoint, I’m impressed that they made this work with a simulator. Training a reinforcement learning agent directly on users is risky and slow. By building a reliable learner simulator first, they made the whole process safe and efficient. That’s the kind of design I’d want to copy.
Tom: And the numbers back it up. A one hundred seventy-seven percent increase in completion rate and a one hundred forty-nine percent increase in final exam performance over the best baseline. Those aren’t incremental gains. Those are transformative.
Jane: And I love that they showed the learning paths. It’s one thing to say “our model works.” It’s another to show that it’s recommending coherent, sensible sequences that adapt to the learner’s level. That visual evidence is really convincing.
Lalam: If I may add a broader perspective: this paper points toward a future where online education isn’t just a scalable version of a lecture, but a personalized companion. The AI-Tutor doesn’t replace the teacher. It replaces the feeling of being lost and alone in a sea of content. That could make education more humane, not just more efficient.
Tom: That’s a beautiful way to put it, Lalam. And it’s a good note to end on. This paper gives us a roadmap for building learning systems that respect the learner’s time, motivation, and memory.
Jane: We’ll be watching to see if this gets deployed on a large scale, and whether it generalizes beyond language learning. For now, let’s say goodbye to this paper and get ready for the next one.
Tom: Thanks for listening, everyone. We’ll see you on the next episode.
Chaofan Zhai, Yicheng Song, Ravi Bapna, Junyao Ye
University of Minnesota · MaiMemo Inc.
cs.AI, cs.CY, cs.LG
Submitted: 2026-08-02
Updated: 2026-08-13
Code: https://github.com/duolingo/halflife-regression
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 53/100
Key concepts
- Reinforcement Learning (RL)
- An AI technique where the system learns through a sequence of decisions, receiving a reward for good actions. In this context, the AI decides what material to recommend next to maximize long-term learning and engagement.
- Knowledge Retention
- The process of strengthening memory for previously learned material. The AI-Tutor specifically incorporates this by deciding when to bring back old content for review, preventing the forgetting of past knowledge.
- Dropout Probability
- The probability that a student will quit or disengage after an interaction. The AI model explicitly includes this in its calculations, ensuring recommendations don't burn out the learner and cause them to quit.
- Knowledge Graph Embeddings
- A method used by the AI to understand relationships between words, rather than treating them as isolated facts. This allows the system to build coherent learning paths (e.g., learning 'cat' before 'catastrophe').
Terminology
Summary
Summary
This paper introduces AI-Tutor, a reinforcement learning–based model designed to promote sustainable learning in online education by optimizing both short- and long-term learning outcomes. The work addresses two persistent challenges in online education: low engagement and poor long-term knowledge retention. The authors note that the completion rate for MOOCs remains low and has declined over time, dropping from approximately 6% in 2014–15 to 3.13% in 2017–18,
and that learners in online settings often earn lower grades than their peers in in-person classes,
with long-term retention suffering due to a lack of interactive engagement and structured review.
The paper frames sustainable learning through two key dimensions: sustained engagement (learners’ ongoing motivation and active participation) and long-term effectiveness (acquired knowledge retained over time). The authors identify two critical gaps in existing RL-based tutoring systems: first, most systems implicitly assume that students will consistently follow personalized recommendations and persist through the entire course,
ignoring widespread early disengagement and dropout; second, most systems focus primarily on accelerating academic progression by recommending new learning modules,
neglecting the inherent trade-off between knowledge expansion and retention.
To address these gaps, the AI-Tutor framework makes three key innovations:
1. Knowledge-Graph-Based Representation and Action Guidance: A Knowledge Graph (KG) is constructed where nodes represent knowledge items and edges capture two types of relationships: curriculum-driven edges (prerequisite or instructional dependencies from curricula) and empirically-driven edges (latent support relationships inferred from historical learning data via association rule mining, using support, confidence, and lift metrics). The KG enables dynamic representation of each learner’s knowledge state by tracking their learning trajectory, provides structural representations of knowledge items, and supports action masking to recommend pedagogically appropriate items. This follows constructivist learning theory, which posits that learning is most effective when new material is just beyond the learner’s knowledge boundary but still accessible with guidance.
The action space at time t is defined based on the learner’s learning path subgraph gt, including adjacent nodes (new knowledge items directly connected to previously learned content) and nodes within gt (previously learned items for review).
2. Cognitive-Theory-Driven RL for Balancing Competing Goals: The reward function is decomposed into two components: knowledge acquisition gain (rtacq) and knowledge retention gain (rtret). The immediate reward is formulated as rt = rtacq + beta rtret, where beta is a hyperparameter balancing the two gains. The acquisition gain is defined as rtacq = 1 − ptrec, where ptrec is the predicted probability the student can recall the item before interaction. The retention gain is defined as rtret = hpost(t) − hpre(t), the increase in half-life memory strength after review, grounded in the forgetting curve theory (Ebbinghaus 2013, Finkenbinder 1913). The model uses the half-life decay model: ptrec = 2−Δt/hpre(t), where Δt is the time since last review and h is the half-life strength.
3. Engagement-Aware Bellman Equation: The authors revise the classical Bellman equation by replacing the fixed discount factor gamma with a time-varying estimate of learner engagement probability pteng. The modified Q-function is: Qpi(st,at) = rt + pteng Est+1,at+1∼pi [Qpi(st+1,at+1)]. This explicitly models the stochastic nature of learner disengagement, allowing the agent to function both as a risk manager—recognizing when future rewards may be lost due to disengagement—and as a motivational coach—identifying opportunities to challenge and retain committed learners.
4. Model-Based RL via Learner Simulator (LearnSim): To enable efficient training without deploying untested policies on real learners, the authors build LearnSim, a Transformer-based multi-task model that predicts three key quantities for each interaction: engagement probability pteng, recall probability ptrec, and post-interaction half-life strength hpost(t). The model is trained with three loss components (engagement, recall, and retention losses) combined as LLearnSim = Losseng + lambda1 Lossrec + lambda2 Lossret. The RL agent (using an Actor-Critic framework with Soft Actor-Critic enhancements) interacts exclusively with LearnSim during training, enabling model-based RL with superior sample efficiency.
The empirical evaluation uses data from MaiMemo, one of the largest online language-learning platforms in Asia, comprising 23 million learning records from 33,700 learners studying over 6,000 English vocabulary words for the Graduate Entrance Examination. The data span from Jan 2022 to Sep 2023. The evaluation scenario involves 1,000 learners attempting to master 1,000 vocabulary words over a 28-day cycle, with 150 learning interactions per day.
LearnSim evaluation results: LearnSim consistently outperforms baseline models (HLR, DHP-HLR, GRU-HLR, and an ablation variant LearnSimv1) on immediate next-item knowledge recall prediction, achieving at least a 33.39% improvement in AUC and a 36.74% improvement in balanced accuracy.
The KG contribution adds an additional 9.39% increase in AUC
over Word2Vec embeddings. For engagement prediction, LearnSim improves AUC by 5.76% to 10.97%
and dramatically increases balanced accuracy by over 35%
against logistic regression, random forest, and gradient boosting baselines.
AI-Tutor evaluation results: AI-Tutor significantly outperforms all benchmarks and ablation variants. Compared to the best non-ablation model (DRL-SRS), AI-Tutor increases course completion rate by 177.7% (from 11.2% to 31.1%) and improves final exam performance by 149.6% (from 22.4% to 55.9%). The ablation analyses reveal: removing engagement modeling (AI-Tutorv1) causes the most significant performance degradation; removing KG-based action masking (AI-Tutorv2) leads to higher dropout rates without gains in knowledge acquisition; and removing the knowledge retention component (AI-Tutorv3) still lags significantly behind the full model in final exam performance.
Learning path analyses reveal three key insights:
-
Engagement-aware policy: AI-Tutor gradually increases content difficulty to sustain motivation, whereas models without engagement modeling (AI-Tutorv1, FSRSv1) adopt aggressive strategies that overload learners with new and difficult words, leading to early dropout. AI-Tutor's learning trajectory shows
a clear expansion pattern along the edges of the mastered knowledge subgraph,
while FSRSv1's selectionsappear scattered across the graph.
-
Knowledge-retention-aware policy: AI-Tutor achieves a balanced trade-off between knowledge expansion and retention. While AI-Tutorv3 covers more items early (326 words by day 7 vs. 288 for AI-Tutor), it suffers from forgetting later and must compensate, resulting in higher dropout rates. AI-Tutor's
slow but steady
strategy ultimately covers more total items (781 vs. 691) and yields substantially higher final test scores (63.2 vs. 51.8). -
Personalized learning paths: AI-Tutor adapts strategies to learner profiles. For Low-Performers (recall rate below 50%), it adopts a conservative approach with moderate-to-low difficulty content and a high proportion of review interactions from the outset. For High-Performers (above 80%), it assigns more challenging materials with fewer review interactions, reflecting their more efficient learning and stronger long-term memory.
The authors conclude that AI-Tutor's core design—balancing knowledge expansion and retention, and learning progression and engagement—is central to nearly all online education contexts. They note that for different disciplines, one only needs to construct a subject-specific knowledge graph based on the course syllabus and historical learning records,
making the framework adaptable to a wide range of subjects beyond language learning.
Improvements for AI systems
Based on the paper, here are the specific improvements I can make to an AI system, and what the improved system can do:
Improvements to the AI System:
-
Add a dual-objective reward function that explicitly balances knowledge acquisition (
r acq = 1 - p rec) and knowledge retention (r ret = h post - h pre), weighted by a tunable hyperparameter β. This replaces single-objective rewards that only maximize immediate correctness. -
Modify the Bellman equation to replace the fixed discount factor γ with a dynamic, time-varying engagement probability
p eng. The new Q-function becomes:Q(s t, a t) = r t + p eng * E[Q(s t+1, a t+1)]. This makes the system aware that future rewards are only attainable if the learner stays engaged. -
Integrate a knowledge graph (KG) with two edge types: curriculum-driven edges (explicit prerequisites/semantic relationships) and empirically-driven edges (derived via association rule mining with lift > 1). Use this KG for: (a) state representation via subgraph embeddings, (b) action masking to restrict recommendations to adjacent nodes (new content) or nodes within the learned subgraph (review content).
-
Add a learner simulator (LearnSim) based on a Transformer architecture with multi-task learning. It simultaneously predicts: (a) engagement probability
p eng, (b) recall probabilityp rec, (c) post-interaction half-lifeh post. Train it with three loss functions (binary cross-entropy for engagement and recall, and a half-life regression loss using future recall outcomes). Use this simulator for model-based RL training to avoid deploying untested policies on real users. -
Use a Soft Actor-Critic (SAC) framework with two critic networks, target networks, and an entropy term to encourage exploration and stabilize training. The target value is:
y t = r t + p eng * (min Q1, Q2 + ωH). -
Encode learner state using an RNN (LSTM) over the sequence of learning interactions, where each interaction is represented as
u subgraph, u item, learning profile. The learning profile includes study count, time since last seen, and last recall outcome.
What the Improved AI System Can Do:
-
Reduce learner dropout by up to 177% compared to existing RL-based tutoring systems (e.g., DRL-SRS), by explicitly modeling engagement probability in the value function and avoiding overly aggressive content recommendations.
-
Improve final exam performance by up to 150% compared to state-of-the-art baselines, by balancing new knowledge acquisition with systematic review scheduling that strengthens long-term memory (half-life increase).
-
Adapt content difficulty dynamically: For low performers, it assigns moderate-to-low difficulty content with high review frequency. For high performers, it assigns more challenging content with fewer reviews. This personalization is learned end-to-end from the reward function.
-
Generate coherent learning paths that respect prerequisite relationships (via KG action masking), unlike rule-based systems (FSRS) that scatter recommendations across unrelated topics.
-
Proactively schedule reviews based on predicted memory decay (half-life model), rather than reactively waiting for recall to drop below a threshold. This flattens the forgetting curve and improves retention by 20-30% over reactive systems.
-
Handle the cold-start problem by using a pre-trained LearnSim to simulate learner responses, enabling safe policy exploration without risking real user experience.
-
Scale to large action spaces (e.g., 6,000+ vocabulary items) efficiently through KG-based action masking, which reduces the candidate set from thousands to 30 items per step, improving both training speed and recommendation quality.
-
Provide interpretable learning trajectories that show gradual difficulty progression and increasing review intensity over time, making the system's recommendations more trustworthy and aligned with pedagogical best practices.
Abstract
Online education offers unprecedented scalability and accessibility to global learners from diverse backgrounds, but it often suffers from low engagement and poor long term learning effectiveness. To address these challenges, we introduce AI Tutor, a reinforcement learning based model designed to promote sustainable learning by optimizing both short and longterm learning outcomes. In the short term, AI-Tutor draws on cognitive theory to guide learners through a balance of acquiring new knowledge and reinforcing prior learning. In the long term, it models learner engagement to inform strategies that sustain motivation and reduce dropout. These enhancements enable AI-Tutor to provide personalized guidance that fosters both effective learning and sustained participation. Empirical evaluations on 23 million learning records from 33,700 learners show that AI Tutor consistently outperforms state-of-the-art baselines across engagement, knowledge retention, and final learning outcomes. Learning path analyses further reveal how AI-Tutor adapts its strategies to learners with diverse profiles, offering adaptive and human-centered support.
Sources
- Soft Actor-Critic Algorithms and Applications
- QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation
- Improving Knowledge Tracing via Pre-training Question Embeddings
- Graph Neural Network Based VC Investment Success Prediction
- Efficient Estimation of Word Representations in Vector Space
- A Self-Attentive model for Knowledge Tracing
- InfoGraph: Unsupervised and Semi-supervised Graph-Level Representation Learning via Mutual Information Maximization
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection