Jointly Predicting Courses and Grades Using a Transformer-Based Model
Paul Savala
St. Edward's University
cs.AI
Submitted: 2026-08-13
Updated: 2026-08-14
Code: https://github.com/paulsavala/TRACE-grade-prediction
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
Importance score: 75/100
The gist: This paper introduces a TRansformer for Academic Course-grade Estimation (TRACE), a novel model that jointly predicts both the set of courses a student will take and their corresponding grades for an
Terminology
Summary
This paper introduces a TRansformer for Academic Course-grade Estimation (TRACE), a novel model that jointly predicts both the set of courses a student will take and their corresponding grades for an upcoming semester. The authors argue that existing predictive models in learning analytics often treat student academic history as a simple sequence, overlooking the concurrent nature of courses taken within a semester, which can lead to inaccurate performance predictions, particularly for students with heavy or challenging course loads.
The study is guided by two primary research questions: (RQ1) To what extent does a Transformer model that encodes semester-level concurrency improve grade prediction accuracy compared to models that treat course history as a flat sequence? (RQ2) Does the joint prediction of courses and grades lead to a significant improvement in grade prediction accuracy compared to an architecture that predicts grades alone?
The model encodes courses on a per-semester basis to capture the effects of course concurrency, utilizing positional encodings where all courses taken in the same semester receive the same temporal vector. This design enforces a permutation-invariant representation within each semester, reflecting the unordered nature of concurrent enrollments. The approach also utilizes a novel loss function combining course-set prediction with grade prediction, specifically using Kullback-Leibler (KL) divergence for course-set prediction and mean squared error (MSE) for grade regression. This set-based multi-task objective avoids order artifacts induced by token-level cross entropy.
The dataset consisted of 5326 students enrolled in 360 different courses across 48 majors and 18 semesters, spanning from Fall 2014 through Fall 2023. Students were split into training and testing sets, consisting of 90% and 10% of all students, respectively. The model was trained on a single NVIDIA RTX 3090 Ti for 20 epochs using the AdamW optimizer with a cosine annealing warm restart learning rate scheduler.
The results demonstrate significant improvements in prediction quality. The TRACE model achieved a mean absolute error of 0.1339, a 46.4% reduction compared to the 'OnlyGradesTransformer''s MAE of 0.2496, and a mean squared error 3.5 times lower. This large difference demonstrates that the Transformer model architecture on its own is not sufficient to guarantee strong performance, and that predicting courses taken in addition to the grades in those courses leads to significant improvements in prediction quality.
Compared to other baseline models, TRACE drastically outperformed the XGBoost model, with the improvement largely because tree-based models are not naturally suited to sequential learning tasks. Compared to the two LSTM models (Encoder-Decoder LSTM with attention and Unidirectional LSTM), the Transformer model showed a significant improvement, with an MSE approximately 30% lower and a MAE approximately 15% lower. Compared to the graph neural network, TRACE had an MSE/MAE 15%-20% lower. The GNN assumes locality in the graph, while the Transformer model makes no such locality assumption, instead learning which historical observations are relevant through attention.
The authors also conducted a feature ablation analysis. When student major was removed from the input, the model showed marginally improved performance (MSE: 0.0360 vs. 0.0392; MAE: 0.1327 vs. 0.1339), suggesting that major provides no additional predictive signal beyond course history and may introduce slight redundancy. The authors interpret this as evidence that a student's major is largely recoverable from their early course selections.
The paper makes three main contributions: (1) formulating next-semester academic prediction as a joint course-set and grade prediction problem in a single model; (2) introducing a semester-level concurrency encoding that is permutation-invariant within a term while preserving temporal order across terms; and (3) proposing a set-based multi-task objective (KL for course-set prediction + MSE for grade regression) that avoids order artifacts induced by token-level cross entropy.
The authors note that the primary purpose of the model is grade prediction, and course prediction is utilized to encourage the model to learn meaningful course vector representations. This is analogous to how auxiliary tasks in multi-task learning improve representation quality for the primary task, similar to how BERT's secondary task of predicting the next sentence leads to better downstream prediction quality.
The paper discusses how this model could be incorporated into early detection systems at institutions of higher education. The model achieves a grade prediction MAE of 0.0392 on a [0, 1] GPA scale, which corresponds to 0.1568 on a typical [0, 4] GPA scale, meaning predictions are accurate, on average, to about half of the difference between a typical +/- grade (e.g. B vs B+).
The authors acknowledge several limitations, including that the model was trained and tested on data from a single, medium-sized private university, the features were limited to student major, course history, and grade history, and the model faces a cold start
problem for new students with no academic history. Future work includes validating the model on datasets from more diverse institutions, integrating richer feature sets such as Learning Management System data, and exploring Graph Transformer models.
Improvements for AI systems
Improvements to AI Systems:
-
Add a permutation-invariant temporal encoding layer to any sequence-based predictive model (e.g., RNNs, LSTMs, standard Transformers) that processes multi-item time steps. This layer assigns the same positional vector to all items within a single time step (e.g., courses in a semester), making the representation invariant to item order within that step while preserving order across steps. This prevents the model from learning spurious order dependencies among concurrent events.
-
Implement a multi-task objective with a set-based loss for auxiliary prediction (e.g., predicting the set of items a user will engage with next) alongside the primary regression or classification task. Use KL divergence for set prediction and MSE for the primary task, avoiding token-level cross-entropy that introduces order artifacts. This forces the model to learn richer, more generalizable item embeddings, improving primary-task accuracy even when the auxiliary task is not the end goal.
-
Add a
course-load concurrency
feature to grade prediction systems: encode the number and difficulty of concurrent items (e.g., courses taken in the same term) as an explicit input or via the positional encoding described above. This captures the effect of cognitive overload or resource competition, which flat-sequence models miss. -
Replace tree-based baselines (e.g., XGBoost) with attention-based architectures for sequential academic or behavioral prediction tasks, since tree models lack sequential memory and underperform on temporal dependencies.
-
Introduce a
major-agnostic
mode that drops categorical demographic features (like declared major) when they are highly correlated with early sequence content. The model can infer these from the sequence itself, reducing input redundancy and slightly improving performance.
What the Improved AI System Can Do:
-
Predict a student’s next-semester course set and grades with 46% lower mean absolute error than a grade-only Transformer, and 30% lower error than LSTM baselines.
-
Accurately forecast academic performance for students with heavy or concurrent course loads, where flat-sequence models fail.
-
Generate course embeddings that capture semantic relationships (e.g., prerequisite chains, difficulty clusters) without explicit feature engineering.
-
Operate without requiring student major as an input, making it more robust to missing demographic data.
-
Serve as an early-warning system in universities, flagging students at risk of low grades with an average error of 0.16 on a 4.0 GPA scale (half a letter grade).
-
Adapt to new students via cold-start handling by leveraging only their initial course selections, and can be extended to other domains (e.g., employee training, healthcare treatment plans) where concurrent item sets and outcome prediction are needed.
Abstract
Existing predictive models in learning analytics often treat student academic history as a simple sequence, overlooking the concurrent nature of courses taken within a semester. This simplification can lead to inaccurate performance predictions, particularly for students with heavy or challenging course loads. This paper introduces a TRansformer for Academic Course-grade Estimation (TRACE) that addresses this limitation by jointly predicting both the set of courses a student will take and their corresponding grades for an upcoming semester. Our approach encodes courses on a per-semester basis to capture the effects of course concurrency and utilizes a novel loss function combining course-set prediction with grade prediction. We demonstrate that predicting courses taken in addition to the grades in those courses leads to significant improvements in prediction quality. Trained on ten years of institutional data, our joint prediction model reduces mean absolute error by nearly 50% compared to an identical architecture that predicts grades alone. The model also outperforms traditional LSTM-based sequential models, as well as graph neural network-based approaches, and offers natural ways to incorporate student attribute data. This work demonstrates the utility of modern neural architectures for creating interpretable models that can be adapted to new institutions via retraining and recalibration, as well as the importance of key techniques, such as predicting courses taken during training. We discuss how this model could be incorporated into early detection systems at institutions of higher education.
Sources
- Context-aware Non-linear and Neural Attentive Knowledge-based Models for Grade Prediction
- Google's Neural Machine Translation System: Bridging the Gap between Human and Machine Translation
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Layer Normalization
- Decoupled Weight Decay Regularization
- SGDR: Stochastic Gradient Descent with Warm Restarts
- Density-aware Chamfer Distance as a Comprehensive Metric for Point Cloud Completion
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection