Not Every Subject Should Stay: Machine Unlearning for Noisy Engagement Recognition
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Not Every Subject Should Stay".
Jane: Engagement recognition datasets are typically subject-indexed and often contain noisy, subjective supervision, making post-hoc dataset revision a practical problem.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: Well, essentially, the paper shows that you can use a model-dependent score to flag subjects that seem harmful based on how much training loss they contribute to the baseline model. Then, by applying a lightweight update process only to those flagged subjects, they managed to recover between eighty-nine point three percent and ninety-two point five percent of the performance boost we see when you retrain from scratch on just the good data <ref:2605.04713#pg0>.
Jane: That recovery rate is quite impressive when you consider that this happens at roughly one quarter of the cost of a full retraining, which makes it a very practical correction mechanism for existing models. It shows that post-hoc subject removal isn't just random noise; it has direction.
Lu: They found something interesting regarding the size of the group they remove. Their sensitivity analysis showed that effectiveness is strongest when you target an intermediate number of subjects, specifically when you select a forget-set size of K=three whereas going up to K=five actually made the gains less favorable <ref:2605.04713#pg2>.
Meng: That finding about the regime dependence is important for us. It tells us that simply removing more data might not help if you're over-correcting; there’s a sweet spot in how much you remove before the benefits start to drop off, which helps me plan our deployment strategy.
Lalam: It’s encouraging because it suggests we don't need a massive cleanup effort every time. We can target the core issues efficiently, which streamlines our model maintenance pipeline and keeps things running smoothly two <ref:2605.04713#pg0>.
Tom: So, they aren't just saying you can unlearn anything; they are providing a specific protocol: use model-dependent scores to pick subjects, apply an approximate update, and evaluate against a fully retrained oracle model. That’s a very concrete roadmap for anyone working in this area.
Jane: Precisely; the paper establishes that using an oracle model—one trained only on the retained subjects—is the best way to measure if our unlearning actually achieves what we want, instead of just seeing random fluctuations in performance <ref:2605.04713#pg2>.
The paper's summary: Tom: They’ve essentially proposed a whole pipeline, starting from training a baseline model on everything, moving to identifying harmful subjects via that mean per-clip training loss score, and then applying that lightweight update specifically to those identified subjects <ref:2605.04713#pg0>.
Lu: The main improvement they introduce is framing subject-level post-hoc sanitization as a distinct problem for engagement recognition and positioning machine unlearning as the right tool for revising trained models after identifying these specific problematic units <ref:2605.04713#pg2>.
Meng: From an engineering standpoint, their use of TCCT-Net as a fixed platform while only updating the final fusion and classification layers is a clever way to keep the heavy feature extractor stable while only modifying what needs to change for the unlearning process <ref:2605.04713#pg0>.
Lalam: This focus on updating just the top layers is very smart because it keeps our massive foundational knowledge intact while we selectively fine-tune away specific undesirable behaviors or data associations two <ref:2605.04713#pg0>.
Jane: Another key improvement they highlight is their oracle-centered evaluation protocol, which compares the unlearned model against a model retrained from scratch on only the retained subjects, making the measurement much more meaningful than just looking at performance loss alone <ref:2605.04713#pg2>.
Tom: And that connects back to their finding about regime dependence; they show that optimizing this process involves selecting an intermediate forget-set size, like K=three because expanding the set beyond that shows a negative marginal utility <ref:2605.04713#pg2>.
The paper's improvements: Jane: I think the main conclusion is that subject-level post-hoc sanitization is a plausible framework for revising trained engagement models, provided you are careful about how you select which subjects to remove and what your criteria are for that selection <ref:2605.04713#pg0>.
Lu: From a theoretical view, the work supports the idea that lightweight approximate unlearning can serve as a useful correction mechanism when the selected forget-set is genuinely beneficial, rather than proving some universal deletion behavior across all datasets or backbones <ref:2605.04713#pg1>.
Meng: For practical implementation, it’s clear that we have a viable low-cost method for sanitizing pre-trained engagement recognition models when we can identify specific problematic subjects using a model-dependent proxy, which is a huge step toward making dataset revision feasible <ref:2605.04713#pg0>.
Lalam: I feel this has big implications for how we deploy AI in real environments; it means our systems can be actively maintained and corrected against specific data issues without needing to halt the entire training cycle for every minor problem two <ref:2605.04713#pg0>.
Conclusion: Tom: So, to wrap up on "Not Every Subject Should Stay: Machine Unlearning for Noisy Engagement Recognition," they showed us how we can surgically remove problematic subjects from trained engagement models without doing a full retraining cycle, and that's a really neat trick.
Jane: It’s certainly neat because it takes something conceptually messy—noisy supervision and subject-specific data—and gives us a practical way to clean up the model afterward, which is exactly what we need for high-stakes recognition systems.
Lu: I think the real excitement lies in that method's ability to recover nearly ninety percent of the oracle gain at only a fraction of the retraining cost; it suggests that targeted unlearning can be much more efficient than brute force data replacement.
Meng: From an engineering standpoint, if we can implement this easily, it really changes how fast we can iterate on models when new data quality issues pop up; you don't have to wait for a full pipeline rebuild.
Lalam: For me, the cultural impact of this is huge because it means that as AI systems become more complex and trained on vast amounts of varied human interaction data, we gain a tool that allows us to actively prune undesirable biases or problematic patterns from the model's learned behavior.
Tom: That’s a big picture thought, Lalam; essentially, we’re talking about making our models more agile and responsive to targeted quality control needs.
Jane: Exactly, Tom; it gives us a way to maintain model integrity while still being able to refine the system based on real-world feedback without drowning in computational overhead.
Lu: And that regime dependence they found with the K=three forget-set size really shows that the effectiveness isn't just about removing more data blindly, but about understanding where the core of the issue lies in those subject removals.
Meng: That means we can actually optimize our cleanup process; we don't just throw everything away at once, which makes sense for a real-world deployment scenario.
Lalam: It’s encouraging because it moves us closer to building AI that isn't just big and complex, but also smart enough to know when it needs to prune itself for better performance and reliability.
Tom: So we’ve seen how they used mean training loss as a heuristic, applied a lightweight update on top of that, and proved it works in practice for improving engagement recognition models.
Jane: That’s the core message; it’s not about perfect deletion, but about achieving significant performance recovery through smart post-hoc correction when we have the right criteria for what constitutes a "harmful" subject.
Lu: I think the future work they mentioned, exploring certified deletion or universal behavior across different datasets, is where the deep theoretical exploration goes next.
Meng: That’s where my practical concerns come in; until we get certified results that work everywhere, we still have to rely on their heuristic-based scoring for now.
Lalam: But even with those limitations acknowledged, the framework itself provides a solid path forward for building more responsible and maintainable AI systems.
Tom: Absolutely; it’s a powerful tool for anyone looking to keep their engagement recognition models sharp and relevant in an evolving data landscape.
cs.CV
Submitted: 2026-05-06
Updated: 2026-10-06
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 79/100
The gist: Engagement recognition datasets are typically subject-indexed and often contain noisy, subjective supervision, making post-hoc dataset revision a practical problem.
Key concepts
- Oracle Model
- This is the ideal reference model created by retraining a model from scratch using only the subjects that were intentionally kept. It serves as the benchmark to measure how much performance can be recovered through subject removal.
- Model-Dependent Subject Scores
- A heuristic used to rank subjects for removal. It scores each subject based on its mean training loss under the already trained baseline model, suggesting that subjects with higher loss are more likely to be problematic or atypical.
- Approximate Unlearning Update
- A lightweight method applied to the trained model specifically targeting a fixed set of identified 'harmful' subjects. This is a practical, post-hoc correction technique rather than a complete retraining process.
- Removal Regime
- Refers to the size of the subject set being removed (the forget-set K). The study found that effectiveness depends on this regime, suggesting that removing only a small core group yields the best results.
Terminology
Summary
Engagement recognition datasets are typically subject-indexed and often contain noisy, subjective supervision, making post-hoc dataset revision a practical problem. The gist: approximate subject-level unlearning can recover 89.3% to 92.5% of the oracle gain on EngageNet and DAiSEE at roughly one quarter of retraining cost when removing problematic subjects after baseline training.
Problem Formulation
The paper addresses the setting where a model has already been trained, and the conceptual clean correction is to remove a problematic subject and retrain on the retained data only, which is often impractical post-hoc. The central research question is: can post-hoc subject-level forgetting recover a meaningful portion of the benefit of oracle data sanitization without paying the full cost of retraining?
This framework frames engagement recognition as a controlled testbed to determine if post-hoc subject removal is a meaningful correction mechanism when supervision is noisy, subjective, and organized by subjects.
Methodology and Protocol
The study instantiates a practical pipeline: (i) train a baseline model on all training subjects; (ii) use model-dependent subject scores
to identify a candidate forget-set of potentially harmful subjects; (iii) apply a lightweight approximate unlearning update
to the trained baseline for that fixed forget-set; and (iv) compare the resulting model against an oracle model retrained from scratch on the retained subjects only.
The protocol is applied on TCCT-Net, where only the final fusion and classification layers
are updated, freezing the lower-level feature extractor.
Subject Harmfulness Scoring
Subjects are ranked using a heuristic based on their training loss under the trained baseline. The main scoring rule assigns each training subject a score equal to the mean per-clip training loss under the trained baseline,
calculated as:
r s = 1/D s Σ(x,y)∈D s lce(fθ0(x), y).
Subjects are then ranked by this score, and the top-K ranked subjects define the forget-set F K.
This score is used as a simple heuristic rather than a calibrated estimate of subject harmfulness,
acknowledging that high loss may reflect ambiguity or atypicality rather than absolute error.
Oracle-Centered Evaluation
The evaluation protocol uses an oracle model, defined as the model retrained from scratch on the retained subjects only.
This reference is considered more informative than measuring forgetting only on the forgotten subset because it separates two issues: whether selected subjects were actually harmful and whether unlearning moves toward the desired retained-data solution. The primary comparison is between baseline on all subjects → unlearned model ↔ oracle retraining without F.
Results and Regime Dependence
The results show that for both DAiSEE and EngageNet, the proposed update moves in the same direction as oracle retraining, recovering 89.3% to 92.5% of the baseline-to-oracle gain under Accuracy. Sensitivity analysis across forget-set sizes (K=1, 3, 5) reveals a small audit regime
where effectiveness is strongest at an intermediate forget-set size (K=3). Expanding the set beyond K=3 to K=5 shows that the marginal utility becomes negative,
suggesting that stronger forgetting does not automatically yield more retained-task benefit once removal extends beyond the small harmful core. The unlearning update is most convincing when its behavior tracks the same qualitative trend as oracle retraining.
Limitations and Conclusion
The study's limitations include using a proxy for harmfulness (mean loss) rather than ground truth, and acknowledging that the proposed update is explicitly approximate
and restricted to a small trainable subset, serving as a practical post-hoc correction mechanism
rather than proof of exact forgetting. The conclusion supports the finding that subject-level post-hoc sanitization is a plausible framework for revising trained engagement models, with effectiveness depending on the removal regime.
The paper does not establish certified deletion or universal behavior across datasets, backbones, and subject-selection rules. Instead, it demonstrates that lightweight approximate unlearning can serve as a useful correction mechanism when the selected forget-set is genuinely beneficial.
The strongest results occur at an intermediate forget-set size, indicating that effectiveness depends on the removal scenario.
The paper does not establish certified deletion or universal behavior across datasets, backbones, and subject-selection rules. Instead, it supports a narrower conclusion: subject-level post-hoc sanitization is a plausible framework for revising trained engagement models, with effectiveness depending on subject selection quality and the removal regime. The claim remains deliberately bounded.
The paper does not establish certified deletion or universal behavior across datasets, backbones, and subject-selection rules. Instead, it supports a narrower conclusion: subject-level post-hoc sanitization is a plausible framework for revising trained engagement models, with effectiveness depending on subject selection quality and the removal regime.
Improvements for AI systems
Here are specific improvements to AI systems based on the findings of this paper:
-
Improve engagement recognition models by implementing a post-hoc, subject-level unlearning mechanism that approximates full retraining. This allows for targeted removal of problematic subjects without incurring the full computational cost of retraining from scratch.
-
The improved system can be used to sanitize pre-trained engagement recognition models (e.g., TCCT-Net) by identifying and removing specific subjects deemed harmful based on a model-dependent proxy (mean per-clip training loss).
-
The improved system enables the creation of an
Oracle Reference
model—a gold standard trained only on the retained, non-problematic subjects—which serves as a precise target for unlearning updates, ensuring the correction moves toward retaining useful knowledge rather than just randomly perturbing it. -
The resulting AI system will be more computationally efficient for dataset revision tasks, achieving performance gains (e.g., 89% to 92% of oracle gain) at roughly one-quarter of the retraining cost, making post-hoc dataset sanitization a viable low-cost correction mechanism for high-stakes engagement recognition systems.
-
The system will exhibit regime dependence awareness: its effectiveness can be optimized by selecting an intermediate forget-set size (K=3), suggesting that harmfulness is often distributed across a small core rather than concentrated in one outlier, and it will avoid discarding useful variability when expanding the removal budget beyond this core.
Abstract
Engagement recognition datasets are typically subject-indexed and often contain noisy, subjective supervision, making post-hoc dataset revision a practical problem. Existing noisy-label and data-cleaning methods largely operate at the sample level before or during training, but do not directly address a different question: once a model has already been trained, can the influence of an entire problematic subject be removed without full retraining? We study this setting through subject-level machine unlearning as a post-hoc sanitization mechanism for engagement recognition. Starting from a baseline trained on all subjects, we rank candidate harmful subjects using a model-dependent proxy, apply a lightweight approximate unlearning update, and compare the result against an oracle model retrained from scratch on the retained subjects only. We instantiate this protocol on DAiSEE and EngageNet using Tensor-Convolution and Convolution-Transformer Network (TCCT-Net) as a fixed platform and evaluate three matched model states under the same removal scenario: baseline, unlearned, and oracle. In representative K=3 forget-set settings, the unlearned model recovers 89.3% and 92.5% of the oracle gain on EngageNet and DAiSEE, respectively, at roughly one quarter of retraining cost. Across the tested small-audit regimes, effectiveness is strongest at an intermediate forget-set size, indicating that approximate subject-level unlearning is a useful low-cost correction mechanism, but one whose benefit depends on subject selection quality and removal regime.
Sources
- DAiSEE: Towards User Engagement Recognition in the Wild
- Computational Analysis of Stress, Depression and Engagement in Mental Health: A Survey
- PriorNet: Prior-Guided Engagement Estimation from Face Video
- Machine Unlearning: A Comprehensive Survey
- Learn to Forget: Machine Unlearning via Neuron Masking
- Federated Unlearning: How to Efficiently Erase a Client in FL?
- Distilling the Knowledge in a Neural Network
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models