Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning".
Jane: The gist Contrastive learning can effectively mitigate bias in automated essay scoring systems while maintaining acceptable accuracy, suggesting that the fairness-accuracy trade-off may be less severe than previously assumed.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we're diving into this paper now, "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning." We've seen how these systems can accidentally penalize students just because they write in a second language, and this study looks at a way to fix that.
Jane: Exactly. The core idea here is using contrastive learning to adjust how the AI scores essays so it treats native speaker writing and ESL writing more equally, even when the essays are actually quite good.
Lu: It’s interesting because they start by showing that the baseline DeBERTa model has this issue where high-proficiency ESL students get scores ten point three percent lower than native speakers for essays of identical human quality <ref:2601.16724#pg1,scores 10.3% lower than native>.
Meng: So the problem isn't that the system is bad at grading low-quality writing, but it’s systematically capping the potential score for high-achieving ESL students regardless of their actual skill level.
Tom: Right, and to tackle that, they propose this contrastive learning approach using a triplet construction strategy to align those latent representations better.
Jane: They built a unified fairness benchmark by combining the ASAP two point zero dataset with the ELLIPSE corpus, which includes essays from English Language Learners, so they’re testing this against a wide range of writing situations <ref:2601.16724#pg1>.
Lu: The study constructed seventeen thousand one hundred sixty-one matched essay pairs and fine-tuned the model using Triplet Margin Loss to align those representations based on quality instead of just language origin <ref:2601.16724#pg1,17,161 matched essay pairs and fine-tuned the model using Triplet>.
Meng: So it's like teaching the AI that when two essays are high quality, they should be close together in its understanding, no matter if one is from a native speaker or an ESL writer.
Tom: And the results show a big improvement: the contrastive learning model reduced that high-proficiency scoring disparity by thirty-nine point nine percent, bringing it down to just a six point two percent gap while keeping the overall quality score, QWK, at zero point seven five six <ref:2601.16724#pg3,scoring disparity by 39.9%, bringing>.
Jane: That reduction is significant because it shows they managed to learn a fairer representation without collapsing the embedding space, which means the model stays reasonably accurate overall.
Lu: They also did an ablation study and found that setting the triplet margin to two point zero actually made things worse, dropping the QWK down to zero point seven one eight without any further reduction in bias, so they found that a margin of one point zero is optimal for this trade-off <ref:2601.16724#pg3>.
Meng: That’s a practical note for us; it tells us exactly where the sweet spot is when we try to balance fairness against the accuracy we need for real-world application.
Tom: So, what does this mean practically? It suggests that instead of just patching the surface issues, you can actively optimize the AI's internal logic to focus on essay quality itself rather than linguistic markers.
Jane: That means we’re moving toward systems that are designed to grade based on what a good essay is, not just what sounds like a native speaker wrote it.
Title and authors: Lu: The authors pointed out that they found sentence complexity was a key trigger for bias in the baseline model, showing a strong negative correlation between sentence length and predicted scores for ESL students.
Meng: That’s telling because it suggests the baseline model was misinterpreting complex second language clause structures as errors or poor reasoning, not just simple mistakes.
Tom: So the contrastive model successfully removed that link, learning that syntactic complexity is actually a stylistic feature in ESL writing rather than a sign of an error.
Jane: It shows that by optimizing the embedding space this way, we can remove demographic information from the decision boundary while still keeping the necessary semantic signals for grading.
Lu: They noted that although the model is fairer, it does show some systematic underrating for both groups, which suggests a more conservative scoring policy than before.
Meng: That's a fair point; it means even with this fix, the model might be slightly more cautious in its scoring decisions to ensure fairness.
Tom: Overall, the implication is that we can build these AI systems to be fairer and still maintain enough accuracy for deployment, provided we use methods like contrastive learning instead of just relying on surface-level features.
Jane: So the paper "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning" gives us a concrete method to address bias by aligning quality representations directly.
Lu: This work lays out how to use triplet construction strategy and margin tuning to achieve a thirty-nine point nine percent reduction in bias, moving the gap down from ten point three percent to six point two percent.
Meng: From an engineering standpoint, it’s valuable because it gives us a clear path on how much alignment we can push before we start hurting the actual predictive power of the model.
Tom: It’s a really helpful blueprint for anyone working on automated essay scoring to move beyond just using surface heuristics and start learning deeper quality signals.
Jane: So, this paper suggests that by optimizing the embedding space to cluster essays based on quality rather than linguistic origin, we can significantly reduce the gap between native and ESL student scores.
Lu: It’s a good demonstration of how domain-specific interventions within contrastive learning can effectively remove those spurious correlations between syntax and score for different groups.
Meng: This approach is grounded because it doesn't just rely on abstract fairness metrics; it’s actually modifying the training process to optimize for a specific, measurable outcome: fairness in essay scoring.
Tom: We’ve seen how this paper tackles the problem of high-proficiency ESL students being systematically penalized by showing that they receive scores ten point three percent lower than native speakers for essays of identical human quality <ref:2601.16724#pg1,receive scores 10.3% lower than native>.
Jane: The contrastive learning method shows that we can actually shrink that gap substantially, achieving a six point two percent gap while keeping the model’s overall agreement score quite high at zero point seven five six <ref:2601.16724#pg2>.
Title and authors: Lu: So, the paper demonstrates how to use contrastive learning with matched essay pairs to explicitly optimize the embedding space to cluster essays based on quality rather than linguistic origin.
Meng: It also highlights that setting the triplet margin at one point zero is the optimal point where we get a good fairness result without sacrificing too much of our predictive accuracy, given the ablation study results <ref:2601.16724#pg2>.
Tom: This whole discussion around "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning" suggests that we don't have to choose between making an AI fair and making it accurate at all.
Jane: It’s about finding a way to align the latent representations of essays so that the model understands what makes an essay good, regardless of which language it was written in.
Lu: This paper gives us a concrete method for using triplet construction strategy to reduce bias by explicitly optimizing the embedding space for quality alignment.
Meng: The practical impact here is showing how we can use contrastive learning to remove surface-level heuristics that cause unfair penalties based on L2 markers or sentence structure.
Tom: We’re going to wrap this up by summarizing how this work addresses the initial problem of constrained score scaling for high-proficiency ESL writing and what it means for future essay scoring AI.
Jane: So, to summarize, the paper "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning" uses contrastive learning with a triplet construction strategy to align essay representations based on quality instead of language origin.
Lu: This resulted in a thirty-nine point nine percent reduction in bias and brought the gap down from ten point three percent to just six point two percent, while maintaining a QWK of zero point seven five six against the aggressive margin of two point zero, which dropped to zero point seven one eight, showing that alpha equals one is optimal for this trade-off.
Meng: For us in the engineering world, this means we have a tested intervention that targets specific latent representations of quality and can be deployed with reasonable accuracy metrics at hand.
Tom: It’s a solid step forward in making these systems more equitable without sacrificing the necessary predictive power for grading essays across different language groups.
Jane: This paper gives us a clear path to address the issue where ESL students face penalties because their valid L2 linguistic markers are misinterpreted as evidence of poor reasoning or low quality.
Lu: It confirms that by optimizing the embedding space, we can successfully remove correlations between syntactic complexity and predicted scores for ESL students, treating it as a feature instead of a fault.
Meng: The limitation they flag is that while the method reduces bias significantly, it still exhibits a systematic underrating for both groups, which suggests even this fairer model might be slightly more conservative in its scoring policy.
Tom: So the final thought on "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning" is that we can build systems that are both fairer and sufficiently accurate for deployment by actively optimizing quality signals within the AI's training process.
The paper's summary: Tom: So we’re circling back to this paper on essay scoring bias for ESL learners. What they’re really saying is that instead of trying to fix every single rule in a massive manual, you can just adjust how the AI learns what makes a good essay, focusing on quality itself.
Jane: Exactly. They used contrastive learning to show that by training the AI to group essays based on their actual quality—not just whether they sound like a native speaker—you can significantly reduce that gap between high-achieving ESL students and native writers.
Lu: The main mechanism they used was creating these fairness triplets, where an anchor essay, like a native one, got paired with positive and negative examples from the ESL group based on their actual human scores. It’s forcing the model to learn a better way to understand what high quality looks like across both groups simultaneously.
Meng: From an engineering standpoint, that means we’re not just tweaking some surface-level filter; we’re changing how the AI builds its internal understanding of 'good writing' by training it on paired examples where language origin is deliberately ignored in favor of semantic quality.
Tom: And the results are pretty telling. They found they could cut that initial ten percent score difference down to about six percent without sacrificing much of the overall accuracy, which is a big win for deployment because we don't want to lose that predictive power just for fairness.
Jane: It’s about moving away from systems that penalize complex L2 structures as errors, and instead learning that those structures are actually stylistic features of advanced writing. That’s a huge shift in how we think about what makes an essay strong.
Lu: They even found that sentence complexity, which they thought was a trigger for bias, is actually treated correctly by the contrastive model because it understands the context better now. It stops misinterpreting long sentences as mistakes and starts seeing them as sophisticated structure.
Meng: So what this means practically is that if we build scoring tools, we can start trusting them more because we’ve shown a method to actively remove those unfair shortcuts based on language markers. We’re targeting the latent representations of quality directly.
Tom: It shows that the fairness-accuracy trade-off isn't as bad as some people thought; you can get thirty-nine point nine percent less bias and still keep a solid score metric at zero point seven five six, which is really encouraging for real-world application.
Jane: And though they did find it’s still a bit more conservative for both groups, meaning the model scores slightly lower overall than before, the trade-off seems manageable when you look at the reduction in that massive proficiency disparity.
Lu: The authors pointed out that while this fixes one huge bias issue, there are still some limitations—they said it doesn't fix every single nuance of scoring perfectly, which is expected for any system.
Meng: So the future work they suggest is to look at score recalibration techniques to keep the fair embedding space clean while trying to restore that absolute top-tier accuracy back up. It’s about fine-tuning the final step.
Tom: That’s a really concrete roadmap—fix the representation first, then tune the output layer later—which gives us a clear path forward for developing more equitable essay scoring AI.
The paper's improvements: Tom: So we’re talking about how this approach actually changes things for essay scoring, moving past just identifying where the bias is happening to actually fixing it systematically. What they suggest is using these alignment tools to create a more balanced internal map of what makes an essay high quality across different language backgrounds.
Jane: Exactly. It’s not just about patching the score; it’s about retraining the AI so its definition of 'high quality' isn't secretly biased toward one linguistic style over another. It forces the model to see that a complex sentence structure from an ESL student is still high quality if it meets the same semantic goals as a native writer.
Lu: The real improvement they highlight is disentangling that syntactic complexity issue we talked about earlier. They showed that by using this contrastive training, the AI learns to treat sentence length as just another feature, not a signal for error when writing in English. That’s a huge unlock for ESL students who write with more complex structures.
Meng: From an engineering angle, that means we can design systems where the model ignores those surface-level linguistic markers that were causing the initial penalties and focuses purely on the underlying reasoning or content quality. It makes the scoring process much more robust against those common heuristics.
Tom: And they found this whole method works best when you tune a specific setting, which is interesting because it shows there’s an optimal balance to hit for fairness without losing too much predictive power. They pinpointed that margin one point zero as the sweet spot for that trade-off.
Jane: So, what this means for someone listening is that we can expect tools to become less biased toward specific writing styles, meaning a student's ability to express complex ideas isn't unfairly penalized just because they aren't native speakers.
Lu: The potential here is huge; imagine an AI tutor that scores essays based on genuine thought process rather than just grammar rules or sentence length. It opens up entirely new ways to assess learning in second languages, which is where my head goes with this research.
Meng: For us building these platforms, it means we can deploy systems that are genuinely fairer while still maintaining the accuracy needed for serious educational use cases. It’s a practical step toward making AI tools usable by everyone without putting high-achieving ESL writers at a disadvantage.
Tom: It really shows that this isn't just theoretical math; it’s a tangible process of refining the model to be more equitable in its judgment, and that tuning that margin is the lever we have to pull for practical deployment.
Conclusion: Tom: So we’ve covered how contrastive learning helps reduce that massive score gap for high-proficiency ESL students in this paper, "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning." It boils down to using careful alignment during training to make the AI learn what good writing is, regardless of who wrote it.
Jane: Right. We’ve seen how they used those triplet construction ideas and found that tuning the margin at one point zero really helped them get that fairness boost without hurting overall accuracy too much. It's a nuanced balancing act.
Lu: The big picture here is that we can start building AI scoring systems where the definition of quality is decoupled from language origin, which could fundamentally change how we assess writing across different cultures and educational settings. It opens up new possibilities for equitable learning tools.
Meng: I see the practical implication as moving away from simply trying to patch surface-level errors in the model and instead training it on quality signals that are inherently more stable, making the tool much more reliable for real deployment. It’s about building systems that can handle real-world variability without getting tripped up by language features.
Lalam: If we look at this from a cultural standpoint, this work means we can build tools that don't unintentionally reinforce biases against certain communication styles; it helps create a more inclusive digital culture where writing skills are valued for what they actually communicate.
Tom: It’s a solid blueprint for how to tackle these kinds of subtle biases in large language models applied to education, and I think we’re going to see more of this type of targeted intervention in the coming months.
Jane: Exactly. The "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning" paper shows that when you focus on aligning the representation space based on true quality, you can achieve a much fairer and still highly accurate system.
Lu: It’s exciting because it proves that domain-specific techniques like contrastive learning are really effective at tackling deep, latent biases in language models without needing a complete overhaul of the entire model architecture.
Meng: I think the next step for us is figuring out how to integrate these alignment strategies into our existing agent planning frameworks to see if we can apply this quality-based alignment across other complex tasks as well.
Lalam: This paper gives us a clear direction for making AI systems that support diverse communication styles, and that’s something I think will really help shape the future of how we interact with intelligent tools.
Tom: Well, that wraps up our deep dive into this research on essay scoring bias, but keep an eye out because we’ve got another interesting paper coming up next about how digital personas handle survey findings.
Kevin Fan, Eric Yun
Georgia Institute of Technology · Georgia State University
cs.CL
Submitted: 2026-01-23
Updated: 2026-10-05
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 87/100
The gist: The gist Contrastive learning can effectively mitigate bias in automated essay scoring systems while maintaining acceptable accuracy, suggesting that the fairness-accuracy trade-off may be less
Key concepts
- Shortcut Learning
- This occurs when neural networks rely on simple, easy-to-learn surface patterns instead of deep, complex semantic understanding. In AES, this means the system might incorrectly focus on superficial markers like specific L2 grammar errors rather than assessing the actual quality of student reasoning.
- Contrastive Learning
- A training technique where the model learns to distinguish between different types of data by comparing them. Here, it pairs an anchor essay (native) with a positive example (a similar ESL essay with a good score) and a negative example (an essay with a low score). This forces the model to learn more meaningful features.
- Fairness Triplets
- A custom dataset created for training. It consists of three parts: an Anchor A (a native-level essay), a Positive P (an ESL essay that should receive a high score), and a Negative N (an essay with a low score). This structure guides the model to learn what constitutes fair scoring across different proficiency groups.
Terminology
Summary
The gist
Contrastive learning can effectively mitigate bias in automated essay scoring systems while maintaining acceptable accuracy, suggesting that the fairness-accuracy trade-off may be less severe than previously assumed.
Introduction and Problem
Automated Essay Scoring (AES) systems are prone to “shortcut learning” where they rely on easy-to-learn surface heuristics instead of complex semantic reasoning. In the context of ESL writing, transformer attention heads often disproportionately attend to distinct L2 markers such as prepositional misuse or specific sentence structures as proxies for predicting lower scores <ref:2601.16724#pg4>. This results in a large penalty for ESL students who face the cognitive load of writing in a second language, and AES systems interpret valid L2 linguistic markers as evidence of poor reasoning <ref:2601.16724#pg4>.
Methodology
The study constructed a unified fairness benchmark by merging two major datasets: 1) ASAP 2.0, which contains Native speakers (N ≈ 17,000) and ESL speakers (N ≈ 2,600), and 2) ELLIPSE, a corpus of essays specifically from English Language Learners. The baseline model used was a fine-tuned microsoft/deberta-v3-base with LoRA rank r = 16, trained using Mean Squared Error (MSE) loss. To implement contrastive learning, the researchers algorithmically generated a dataset of “Fairness Triplets” (A, P, N) by selecting an Anchor A (Native essay) and pairing it with a Positive P (an ESL essay with a human score in the range [S − 0.02, S + 0.02]) and a Negative N (an essay with a score in the range [S ± 0.20]). The contrastive training utilized Triplet Margin Loss: L = max(0, d(A, P) − d(A, N) + α) (1), with a margin of α = 1.0.
Results and Performance
The baseline DeBERTa model achieved a Quadratic Weighted Kappa (QWK) of 0.79, but stratified analysis revealed a massive disparity where the ESL residual was-0.11 compared to-0.01 for Native Residual <ref:2601.16724#pg4>. The Contrastive Learning model demonstrated a significant reduction in bias, reducing the high-proficiency scoring disparity by 39.9% (to a 6.2% gap) while maintaining a QWK of 0.756. An ablation study showed that increasing the margin to α = 2.0 resulted in a QWK drop to 0.718 without further bias reduction, suggesting that α = 1.0 represents the optimal Pareto frontier.
Linguistic Triggers and Conclusion
Post-hoc linguistic analysis identified that sentence complexity was a key trigger for bias, as the baseline model showed a strong negative correlation between sentence length and predicted score for ESL students, likely misinterpreting complex L2 clause structures as syntactic errors. The contrastive model successfully removed this correlation, effectively learning that syntactic complexity in ESL writing is a stylistic feature rather than an error. Although the model exhibits systematic underscore for both groups, suggesting a more conservative scoring policy, the approach targets specific latent representations of quality, resulting in a model that is both fairer and sufficiently accurate for deployment. Future work should explore score recalibration techniques to maintain the fair embedding space while restoring absolute accuracy.
References
[1] Evelin Amorim, Marcia Can ´ c¸ado, and Adriano Veloso. Automated essay scoring in the presence of biased ratings. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 229–237, 2018
[5] Robert Geirhos, Jorn-Henrik Jacobsen, Claudio ¨ Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2(11):665–673, 2020
[8] Elijah Mayfield, Michael Madaio, Shrimai Prabhumoye, David Gerritsen, Brittany McLaughlin, Ezekiel DixonRoman, and Alan W Black. Equity beyond bias in ´ automated essay scoring. In Proceedings of the Fourteenth Workshop on Innovative Use of NLP for Building Educational Applications, pages 444–460, 2019
[9] Hansen Susanto, Alexander AS Gunawan, and M Fikri Hasani. Development of automated essay scoring system using DeBERTa as a transformer-based language model. In Computational Methods in Systems and Software, pages 202–215. Springer, 2024
[10] Kai Yang, Xinyu Liu, Qun Liu, C Guan, and L Wei. Unveiling the tapestry of automated essay scoring: A comprehensive investigation of accuracy, fairness, and generalizability. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pages 22466–22474, 2024
[11] Brian Hu Zhang, Blake Lemoine, and Margaret Mitchell. Mitigating unwanted biases with adversarial learning. In Proceedings of the 2018 AAAI/ACM Conference on AI, Ethics, and Society, pages 335–340, 2018
[4] Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, pages 4171–4186, 2019
[6] Pengcheng He, Xiaodong Liu, Jianfeng Gao, and Weizhu Chen. DeBERTa: Decoding-enhanced BERT with disentangled attention. In International Conference on Learning Representations, 2021
[7] Edward J Hu, Yelong Shen, Phil Wallis, Zeyuan AllenZhu, Yuanzhi Li, Shean Wang, Lu Wang, and Weizhu Chen. LoRA: Low-rank adaptation of large language models. In International Conference on Learning Representations, 2022
[3] Scott Crossley, Yingjian Tian, Pae Baffour, A Franklin, Y Kim, W Morris, B Benner, A Picou, and U Boser. The English Language Learner Insight, Proficiency and Skills Evaluation (ELLIPSE) corpus. International Journal of Learner Corpus Research, 9(2):248–269, 2024
[1] Evelin Amorim, Marcia Can ´ c¸ado, and Adriano Veloso. Automated essay scoring in the presence of biased ratings. In Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), pages 229–237, 2018
[5] Robert Geirhos, Jorn-Henrik Jacobsen, Claudio ¨ Michaelis, Richard Zemel, Wieland Brendel, Matthias Bethge, and Felix A Wichmann. Shortcut learning in deep neural networks. Nature Machine Intelligence, 2(11):665–673, 2020
[9] Hansen Susanto, Alexander AS Gunawan, and M Fikri Hasani. Development of automated essay scoring system using DeBERTa as a transformer-based language model. In Computational Methods in Systems and Software, pages 202–215.
Improvements for AI systems
-
Robustness against high-proficiency ESL bias through contrastive learning: The improved system can achieve a
constrained score scaling for high-proficiency ESL writing where high-proficiency ESL essays receive scores 10.3% lower than Native speaker essays of identical human-rated quality
by aligning the latent representations of ESL and Native writing, thereby reducing the disparity by39.9% (to a 6.2% gap).
-
Disentanglement of syntactic complexity from grammatical error: The system can correct for surface-level heuristics by removing the correlation between sentence structure and score, as demonstrated by the finding that
The baseline model showed a strong negative correlation between sentence length and predicted score for ESL students.
-
Optimal margin selection for fairness-accuracy trade-off: The system can maintain high accuracy while achieving fairer results by employing an optimal alignment strategy, as indicated by the ablation study showing that
α = 1.0 represents the optimal Pareto frontier
yielding a QWK of 0.756 against an aggressive margin of α = 2.0 which resulted in a QWK drop to 0.718 without further bias reduction.
Abstract
Automated Essay Scoring systems disproportionately penalize high-proficiency English as a Second Language (ESL) learners. We propose Contrastive Learning with Matched Essay Pairs (CL-MEP), a bi-directional alignment strategy. CL-MEP reduces this scoring bias by 39.9% while improving overall accuracy, successfully disentangling valid syntactic complexity from surface-level grammatical errors.
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering