Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning
summary
The gist
The gist Contrastive learning can effectively mitigate bias in automated essay scoring systems while maintaining acceptable accuracy, suggesting that the fairness-accuracy trade-off may be less
In short
The study used contrastive learning to reduce bias in automated essay scoring (AES) systems for English as a Second Language (ESL) learners. By training a model with 'Fairness Triplets,' the researchers significantly decreased the score disparity between native and ESL students by nearly 40%, showing that contrastive learning effectively mitigates unfair linguistic penalties.
Key concepts
- Shortcut Learning
- This occurs when neural networks rely on simple, easy-to-learn surface patterns instead of deep, complex semantic understanding. In AES, this means the system might incorrectly focus on superficial markers like specific L2 grammar errors rather than assessing the actual quality of student reasoning.
- Contrastive Learning
- A training technique where the model learns to distinguish between different types of data by comparing them. Here, it pairs an anchor essay (native) with a positive example (a similar ESL essay with a good score) and a negative example (an essay with a low score). This forces the model to learn more meaningful features.
- Fairness Triplets
- A custom dataset created for training. It consists of three parts: an Anchor A (a native-level essay), a Positive P (an ESL essay that should receive a high score), and a Negative N (an essay with a low score). This structure guides the model to learn what constitutes fair scoring across different proficiency groups.
Terminology used across episodes
This episode discusses
The paper
Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning · Read on arXiv
Kevin Fan, Eric Yun
Georgia Institute of Technology · Georgia State University
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning".
Jane: The gist Contrastive learning can effectively mitigate bias in automated essay scoring systems while maintaining acceptable accuracy, suggesting that the fairness-accuracy trade-off may be less severe than previously assumed.
Tom: First, who's behind it and why it matters.
Title and authors: Tom: So we're diving into this paper now, "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning." We've seen how these systems can accidentally penalize students just because they write in a second language, and this study looks at a way to fix that.
Jane: Exactly. The core idea here is using contrastive learning to adjust how the AI scores essays so it treats native speaker writing and ESL writing more equally, even when the essays are actually quite good.
Lu: It’s interesting because they start by showing that the baseline DeBERTa model has this issue where high-proficiency ESL students get scores ten point three percent lower than native speakers for essays of identical human quality <ref:2601.16724#pg1,scores 10.3% lower than native>.
Meng: So the problem isn't that the system is bad at grading low-quality writing, but it’s systematically capping the potential score for high-achieving ESL students regardless of their actual skill level.
Tom: Right, and to tackle that, they propose this contrastive learning approach using a triplet construction strategy to align those latent representations better.
Jane: They built a unified fairness benchmark by combining the ASAP two point zero dataset with the ELLIPSE corpus, which includes essays from English Language Learners, so they’re testing this against a wide range of writing situations <ref:2601.16724#pg1>.
Lu: The study constructed seventeen thousand one hundred sixty-one matched essay pairs and fine-tuned the model using Triplet Margin Loss to align those representations based on quality instead of just language origin <ref:2601.16724#pg1,17,161 matched essay pairs and fine-tuned the model using Triplet>.
Meng: So it's like teaching the AI that when two essays are high quality, they should be close together in its understanding, no matter if one is from a native speaker or an ESL writer.
Tom: And the results show a big improvement: the contrastive learning model reduced that high-proficiency scoring disparity by thirty-nine point nine percent, bringing it down to just a six point two percent gap while keeping the overall quality score, QWK, at zero point seven five six <ref:2601.16724#pg3,scoring disparity by 39.9%, bringing>.
Jane: That reduction is significant because it shows they managed to learn a fairer representation without collapsing the embedding space, which means the model stays reasonably accurate overall.
Lu: They also did an ablation study and found that setting the triplet margin to two point zero actually made things worse, dropping the QWK down to zero point seven one eight without any further reduction in bias, so they found that a margin of one point zero is optimal for this trade-off <ref:2601.16724#pg3>.
Meng: That’s a practical note for us; it tells us exactly where the sweet spot is when we try to balance fairness against the accuracy we need for real-world application.
Tom: So, what does this mean practically? It suggests that instead of just patching the surface issues, you can actively optimize the AI's internal logic to focus on essay quality itself rather than linguistic markers.
Jane: That means we’re moving toward systems that are designed to grade based on what a good essay is, not just what sounds like a native speaker wrote it.
Title and authors: Lu: The authors pointed out that they found sentence complexity was a key trigger for bias in the baseline model, showing a strong negative correlation between sentence length and predicted scores for ESL students.
Meng: That’s telling because it suggests the baseline model was misinterpreting complex second language clause structures as errors or poor reasoning, not just simple mistakes.
Tom: So the contrastive model successfully removed that link, learning that syntactic complexity is actually a stylistic feature in ESL writing rather than a sign of an error.
Jane: It shows that by optimizing the embedding space this way, we can remove demographic information from the decision boundary while still keeping the necessary semantic signals for grading.
Lu: They noted that although the model is fairer, it does show some systematic underrating for both groups, which suggests a more conservative scoring policy than before.
Meng: That's a fair point; it means even with this fix, the model might be slightly more cautious in its scoring decisions to ensure fairness.
Tom: Overall, the implication is that we can build these AI systems to be fairer and still maintain enough accuracy for deployment, provided we use methods like contrastive learning instead of just relying on surface-level features.
Jane: So the paper "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning" gives us a concrete method to address bias by aligning quality representations directly.
Lu: This work lays out how to use triplet construction strategy and margin tuning to achieve a thirty-nine point nine percent reduction in bias, moving the gap down from ten point three percent to six point two percent.
Meng: From an engineering standpoint, it’s valuable because it gives us a clear path on how much alignment we can push before we start hurting the actual predictive power of the model.
Tom: It’s a really helpful blueprint for anyone working on automated essay scoring to move beyond just using surface heuristics and start learning deeper quality signals.
Jane: So, this paper suggests that by optimizing the embedding space to cluster essays based on quality rather than linguistic origin, we can significantly reduce the gap between native and ESL student scores.
Lu: It’s a good demonstration of how domain-specific interventions within contrastive learning can effectively remove those spurious correlations between syntax and score for different groups.
Meng: This approach is grounded because it doesn't just rely on abstract fairness metrics; it’s actually modifying the training process to optimize for a specific, measurable outcome: fairness in essay scoring.
Tom: We’ve seen how this paper tackles the problem of high-proficiency ESL students being systematically penalized by showing that they receive scores ten point three percent lower than native speakers for essays of identical human quality <ref:2601.16724#pg1,receive scores 10.3% lower than native>.
Jane: The contrastive learning method shows that we can actually shrink that gap substantially, achieving a six point two percent gap while keeping the model’s overall agreement score quite high at zero point seven five six <ref:2601.16724#pg2>.
Title and authors: Lu: So, the paper demonstrates how to use contrastive learning with matched essay pairs to explicitly optimize the embedding space to cluster essays based on quality rather than linguistic origin.
Meng: It also highlights that setting the triplet margin at one point zero is the optimal point where we get a good fairness result without sacrificing too much of our predictive accuracy, given the ablation study results <ref:2601.16724#pg2>.
Tom: This whole discussion around "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning" suggests that we don't have to choose between making an AI fair and making it accurate at all.
Jane: It’s about finding a way to align the latent representations of essays so that the model understands what makes an essay good, regardless of which language it was written in.
Lu: This paper gives us a concrete method for using triplet construction strategy to reduce bias by explicitly optimizing the embedding space for quality alignment.
Meng: The practical impact here is showing how we can use contrastive learning to remove surface-level heuristics that cause unfair penalties based on L2 markers or sentence structure.
Tom: We’re going to wrap this up by summarizing how this work addresses the initial problem of constrained score scaling for high-proficiency ESL writing and what it means for future essay scoring AI.
Jane: So, to summarize, the paper "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning" uses contrastive learning with a triplet construction strategy to align essay representations based on quality instead of language origin.
Lu: This resulted in a thirty-nine point nine percent reduction in bias and brought the gap down from ten point three percent to just six point two percent, while maintaining a QWK of zero point seven five six against the aggressive margin of two point zero, which dropped to zero point seven one eight, showing that alpha equals one is optimal for this trade-off.
Meng: For us in the engineering world, this means we have a tested intervention that targets specific latent representations of quality and can be deployed with reasonable accuracy metrics at hand.
Tom: It’s a solid step forward in making these systems more equitable without sacrificing the necessary predictive power for grading essays across different language groups.
Jane: This paper gives us a clear path to address the issue where ESL students face penalties because their valid L2 linguistic markers are misinterpreted as evidence of poor reasoning or low quality.
Lu: It confirms that by optimizing the embedding space, we can successfully remove correlations between syntactic complexity and predicted scores for ESL students, treating it as a feature instead of a fault.
Meng: The limitation they flag is that while the method reduces bias significantly, it still exhibits a systematic underrating for both groups, which suggests even this fairer model might be slightly more conservative in its scoring policy.
Tom: So the final thought on "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning" is that we can build systems that are both fairer and sufficiently accurate for deployment by actively optimizing quality signals within the AI's training process.
The paper's summary: Tom: So we’re circling back to this paper on essay scoring bias for ESL learners. What they’re really saying is that instead of trying to fix every single rule in a massive manual, you can just adjust how the AI learns what makes a good essay, focusing on quality itself.
Jane: Exactly. They used contrastive learning to show that by training the AI to group essays based on their actual quality—not just whether they sound like a native speaker—you can significantly reduce that gap between high-achieving ESL students and native writers.
Lu: The main mechanism they used was creating these fairness triplets, where an anchor essay, like a native one, got paired with positive and negative examples from the ESL group based on their actual human scores. It’s forcing the model to learn a better way to understand what high quality looks like across both groups simultaneously.
Meng: From an engineering standpoint, that means we’re not just tweaking some surface-level filter; we’re changing how the AI builds its internal understanding of 'good writing' by training it on paired examples where language origin is deliberately ignored in favor of semantic quality.
Tom: And the results are pretty telling. They found they could cut that initial ten percent score difference down to about six percent without sacrificing much of the overall accuracy, which is a big win for deployment because we don't want to lose that predictive power just for fairness.
Jane: It’s about moving away from systems that penalize complex L2 structures as errors, and instead learning that those structures are actually stylistic features of advanced writing. That’s a huge shift in how we think about what makes an essay strong.
Lu: They even found that sentence complexity, which they thought was a trigger for bias, is actually treated correctly by the contrastive model because it understands the context better now. It stops misinterpreting long sentences as mistakes and starts seeing them as sophisticated structure.
Meng: So what this means practically is that if we build scoring tools, we can start trusting them more because we’ve shown a method to actively remove those unfair shortcuts based on language markers. We’re targeting the latent representations of quality directly.
Tom: It shows that the fairness-accuracy trade-off isn't as bad as some people thought; you can get thirty-nine point nine percent less bias and still keep a solid score metric at zero point seven five six, which is really encouraging for real-world application.
Jane: And though they did find it’s still a bit more conservative for both groups, meaning the model scores slightly lower overall than before, the trade-off seems manageable when you look at the reduction in that massive proficiency disparity.
Lu: The authors pointed out that while this fixes one huge bias issue, there are still some limitations—they said it doesn't fix every single nuance of scoring perfectly, which is expected for any system.
Meng: So the future work they suggest is to look at score recalibration techniques to keep the fair embedding space clean while trying to restore that absolute top-tier accuracy back up. It’s about fine-tuning the final step.
Tom: That’s a really concrete roadmap—fix the representation first, then tune the output layer later—which gives us a clear path forward for developing more equitable essay scoring AI.
The paper's improvements: Tom: So we’re talking about how this approach actually changes things for essay scoring, moving past just identifying where the bias is happening to actually fixing it systematically. What they suggest is using these alignment tools to create a more balanced internal map of what makes an essay high quality across different language backgrounds.
Jane: Exactly. It’s not just about patching the score; it’s about retraining the AI so its definition of 'high quality' isn't secretly biased toward one linguistic style over another. It forces the model to see that a complex sentence structure from an ESL student is still high quality if it meets the same semantic goals as a native writer.
Lu: The real improvement they highlight is disentangling that syntactic complexity issue we talked about earlier. They showed that by using this contrastive training, the AI learns to treat sentence length as just another feature, not a signal for error when writing in English. That’s a huge unlock for ESL students who write with more complex structures.
Meng: From an engineering angle, that means we can design systems where the model ignores those surface-level linguistic markers that were causing the initial penalties and focuses purely on the underlying reasoning or content quality. It makes the scoring process much more robust against those common heuristics.
Tom: And they found this whole method works best when you tune a specific setting, which is interesting because it shows there’s an optimal balance to hit for fairness without losing too much predictive power. They pinpointed that margin one point zero as the sweet spot for that trade-off.
Jane: So, what this means for someone listening is that we can expect tools to become less biased toward specific writing styles, meaning a student's ability to express complex ideas isn't unfairly penalized just because they aren't native speakers.
Lu: The potential here is huge; imagine an AI tutor that scores essays based on genuine thought process rather than just grammar rules or sentence length. It opens up entirely new ways to assess learning in second languages, which is where my head goes with this research.
Meng: For us building these platforms, it means we can deploy systems that are genuinely fairer while still maintaining the accuracy needed for serious educational use cases. It’s a practical step toward making AI tools usable by everyone without putting high-achieving ESL writers at a disadvantage.
Tom: It really shows that this isn't just theoretical math; it’s a tangible process of refining the model to be more equitable in its judgment, and that tuning that margin is the lever we have to pull for practical deployment.
Conclusion: Tom: So we’ve covered how contrastive learning helps reduce that massive score gap for high-proficiency ESL students in this paper, "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning." It boils down to using careful alignment during training to make the AI learn what good writing is, regardless of who wrote it.
Jane: Right. We’ve seen how they used those triplet construction ideas and found that tuning the margin at one point zero really helped them get that fairness boost without hurting overall accuracy too much. It's a nuanced balancing act.
Lu: The big picture here is that we can start building AI scoring systems where the definition of quality is decoupled from language origin, which could fundamentally change how we assess writing across different cultures and educational settings. It opens up new possibilities for equitable learning tools.
Meng: I see the practical implication as moving away from simply trying to patch surface-level errors in the model and instead training it on quality signals that are inherently more stable, making the tool much more reliable for real deployment. It’s about building systems that can handle real-world variability without getting tripped up by language features.
Lalam: If we look at this from a cultural standpoint, this work means we can build tools that don't unintentionally reinforce biases against certain communication styles; it helps create a more inclusive digital culture where writing skills are valued for what they actually communicate.
Tom: It’s a solid blueprint for how to tackle these kinds of subtle biases in large language models applied to education, and I think we’re going to see more of this type of targeted intervention in the coming months.
Jane: Exactly. The "Mitigating Bias in Automated Essay Scoring for ESL Learners via Contrastive Learning" paper shows that when you focus on aligning the representation space based on true quality, you can achieve a much fairer and still highly accurate system.
Lu: It’s exciting because it proves that domain-specific techniques like contrastive learning are really effective at tackling deep, latent biases in language models without needing a complete overhaul of the entire model architecture.
Meng: I think the next step for us is figuring out how to integrate these alignment strategies into our existing agent planning frameworks to see if we can apply this quality-based alignment across other complex tasks as well.
Lalam: This paper gives us a clear direction for making AI systems that support diverse communication styles, and that’s something I think will really help shape the future of how we interact with intelligent tools.
Tom: Well, that wraps up our deep dive into this research on essay scoring bias, but keep an eye out because we’ve got another interesting paper coming up next about how digital personas handle survey findings.
More episodes
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language
- 2508.08833-An Investigation of Robustness of LLMs in Mathematical Reasoning: Benchmarking with Mathematically-Equivalent Transformation of Advanced Mathematical Problems
- 2405.04118-Policy Learning with a Language Bottleneck