Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring
summary
The gist
The paper proposes RA-DPO (Reliability-Aware Direct Preference Optimization), an extension of DPO (Rafailov et al., 2023) that integrates annotator agreement, model confidence, and token-level
In short
The episode discusses a paper titled "Reliability-Aware Sexism Detection," which combines DPO with annotator agreement and token-level confidence scoring. The authors propose a unified framework for AI that is reliable in subjective tasks like sexism detection. Key findings include efficient training using only the most reliable data and implementing principled abstention at inference, resulting in significantly higher accuracy.
Key concepts
- Reliability-Aware DPO (RA-DPO)
- This is a unified framework combining model confidence, human annotator agreement, and token-level uncertainty. It allows the AI to calculate a single reliability score for each post. This score determines whether the model should provide an answer or abstain, ensuring the AI knows its own limits.
- Annotator Agreement
- This measures how much human experts agree on a subjective task, like whether a tweet is sexist. The paper uses this agreement as a signal: high agreement indicates reliability, while disagreement signals ambiguity.
Terminology used across episodes
This episode discusses
- Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring · Paper Radio
- The Llama 3 Herd of Models · Paper Radio
- GPT-4o System Card
- Exploring the Influence of Label Aggregation on Minority Voices: Implications for Dataset Bias and Model Training
- Qwen2.5 Technical Report
The paper
Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring · Read on arXiv
Hadi Mohammadi, Shihan Wang, Masoume M. Raeissi, Anastasia Giachanou
Utrecht University · Wageningen University & Research
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring".
Jane: The paper was written by Hadi Mohammadi, Shihan Wang, Masoume M. Raeissi and Anastasia Giachanou from Utrecht University and Wageningen University & Research.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title: Tom: Welcome back, everyone. Today we're looking at a paper with a real mouthful of a title — "Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring." Jane, I have to say, just reading that title tells me this is about making AI more careful, not just smarter.
Jane: Exactly, Tom. And that's such an important shift. For a long time, we've been obsessed with getting AI to be more accurate — but this paper is asking a different question: when should the AI admit it doesn't know? And that's a huge deal for something as subjective as detecting sexism online.
Tom: Right, because sexism isn't like math. Two people can read the same tweet and genuinely disagree on whether it's offensive. The paper uses a dataset called EXIST two thousand twenty-three where every post is labeled by six different people.
Jane: And that's the key insight — instead of treating those disagreements as noise and just taking a majority vote, the researchers treat disagreement as information. If all six annotators agree, that's a strong signal. If they're split three to three, that tells you the post is genuinely ambiguous.
Lu: And that's where the title's "reliability-aware" part comes in. I'm Lu, by the way. The authors build a single score for each post that combines three things: how confident the model is, how much the human annotators agreed, and a token-level uncertainty signal. It's like asking three different experts before making a decision.
Meng: But here's the practical question I always have — does this actually change how the system behaves? Because I can build a fancy score, but if it doesn't improve results, it's just academic.
Jane: Great question, Meng. And the answer is yes, it does. The paper shows that this reliability score can be used in two ways. First, during training, to pick only the most reliable examples — and they found you can use just thirty percent of the data and match the performance of using everything.
Tom: That's a massive efficiency gain. And the second use is even cooler — at inference time, the model can choose to abstain. If the reliability score is low, it just says "I'm not sure" instead of forcing an answer.
Lu: Which is exactly what you want in a real-world moderation system. You'd rather have a human review a borderline case than have the AI make a confident wrong call.
Meng: So the model gets to say "pass" — and that trades coverage for accuracy. The paper reports ninety-six point two percent accuracy at fifty percent coverage, which is a huge jump from the eighty-two point eight percent at full coverage.
Jane: And that's the promise here — not just a better classifier, but a classifier that knows its own limits. Stick around, because next we're going to dig into the actual summary and what the authors claim to have achieved.
Summary: Tom: So we're back with "Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring." Jane, let's talk about what the authors actually set out to do — the core problem they're attacking.
Jane: The core problem is that sexism detection is subjective, but most AI systems pretend it isn't. They take six annotators, collapse them into one majority label, and then train the model as if that label is ground truth. The authors say that throws away useful information.
Lu: And it's not just about the training data. The paper points out that most existing methods use either annotator agreement OR model confidence, but never both together. And they definitely don't use that signal at inference time to decide when to abstain.
Meng: So they're proposing a unified framework — one score that combines everything and gets used consistently throughout. That's elegant, I'll give them that. But what does the actual pipeline look like?
Jane: So they call it RA-DPO — Reliability-Aware Direct Preference Optimization. DPO is a training method where you show the model a good answer and a bad answer, and teach it to prefer the good one. Here, the "good" answer is the majority label, and the "bad" one is the opposite.
Tom: But instead of training on all five thousand five hundred thirty-six pairs equally, they rank them by that reliability score and only train on the top subset. Smart-thirty percent means the top one thousand six hundred sixty-one most reliable pairs.
Lu: And the remarkable finding is that Smart-thirty percent matches full-data DPO exactly — both hit an F1 of zero point eight two one. So you get the same performance with seventy percent less training data.
Meng: That's a real cost saving. But I'm curious — what makes those top pairs so special? Is it the model confidence, the annotator agreement, or something else?
Jane: The paper does a composition analysis, and it's striking. The top thirty percent of pairs are exclusively unanimous — all six annotators agreed. Not a single five-to-one or four-to-two split in there.
Tom: And that explains why adding more data doesn't help. Once you have all the unanimous examples, adding the noisier ones just cancels out the benefit. The signal is already saturated.
Lu: They even test this with a negative control — training only on the ambiguous three-to-three split examples. That model performs terribly, dropping from zero point seven two three to zero point six five three F1. So ambiguous examples aren't just unhelpful, they're actively harmful for training.
Meng: So the takeaway is that high-agreement examples carry the real signal, and the reliability score is a simple way to find them.
Jane: Exactly. And that's the training side. But the bigger payoff comes at inference, which we'll dig into next — how the model decides when to abstain.
Improvements: Tom: Welcome back. We're still on "Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring." Jane, we talked about the training improvements — now let's get into the inference side, which I think is the real star here.
Jane: Absolutely, Tom. The training efficiency is nice, but the inference-time abstention is the game-changer. The model computes that same reliability score R(x) for each new post, and if it falls below a threshold, the model simply refuses to answer.
Meng: So instead of forcing a label on something ambiguous, it flags it for human review. That's how you'd actually deploy this in production — you don't want an AI confidently calling a borderline post sexist when six humans couldn't agree on it.
Lu: And the numbers back that up. At fifty percent coverage — meaning the model answers half the posts and abstains on the other half — accuracy jumps to ninety-six point two percent in the true-agreement setting. That's a thirteen point four percentage point gain over full coverage.
Tom: But here's the catch — in the real world, you don't have six annotators sitting around to compute that agreement score. So the paper tests a predicted-agreement setting, where a separate model guesses the agreement from the text alone.
Jane: And even then, it hits eighty-eight point seven percent accuracy at fifty percent coverage, which is still three point four points above the no-agreement baseline of eighty-five point three percent. So even with a noisy estimate of agreement, the system is more reliable than ignoring agreement entirely.
Meng: That's the deployable version. But I want to know — how do they set that abstention threshold? Is it arbitrary or learned?
Lu: They sweep the threshold from zero point three to zero point nine five and pick the value that maximizes the harmonic mean of accuracy and coverage. So it's a principled trade-off — you tell the system how much you value accuracy versus how many posts you want it to handle.
Tom: And the really interesting part is that this works across different model families. They tested OpenAI's GPT-4o, plus two open-weight 3B models — Llama and Qwen. The pattern holds on all three.
Jane: Right, the absolute numbers are lower on the smaller models — they start with weaker bases — but the reliability filtering still gives a solid boost. Qwen goes from zero point six three four to zero point six nine seven at fifty percent coverage, and Llama from zero point six three six to zero point seven one one.
Meng: So the method is model-agnostic. That's what you want in a real system — you're not locked into one vendor.
Lu: And there's a deeper point here. The paper shows that the model's confidence becomes better calibrated after this training. The gap between reported confidence and actual F1 shrinks from zero point two four on the base model to zero point one zero on RA-DPO. The model is learning to be honest about its uncertainty.
Tom: And that honesty is exactly what we need to talk about next — the first page of the paper lays out this whole vision. Let's get into that.
First Page: Tom: We're back on "Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring." Jane, let's look at the opening page — the abstract and the figure that sets up the whole framework.
Jane: The abstract makes a really clear statement: most systems reduce multi-annotator labels to a single majority decision and treat all instances uniformly. And that ignores two informative signals — annotator agreement and model uncertainty. That's the whole motivation in one sentence.
Lu: And the figure on that first page is a nice visual summary. It shows a tweet — "Women in politics are too emotional to lead" — with five annotators saying YES and one saying NO. The model says YES with zero point eight seven confidence, and the agreement is zero point eight three.
Meng: And then they compute that reliability score R(x) as a weighted combination — confidence, agreement, and one minus the token uncertainty. In that example it comes to zero point eight three. So the system would answer, because that's above the threshold.
Tom: But what if the agreement had been lower? Say three YES and three NO — then the reliability score would drop, and the system would abstain. That's the beauty of the design.
Jane: And the abstract makes another bold claim — that training on the top thirty percent most reliable pairs matches full-data DPO. That's the efficiency finding we talked about, and it's right there on the first page.
Lu: I also like that they're explicit about the three settings they evaluate — true agreement as an upper bound, predicted agreement for real deployment, and no agreement as a floor. That's honest science — they're not hiding the fact that the perfect signal isn't available in practice.
Meng: And the predicted-agreement setting uses a Twitter-pretrained XLM-RoBERTa model to guess agreement from text. It only gets a Pearson correlation of zero point three five one with true agreement — so it's not great — but it's still enough to improve abstention decisions.
Tom: That's a really important point, Meng. Even a weak proxy for agreement helps. You don't need perfect information to make better decisions.
Jane: And the paper frames this as a step toward more reliable deployment in subjective classification. It's not just about sexism — this framework could apply to any task where human judgment varies.
Lu: Absolutely. Hate speech, misinformation, sentiment — anywhere you have disagreement, this reliability-aware approach has value.
Tom: So we've got efficient training, calibrated confidence, and principled abstention. Next up, we're going to wrap this up and see what it all means for the bigger picture.
Conclusion: Tom: Alright, we're wrapping up our discussion of "Reliability-Aware Sexism Detection: Combining DPO with Annotator Agreement and Token-Level Confidence Scoring." Jane, what's the big picture here?
Jane: The big picture is that this paper gives us a complete framework — one reliability score that works for both training and inference. You use it to pick the best training data, and you use it to decide when the model should abstain. That's rare — most papers focus on one or the other.
Lu: And the results are genuinely practical. You can cut training data by seventy percent without losing performance. At inference, you can hit ninety-six percent accuracy by answering only half the posts. And the method works across three different model families — OpenAI, Llama, and Qwen.
Meng: From an engineering standpoint, the predicted-agreement setting is what makes this deployable. You don't need human annotators at runtime — you just need a small model that estimates agreement from text. Even a rough estimate gives you a meaningful boost.
Tom: And let's not forget the negative result — training on ambiguous examples actually hurts. That's a warning to anyone who thinks more data is always better.
Jane: Exactly. And there's a broader cultural implication here. This paper is about building AI that knows its limits. In a world where AI is increasingly used to moderate content, having a system that says "I'm not sure, let a human decide" is a safeguard against over-confident mistakes.
Lu: The authors even mention that in their ethical considerations — abstention acts as a built-in safeguard. For consequential decisions, you route uncertain cases to human review.
Meng: And the limitations are honest too — they only tested on short social media posts. The token-level uncertainty signal might behave differently on longer text. So there's clear room for future work.
Tom: Well said. So we've got efficient training, calibrated confidence, and principled abstention — all wrapped into one framework. That's a solid contribution to making AI more trustworthy in subjective tasks.
Jane: And with that, we're saying goodbye to "Reliability-Aware Sexism Detection." Thanks for joining us, and we'll see you next time with another paper from the arXiv.
Tom: Take care, everyone.
More episodes
- 2610.10768-Strategic Investment Decision Making for Value Creation in Energy Transition: A Reinforcement Learning Approach
- 2610.10858-RFChipAgent: Multi-Agentic AI Flow for Analog/RF Chip Design
- 2610.10613-Temporal transformer CAN encoder with federated lightweight heads for anomaly detection
- 2610.10616-When Routing Reveals Membership: Privacy Leakage from MoE Router Telemetry
- 2610.10655-Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
- 2610.11031-Language Modeling is Monotone Compression
- 2610.01253-Context-Aware Error Mitigation Orchestration for Hybrid Quantum Reinforcement Learning on NISQ Systems
- 2604.24201-CMGL: Confidence-guided Multi-omics Graph Learning for Cancer Subtype Classification
- 2609.34069-Towards Certificate-Driven Software Porting: A Self-Improving Agentic Harness for Scientific Program Optimization
- 2312.01221-Enabling Quantum Natural Language Processing for Hindi Language