AIWizards at MULTIPRIDE: A Hierarchical Approach to Slur Reclamation Detection
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.
Jane: Today's paper: "AIWizards at MULTIPRIDE".
Tom: Detecting reclaimed slurs represents a fundamental challenge for hate speech detection systems,
Jane: First, who's behind it and why it matters.
Title and authors: Tom: So Jane, let's talk a bit more about the title and who cooked this up with "AIWizards at MULTIPRIDE: A Hierarchical Approach to Slur Reclamation Detection." It really tells you that they aren't just looking at the text in isolation; they are using a structured method to tackle reclaimed slurs.
Jane: Exactly, Tom, it highlights that the core challenge is recognizing when a slur shifts from being abusive to being an affirmation based on context. The authors are proposing this hierarchical approach as their main solution to solve that ambiguity.
Lu: The team includes Luca Tedeschini and Matteo Fasulo, and they’re coming from some really strong academic backgrounds in data science and computer engineering at places like ETH Zürich and the University of Bologna, which tells you the research has a solid technical foundation.
Meng: From an engineering standpoint, seeing that they're tackling Subtask B of the MultiPRIDE shared task shows they are working on real-world data challenges from actual moderation efforts, which is what matters most for practical impact.
Lalam: It’s exciting to see researchers focusing on these specific, subtle social cues; it suggests that the future of AI in online spaces needs to be much more sensitive to identity and context than just looking at word frequency.
The paper's summary: Tom: Now, let's get into what the paper actually does. Basically, they outline a two-stage process: first, they use a weakly supervised Large Language Model to assign fuzzy labels about whether a user might belong to the LGBTQ+ community based on their tweets and bios.
Jane: So it’s not just reading the tweet; they’re using contextual information from the user's profile data to predict community membership. That output then feeds into a second stage where they train a BERT-like model to decide if the slur is reclaimed or offensive.
Lu: The paper strongly suggests that this decomposition is inspired by existing research showing a direct link between who you are as a user and how you use certain language, especially regarding slurs. They hypothesize that getting identity signals from biographies is simpler than trying to infer deep social context just from the short tweets themselves.
Meng: That makes sense; using biographical data for identity clues is often more direct than trying to find those deep social signals solely within short text snippets, which is a practical consideration for building systems that can actually run smoothly.
Lalam: It’s like having two specialized experts working together: one who figures out the user's background and another who analyzes the language in relation to that background, and that collaboration should create a much better final result than either model could achieve alone.
The paper's improvements: Tom: The real clever part of this work is how they put those two separate ideas together. They don't just run the user identification model and the detection model separately; they integrate the user identity information into the slur detection model using a learned gating mechanism.
Jane: That gating mechanism is really smart, Tom; it dynamically decides whether the system should lean more on what it learned about the tweet itself or more on what it learned about that specific user's context. It’s like giving each piece of information a different volume control during classification.
Lu: They use a dual-encoder architecture where one encoder is trained specifically for slur detection, and the other is trained to pick up broader social signals through that auxiliary task we just discussed. Then they combine those two representations using this learned gate.
Meng: From an implementation standpoint, having this fusion layer means we don't have to choose one signal or the other; the system learns how to weigh them best for any given input, which is a huge step toward building flexible applications that can adapt.
Lalam: This modular design is fantastic because it means if we later find better ways to infer user identity signals, we can swap out that part easily without having to rebuild the entire slur detection engine from scratch.
Conclusion: Tom: So, wrapping up this discussion on "AIWizards at MULTIPRIDE: A Hierarchical Approach to Slur Reclamation Detection," the main point is that breaking down a complex problem into stages based on user identity is a highly effective way to detect reclaimed slurs. They found it works statistically equivalent to strong BERT-based baselines, which is impressive given the complexity involved.
Jane: Exactly, Tom; they demonstrated that by incorporating author-level context derived from self-descriptions, we can significantly improve accuracy in telling the difference between harmful use and reclaimed usage. The implication for us is that future AI systems shouldn't just look at words in isolation when dealing with sensitive topics like hate speech.
Lu: I think the biggest impact here is demonstrating a viable architectural framework for modeling sociolinguistic context; it shows that hierarchical modeling isn't just theoretical, it’s practical and powerful for handling this kind of nuance.
Meng: For practical deployment, this means we can build systems that are much less likely to flag harmless in-group language as abuse, which could drastically reduce false positives in online moderation tools across the board.
Lalam: I’m really excited about how this work can improve culture because if AI systems can better understand the subtle social signals behind language, it helps create safer and more inclusive digital spaces for everyone to use without fear.
Tom: And that's a fantastic summary of what makes "AIWizards at MULTIPRIDE: A Hierarchical Approach to Slur Reclamation Detection" so impactful! We'll keep an eye on how this hierarchical approach evolves in other detection tasks.
Jane: Agreed, Tom; it’s a major step forward for making AI interactions feel more human and contextually aware. Thanks for joining us today!
Luca Tedeschini, Matteo Fasulo
Villanova.ai S.P.A. · Swiss Data Science Center · Department of Computer Science and Engineering (DISI), University of Bologna
cs.CL
Submitted: 2026-02-13
Updated: 2026-02-13
Code: https://github.com/LucaTedeschini/multipride
Importance score: 79/100
The gist: Detecting reclaimed slurs represents a fundamental challenge for hate speech detection systems, as "the same lexical items can function either as abusive expressions or as in-group affirmations
Key concepts
- Hierarchical Approach
- This method breaks down the complex problem of detecting reclaimed slurs into stages based on user identity. It involves using contextual information from a user's profile data to predict community membership first, which then informs the second stage of slur detection.
- Weakly Supervised LLM
- A large language model is used in the first stage to assign fuzzy labels about whether a user belongs to an LGBTQ+ community based on their tweets and bios. This step uses contextual information from profile data rather than just reading the text alone.
- Learned Gating Mechanism
- This is an integration technique where user identity information is combined with the slur detection model. The mechanism dynamically decides whether the system should prioritize information learned from the tweet itself or more on that specific user's context, allowing for flexible weighting.
Terminology
Summary
Detecting reclaimed slurs represents a fundamental challenge for hate speech detection systems, as the same lexical items can function either as abusive expressions or as in-group affirmations depending on social identity and context.
This work addresses Subtask B of the MultiPRIDE shared task at EVALITA 2026 by proposing a hierarchical approach to modeling the slur reclamation process.
The core assumption is that members of the LGBTQ+ community are more likely, on average, to employ certain slurs in a reclamatory manner.
Based on this hypothesis, the task is decomposed into two stages. First, using a weakly supervised LLM-based annotation, we assign fuzzy labels to users indicating the likelihood of belonging to the LGBTQ+ community,
which is inferred from the tweet and user bio. These soft labels are then used to train a BERT-like model to predict community membership, encouraging the model to learn latent representations associated with LGBTQ+ identity.
In the second stage, we integrate this latent space with a newly initialized model for the downstream slur reclamation detection task.
The intuition behind this is that the first model encodes user-oriented sociolinguistic signals, which are then fused with representations learned by a model pretrained for hate speech detection.
The approach is motivated by the hypothesis that "there exists a strong association between user identity and the reclamatory use of slurs. In particular, users who self-identify as part of, or closely aligned with, the LGBTQ+ community may be more likely to employ certain slurs in a reclamatory manner compared to users who are not. Furthermore, it is hypothesized that
inferring such identity-related signals is comparatively easier when leveraging user biographies, where individuals often explicitly express aspects of their identity through self-descriptions, pronouns, symbols, or community-related markers."
The methodology involves defining an auxiliary proxy variable called the LGBTQ+ affiliation signal,
which is used exclusively to train the user encoder. This label is automatically assign[ed] based on the concatenation of a user’s tweet and biography using an instruction-tuned Large Language Model (LLM), specifically DeepSeek-V3.2.
This process is described as using weak supervision for representation learning, not as a user-level prediction target,
with the objective being to encourage the model to encode latent social and contextual signals useful for reclamation detection, without claiming accurate identification of individuals or verification of protected attributes.
The resulting auxiliary label distribution is reported in Table 3.
The system description involves formulating slur reclamation detection as a binary classification task whose interpretation depends on both (i) the local linguistic context of the tweet and (ii) author-level contextual cues derived from user selfdescriptions,
modeled through a dual-encoder architecture with gated representation level fusion.
This architecture consists of two parallel BERT-like encoders that process identical inputs, optimized under different training signals. One encoder is optimized directly for slur reclamation detection, while the other is encouraged to encode broader user-related sociolinguistic signals via an auxiliary proxy task.
The representations are subsequently combined through a learned gating mechanism
defined by:
g = sigma (W2 · tanh (W1 · [htext:: huser]))
hfused = g ⊙ htext + (1 − g) ⊙ huser
The training procedure consists of three stages: "(i) fine-tuning a baseline text encoder on the slur reclamation task, (ii) training the user encoder on the auxiliary proxy task to induce user-oriented representations, and (iii) training the fused dual-encoder model on the primary reclamation task, jointly optimizing the gating mechanism, classifier, and both encoders."
Experimental results show that the proposed architecture achieves performance statistically equivalent to a strong BERT-based baseline
on both Italian and Spanish. Specifically, Table 6 shows that for the development set (macro F1), the Dual Encoder model's F1 scores are 0.90 ± 0.03
for Italian and 0.67 ± 0.04
for Spanish, which are statistically not significantly different from the Baseline scores (p-values > 0.28 and p-values > 0.17, respectively). Performance on the official test set also shows that the dual-encoder model maintains performance comparable to the baseline.
The authors conclude that this suggests that hierarchical and representation-based modeling of sociolinguistic context is a viable direction for slur reclamation detection
and emphasize that the primary contribution of this work lies in the architectural framework rather than in immediate performance gains.
They further argue that the modular nature of the approach allows for future extensions, such as alternative fusion mechanisms, more refined proxy tasks, or the integration of additional contextual encoders.
A limitation identified is that contextual signals, if not sufficiently constrained, may override strong local indicators of abuse,
leading to occasional misclassifications where models "overestimate the likelihood of reclamatory usage when specific slur terms are present.
Improvements for AI systems
Based on the scientific paper AIWizards at MULTIPRIDE: A Hierarchical Approach to Slur Reclamation Detection,
here are specific improvements that can be made to existing AI systems, along with what those improved systems could achieve:
-
The system should implement a hierarchical modeling approach for hate speech detection by decomposing the task into two distinct stages:
-
Inferring user-level sociolinguistic signals (e.g., potential LGBTQ+ community membership) using weakly supervised Large Language Models (LLMs) on user biographies and tweets, resulting in
soft labels.
-
Integrating these latent representations from the user encoder with a secondary model pre-trained for hate speech detection.
-
Employ a learned gating mechanism to dynamically balance reliance between the tweet's local linguistic context (encoded by the primary text encoder) and the user's sociolinguistic context (encoded by the user encoder).
This improved AI system can:
-
Detect reclaimed slurs with higher contextual accuracy because it moves beyond surface-level lexical analysis.
-
Identify instances where a slur is used in an in-group, non-derogatory, or empowering manner, which standard models often misclassify as abusive due to the term's inherent toxicity.
-
Provide a modular and extensible framework for incorporating future contextual signals (e.g., multimodal content or social network information) without needing to retrain the core hate speech detector from scratch.
-
Be more robust against ambiguities in text-only classification by leveraging author-level cues, leading to fewer false positives on reclaimed language while maintaining high recall for harmful uses.
Sources
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering