AI Wizards at CheckThat! 2025: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles

arXiv:2507.11764 · cs.CL, cs.IR · Submitted 2025-07-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "AI Wizards at CheckThat! 2025".

Jane: This paper presents "AI Wizards’ participation in CLEF 2025 CheckThat! Lab Task 1: Subjectivity Detection in News Articles, classifying sentences as subjective/objective in monolingual, multilingual,

Tom: First, who's behind it and why it matters.

Title and authors: Tom: Welcome back everyone! We've got a fascinating paper today, "AI Wizards at CheckThat! two thousand twenty-five: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles," and we are absolutely buzzing about what they've done to tackle language understanding.

Jane: It really is a deep dive into how AI can understand human opinion in news, and I think the focus on multilingual settings is what makes this so relevant for us all right now.

Lu: This paper shows how we can move beyond just looking at words and start understanding the actual feeling behind those words, which opens up huge possibilities for creative applications.

Meng: From an engineering side, it’s interesting to see them test models like mDeBERTaV3-base alongside Llama3 point 2-1B to see which architectures actually benefit most from this sentiment fusion technique.

Lalam: I think the core idea here is building a bridge between deep language understanding and affective computing, which is incredibly important for creating systems that can grasp complex human communication patterns.

Tom: So, to summarize what they did in this paper, the authors show that their main strategy was to enhance standard transformer classifiers by fusing sentiment scores from an external auxiliary model directly with the sentence representations before the final classification step.

Jane: That means they are essentially giving the AI an emotional compass alongside its dictionary to help it make a better judgment on whether a sentence is subjective or objective. It’s a really intuitive way to think about improving classification accuracy.

Lu: I find that fusion strategy really clever because it allows the model to consider both what the words literally mean and how those words actually feel, which should make it much more robust when dealing with subtle human language. It moves us closer to capturing genuine human nuance instead of just matching patterns.

Meng: The summary also points out that they systematically integrated those sentiment scores as a key feature engineering step, which really shows that careful feature design is just as important as the model architecture itself for getting better results. It’s about making the AI smarter by giving it richer input data to process.

Lalam: I think the summary highlights how they systematically integrated those sentiment scores as a key feature engineering step, which shows that careful feature design is just as important as the model architecture itself for getting better results. It’s about making the AI smarter by giving it richer input data to process.

Title and authors: Tom: And they didn't stop there with their experiments, showing that this feature fusion significantly boosts performance, especially in the subjective F1 score, leading to very high rankings. This is the payoff for all that thoughtful design work we just discussed.

Jane: So, what really stood out in their methodology were the specific improvements they focused on beyond just adding those sentiment scores, which included a post-hoc decision threshold optimization strategy to handle class imbalance. They even used a grid search over thresholds ranging from zero point one to zero point nine to maximize the macro F1 score on the development set.

Lu: That threshold calibration is a really smart maneuver because it directly addresses the uneven distribution of subjective and objective sentences they found in languages like Arabic and Italian, which can seriously skew model performance. It’s about tuning the final decision boundary to be fairer for both classes across different contexts.

Meng: From a practical standpoint, that threshold calibration means the system won't just output a raw probability score; it will give us a more reliable classification based on what the data actually tells us about how opinions are distributed in that specific news context. That reliability is exactly what we need when deploying these tools.

Tom: I totally agree with Meng; that refinement shows they understand that raw probabilities aren't always enough for real-world deployment, especially when class imbalance is a problem. And it’s great to see how they achieved such solid results, like the subjective F1 score gaining significantly for English and Italian.

Jane: It really validates the idea that combining strong base models with task-relevant feature engineering and post-processing is a powerful strategy for solving these kinds of nuanced NLP problems. It proves that we don't have to rely on one single approach to get high quality results.

Lu: And seeing those results, like the Greek result achieving a Macro F1 of zero point five one, is a major success for understanding subjectivity across different languages without needing extensive retraining, which is pretty impressive. It shows the power of this method for cross-lingual generalization.

Title and authors: Lalam: The paper’s suggestion to combine strong base models with task-relevant feature engineering and post-processing really underscores a powerful philosophy: sometimes the best AI isn't just about having the deepest model, but about intelligently augmenting its inputs to handle real-world messiness. It’s an excellent blueprint for how we can build better systems.

Tom: So, wrapping up our discussion on "AI Wizards at CheckThat! two thousand twenty-five: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles," it seems the main point is that combining sentiment augmentation with careful threshold calibration gives us a way to significantly boost performance in multilingual subjectivity detection. The results, especially seeing those gains in English and Italian F1 scores, are really impressive for what this method can achieve.

Jane: Absolutely; the implications here are huge because it shows that we don't have to build separate, specialized models for every language if we can use a unified framework that intelligently incorporates sentiment signals. It opens up possibilities for much more efficient and context-aware content analysis systems.

Lu: I think the future work they suggest—exploring finetuning sentiment models on actual news corpora or using even larger foundation models—points toward a future where AI doesn't just classify; it actually learns the meaning of emotion within journalistic language. That level of learning is what we’re really aiming for in terms of building more sophisticated AI.

Meng: I agree with Lu; if we can train these sentiment models on domain-specific news data, the practical impact on content moderation and automated reporting will be enormous. We need that contextual understanding to build trustworthy AI tools for real use cases.

Lalam: Ultimately, this work demonstrates that thoughtful feature engineering combined with adaptive post-processing is a vital strategy for making AI systems more nuanced and culturally aware, which really enhances the overall quality of information processing in our world. What a significant advancement!

Tom: Fantastic talk on "AI Wizards at CheckThat! two thousand twenty-five: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles". We've seen how this combination of sentiment and threshold calibration can give us better performance, especially in tricky languages like Arabic and Italian.

Jane: It’s a powerful concept—using the insights from this paper to move toward more efficient, context-aware systems that understand the subtleties of human expression better. We should definitely keep an eye on how these ideas translate into real applications.

The paper's summary: Tom: So, to recap what we just went over, this paper is all about taking existing language models and supercharging them by adding an emotional layer—sentiment scores—to help them decide if a news sentence is subjective or objective. It’s not just about reading the words; it's also about understanding the tone behind them.

Jane: Exactly, Tom; think of it like giving a brilliant translator a little emotional context so they can choose the right nuance when translating complex ideas into clear language. The authors show that this extra input makes the AI much better at picking up on those tricky personal opinions or sarcasm in news.

Lu: What really excites me is how they handled the multilingual aspect, proving that this technique isn't just a trick for one language; it works across Arabic, Bulgarian, English, German, and Italian simultaneously. This suggests we can build more universally smart AI tools that don't need to be rebuilt from scratch every time we talk about a new language.

Meng: From an engineering view, the fact that they used mDeBERTaV3-base and Llama3 point two-1B shows us that this sentiment feature is surprisingly effective across different model families, which is a huge win for developers. I'm curious if we can use this same fusion method to speed up the training process for other NLP tasks, not just classification.

Lalam: I think the real impact here is on how we communicate globally; if AI can better grasp the emotional weight of different cultural expressions, it helps bridge those gaps in understanding between people. It’s about making information processing more empathetic and less prone to misinterpretation across borders.

Tom: That's a huge point, Lalam; moving beyond just factual accuracy to capturing the emotional texture of news makes the whole system feel much more human-centric. And their results really back this up with those impressive F1 scores for English and Italian, showing tangible gains in performance.

Jane: It’s so encouraging to see that tangible performance boost; it proves that augmenting a model's input data can lead to real, measurable improvements in how we understand human expression. This isn't just academic research; this is practical improvement for content analysis systems.

Lu: And the threshold calibration part really shows they thought about the real-world messiness of data, specifically dealing with class imbalance in languages where opinions are heavily skewed. That level of detail in handling data distribution is what pushes AI systems from just working to truly being robust and reliable across diverse contexts.

Meng: Exactly; it moves the conversation from theoretical potential to practical deployment readiness, because you can't just throw a model at real-world data without dealing with those distribution issues first. It’s about making sure the AI actually performs well when it matters most in a live environment.

Lalam: Ultimately, this research shows that thoughtful feature engineering combined with smart post-processing is a vital strategy for making AI systems more nuanced and culturally aware, which really enhances the overall quality of information processing in our world. This is where culture and language meet cutting-edge computation to build something genuinely helpful.

Tom: So, we’ve seen how this combination of sentiment and calibration can give us better performance, especially in tricky languages like Arabic and Italian. We've got a lot of exciting implications here for making content analysis systems much more empathetic.

The paper's improvements: Tom: So, to wrap up on the improvements section of "AI Wizards at CheckThat! two thousand twenty-five" it really boils down to two major enhancements: first, they implemented a decision threshold calibration strategy to manage that class imbalance we talked about earlier, and second, they set the stage for future work like fine-tuning those sentiment models on actual news data.

Jane: That threshold calibration is super practical; it means the AI gets smarter about when to trust its own prediction based on how many subjective versus objective sentences it's seeing in a specific context. It’s like giving the system a built-in sense of fairness so it doesn't accidentally favor one type of sentence over another.

Lu: I think that future work on finetuning sentiment models is where things get really wild; if we train those emotion detectors specifically on news corpora, we could unlock entirely new layers of understanding about how different emotions manifest in journalistic writing. That opens up possibilities for AI to truly learn the *meaning* of emotion within a specific domain.

Meng: From my side, that's exactly what I need to see; if we can train those sentiment models on domain-specific data, the practical impact on content moderation and automated reporting will be enormous because it will give us trustworthy tools that understand context better than general sentiment detectors.

Lalam: I think the vision here is profound: AI moving from just classifying text to actually learning the cultural and emotional context embedded in human communication. This progress means we are building systems that can navigate our world with a much deeper sense of understanding and empathy.

Tom: That’s a massive leap, Lalam; it takes us from pattern recognition to true contextual awareness in language processing. And the way they framed those future steps shows they're thinking about building something long-term, not just solving one task for today.

Jane: It really validates that combining a strong base model with these kinds of adaptive strategies—both during training and after—is the way forward for tackling complex NLP problems in a practical setting. It’s about making the AI adaptable to real-world noise.

Lu: And I think this whole paper lays down a great blueprint for how we should approach other nuanced tasks, showing that feature engineering combined with post-processing is a solid path toward more sophisticated, context-aware AI solutions.

Meng: That blueprint is useful because it shows us the pipeline: model choice, feature fusion, tuning the threshold, and then adapting it to specific data distributions. It’s a clear roadmap for building reliable systems in production environments.

Lalam: Ultimately, this research demonstrates that thoughtful feature engineering combined with adaptive post-processing is a vital strategy for making AI systems more nuanced and culturally aware, which really enhances the overall quality of information processing in our world. That’s the kind of advancement we want to see happening everywhere.

Tom: Fantastic summary; it seems like the future isn't just about bigger models, but about smarter ways to feed those models richer, context-aware data and tuning how they make their final decisions. We’ve seen how this combination of sentiment and threshold calibration can give us better performance, especially in tricky languages like Arabic and Italian.

Conclusion: Tom: Alright team, we're wrapping up our discussion on "AI Wizards at CheckThat! two thousand twenty-five: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles." We’ve seen how adding sentiment scores and calibrating the threshold gives us a way to significantly boost performance in multilingual settings.

Jane: It really shows that we can build systems that are more efficient and context-aware by intelligently incorporating those sentiment signals, opening up possibilities for much better content analysis.

Lu: I think the future work they suggest about finetuning those sentiment models on real news data points toward a future where AI doesn't just classify; it actually learns the meaning of emotion within journalistic language. That level of learning is what we’re aiming for in terms of building more sophisticated AI.

Meng: I agree with Lu; if we can train those sentiment models on domain-specific news data, the practical impact on content moderation and automated reporting will be enormous because we need that contextual understanding to build trustworthy tools.

Lalam: Ultimately, this work demonstrates that thoughtful feature engineering combined with adaptive post-processing is a vital strategy for making AI systems more nuanced and culturally aware, which really enhances the overall quality of information processing in our world. What a significant contribution!

Tom: Fantastic talk on "AI Wizards at CheckThat! two thousand twenty-five: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles." We've seen how this combination of sentiment and threshold calibration can give us better performance, especially in tricky languages like Arabic and Italian.

Jane: It’s a powerful concept—using the insights from this paper to move toward more efficient, context-aware systems that understand the subtleties of human expression better.

Lu: I’m really looking forward to seeing how these ideas translate into other areas, like applying sentiment fusion to visual or time-series data we're seeing in other papers on arXiv.

Meng: I hope we see more of this kind of robust feature engineering applied to real-world engineering problems where reliability is everything.

Lalam: I think the paper’s focus on cultural awareness through language processing is a huge step forward for how AI interacts with human communities globally.

Matteo Fasulo, Luca Babboni, Luca Tedeschini

Department of Computer Science and Engineering - University of Bologna

cs.CL, cs.IR

Submitted: 2025-07-15

Updated: 2025-07-15

Code: https://github.com/MatteoFasulo/clef2025-checkthat

Importance score: 72/100

The gist: This paper presents "AI Wizards’ participation in CLEF 2025 CheckThat! Lab Task 1: Subjectivity Detection in News Articles, classifying sentences as subjective/objective in monolingual,

Key concepts

Sentiment Fusion
This technique involves adding sentiment scores from an external auxiliary model directly into the sentence representations before the final classification step. It gives the AI an emotional compass alongside its word meanings to help it better judge if a sentence is subjective or objective.
Decision Threshold Calibration
This is a strategy where models are tuned by adjusting decision thresholds, such as searching between 0.1 and 0.9, to handle uneven distribution of classes. This helps ensure the system makes fairer classifications across different contexts by managing class imbalance.
Feature Engineering
This refers to the process of systematically integrating sentiment scores as a key feature engineering step. It shows that careful design of richer input data is crucial for making AI smarter, regardless of the underlying model architecture.
Post-hoc Decision Threshold Optimization
This involves optimizing the final decision boundary after initial classification by using grid search over thresholds. This refinement addresses real-world data distribution issues, leading to more reliable classifications based on actual context.

Terminology

Summary

This paper presents AI Wizards’ participation in CLEF 2025 CheckThat! Lab Task 1: Subjectivity Detection in News Articles, classifying sentences as subjective/objective in monolingual, multilingual, and zero-shot settings. The primary strategy explored was to enhance transformer-based classifiers by integrating sentiment scores derived from an auxiliary model with sentence representations. This sentiment-augmented architecture was tested on mDeBERTaV3-base, ModernBERT-base (English), and Llama3.2-1B. To address class imbalance prevalent across languages, decision threshold calibration optimized on the development set was employed. Experiments showed that sentiment feature integration significantly boosts performance, especially the subjective F1 score, leading to high rankings; notably 1st for Greek (Macro F1 = 0.51).

The work addresses subjectivity detection as defined in Task 1 [1] of the CLEF 2025 CheckThat! Lab, which challenges systems to classify sentences from news articles as subjective (SUBJ) or objective (OBJ). The system fine-tunes transformer-based models by strategically fusing external sentiment information with sentence representations before classification. This strategy was evaluated on mDeBERTaV3-base [5, 6], ModernBERT-base [7], and Llama3.2-1B [8] fine-tuned with and without the sentiment feature fusion for multilingual subjectivity detection. The systematic integration of sentiment scores as a key feature engineering step demonstrated its impact on improving subjective content classification. The application of decision threshold calibration was used to mitigate class imbalance inherent in the provided datasets, further refining performance.

The dataset consists of sentences extracted from news articles across five languages: Arabic (AR), Bulgarian (BG), English (EN), German (DE), and Italian (IT). The annotation guidelines define subjective sentences as those expressing personal opinions, sarcasm, exhortations, discriminatory language, or rhetorical figures conveying an opinion, while objective sentences include factual statements, reported third-party opinions, open-ended comments, and factual conclusions. An analysis of the label distribution reveals a notable class imbalance across all languages.

The model architectures explored include:

mDeBERTaV3-base:

ModernBERT-base:

Llama3.2-1B:

Sentiment augmentation involved predicting sentiment using an external pre-trained multilingual sentiment analysis model, twitter-xlmroberta-base-sentiment [18], which outputs a three-dimensional vector for positive, neutral, and negative sentiment. These three scores were then concatenated with the [CLS] token embedding before passing it to the final classification layer.

Data preprocessing involved tokenization using model-specific tokenizers and applying padding/truncation to a maximum sequence length of 256 tokens. For Arabic experiments, an additional strategy involved translating the Arabic data into English using Helsinki-NLP/opus-mt-ar-en [19, 20] prior to fine-tuning, though this did not ultimately lead to improved performance.

Training utilized the AdamW optimizer with a linear learning rate scheduler and warmup, employing Cross-Entropy Loss with class weights initially. The best checkpoint was selected based on development set performance. To address class imbalance, a post-hoc decision threshold optimization strategy was implemented: the model is trained on the training set using cross-entropy loss, and an optimal decision threshold is determined by conducting a grid search over values ranging from 0.1 to 0.9 (0.01 increment) aiming to maximize the macro F1 score on the development set, which is then applied to the model’s softmax outputs for classification on the test set.

Experiments were conducted for monolingual, multilingual, and zero-shot subjectivity detection subtasks defined by CLEF 2025 CheckThat! Lab Task 1. Evaluation primarily focused on macroaverage F1 and SUBJ F1 scores. The results demonstrated that adding sentiment features consistently improved SUBJ F1 scores across most languages, with notable gains for English (0.4046 to 0.5279) and Italian (0.6291 to 0.6804). Decision threshold calibration led to substantial improvements in both Macro F1 and SUBJ F1 scores for languages with significant class imbalance like Arabic and Italian, while for more balanced languages, the gains were marginal or standard thresholding performed slightly better by one metric. Performance on Arabic remained a consistent challenge across all tasks. The work concluded that combining strong base models with task-relevant feature engineering (sentiment augmentation) and post-processing (threshold calibration) is valuable for nuanced NLP problems in multilingual contexts. The team achieved high rankings, notably 1st place for Greek (Macro F1 = 0.51). Future work suggested exploring finetuning sentiment models on news corpora, leveraging larger LLMs, and investigating more sophisticated attention-based fusion mechanisms.

Improvements for AI systems

As a fastidious and diligent AI researcher, I have analyzed this paper, AI Wizards at CheckThat! 2025: Enhancing Transformer-Based Embeddings with Sentiment for Subjectivity Detection in News Articles.

The core contribution of this work is the successful integration of sentiment scores as an auxiliary feature to enhance transformer-based models (specifically mDeBERTaV3) for multilingual subjectivity detection, coupled with a robust decision threshold calibration strategy.

Here are the specific improvements that can be made to AI systems based on this research, and what the improved system could achieve:


)

The improved AI system would be a specialized, high-performance Multilingual Subjectivity Detection Pipeline capable of classifying news sentences as Subjective (SUBJ) or Objective (OBJ).

Specific improvements include:

  1. [Sentiment Feature Fusion Architecture]: Implement a late-fusion architecture where the output embeddings from the base transformer (e.g., mDeBERTaV3) are concatenated with explicit sentiment scores derived from a pre-trained multilingual sentiment model.

  2. [Language-Specific Sentiment Modeling]: Develop language-specific adaptations for the auxiliary sentiment model, potentially through fine-tuning it on domain-specific news corpora to capture nuances missed by general Twitter/web data (as suggested in Section 8).

  3. [Adaptive Threshold Calibration Module]: Integrate a dynamic decision threshold calibration module that optimizes the classification boundary specifically for each language and dataset split (Training, Dev, Test) using a grid search over the Macro F1 score on the development set.

  4. [Sentiment-Aware Attention Mechanism (Future Work)]: Transition from simple feature concatenation to an attention-based fusion mechanism where the model learns to dynamically weigh the importance of sentiment signals versus semantic content during classification.

The improved AI system can achieve the following:

  1. [Significantly Higher Accuracy in Imbalanced Scenarios]: By leveraging sentiment features and adaptive threshold calibration, it will substantially boost performance (as shown by the reported gains in SUBJ F1 scores, e.g., English: 0.4843 to 0.5279), particularly for minority classes across languages where class imbalance is pronounced (Arabic, Italian).

  2. [Superior Multilingual Generalization]: The system will demonstrate enhanced cross-lingual transfer capabilities by successfully leveraging sentiment signals in zero-shot settings (as seen in Table 5 results) and performing better when excluding known challenging languages (like Arabic) from the training set.

  3. [Contextual Nuance Detection]: It will be able to correctly identify subjective statements that rely on emotional cues, such as strong negative sentiments or rhetorical figures, by utilizing the explicit sentiment signal as a powerful disambiguating feature for the transformer's internal representation.

  4. [Robustness Against Linguistic Variation]: The system will maintain high performance across diverse linguistic expressions within a single model framework (mDeBERTaV3), reducing the need to maintain separate, language-specific fine-tuned models for every language.

Sources

Related papers