Multilingual Models for Check-Worthy Social Media Posts Detection
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Multilingual Models for Check-Worthy Social Media Posts Detection".
Jane: The paper was written by Sebastian Kula and Michal Gregor from Kempelen Institute of Intelligent Technologies and Kazimierz Wielki University of Bydgoszcz, Poland.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary: Tom: We were just talking about the massive scope, and now we're moving into the summary part of "Multilingual Models for Check-Worthy Social Media Posts Detection." The paper seems to dig into how these models actually process and summarize the content to find those problematic claims.
Jane: Right. If you look at what they summarized, it emphasizes that simply having a big model isn't enough; you have to structure the analysis around *why* something might be dubious—is it verifiable? Is it harmful?
Lu: What I found really fascinating in the summary is how they are using these models not just as binary classifiers, but almost as reasoning engines to flag potential issues. It suggests a multi-faceted approach to detection.
Meng: That brings up a practical concern for me, Lu. When you talk about flagging potential issues, how do they handle ambiguity? A post that is technically unverified might just be an opinion, and the model needs to distinguish between those two things without heavy human labeling.
Lalam: The goal of improving information flow isn't to silence speech; it's to improve the quality of public discourse. The summary suggests a pathway toward empowering users with better contextual understanding, which is vital for a healthy culture.
Tom: So, Jane, stepping back from the tech talk for a second, what’s the core message they are conveying in this summary regarding how to actually use these models?
Jane: They're showing that by combining different types of detection—say, checking if a claim is factually verifiable versus whether it promotes harm—you get a much richer picture than just looking at one dimension alone.
Tom: It sounds like they are treating misinformation as a combination of factors, not just one single error type.
Lu: Precisely. The model needs to be sophisticated enough to weigh those different flags against each other, rather than just outputting a single risk score based on the easiest trigger it finds.
Meng: From an implementation standpoint, that layered approach means the system has to be incredibly robust at maintaining state across different checks—if one check fails or times out, the others can't suddenly collapse.
Lalam: Thinking about culture, if we adopt this summarized understanding of content risk, it changes how communities interact online; they become more critically aware participants rather than just passive consumers of information.
Improvements: Tom: Okay, so we've seen the scope and the summary; now we’re looking at what improvements the paper suggests for "Multilingual Models for Check-Worthy Social Media Posts Detection." This section feels like where the real engineering headaches are addressed.
Jane: The main takeaway here, if I'm reading it right, is that they aren't just slapping a model onto existing data; they’re suggesting architectural tweaks to make the whole detection process more resilient and accurate.
Lu: The improvements section really highlights the importance of granularity in error analysis. They look at how performance changes based on input
Paper discussion segment 3: Tom: So, we’ve seen how these models tackle claims across different languages and topics, but what really sets this work apart for us isn't just the results—it’s the engineering improvements they introduced to make it a viable tool.
Jane: That’s right, Tom. The way they built this system is incredibly elegant; instead of having separate models for checking factual claims and then another model for spotting harmful content, they integrated both into a single multi-label architecture.
Lu: That simultaneous approach is the big creative leap here, Jane. By training one set of weights to predict two distinct outcomes—Verifiable Factuality and Harmfulness—they’ are forcing the model to learn complex relationships between the data points that are far more nuanced than if we just used two separate classifiers.
Meng: From an engineering perspective, that combined approach also significantly streamlines the pipeline. We’re talking about one inference step for a single batch of posts instead of two or three, which is a massive gain in efficiency for deployment on standard hardware like a Tesla T4 GPU or even powerful CPU units.
Lalam: The implications for global discourse are huge because of how they handled that multilingual setup; by designing the model to natively process low-resource languages like Arabic and Polish within the same architecture, we’ve essentially given a unified voice to those communities.
Jane: And it’s not just about being capable; it' also about being fast. The authors demonstrated that this complex multi-label task allows them to classify thousands of posts per hour, which is a massive speedup compared to human fact-checkers, as they showed in the inference time analysis.
Tom: It’s wild that you're talking about a machine processing thousands of posts while humans are still reading them sentence by by sentence.
Lu: That speaks directly to the automation potential; we aren't replacing the human expert, but we are giving them an incredibly powerful initial triage tool.
Meng: And it provides a robust baseline for when the model flags something that isn't just a simple factual error, allowing us to route those flagged posts straight to check-worthy experts based on both criteria.
Lalam: It’s about enabling human review at scale, ensuring that these powerful tools are not just speeding up the process but improving the quality of public discourse by identifying verifiable risk in every language.
Jane: Exactly, Lu. We've got a system that is faster, smarter, and handles both factual and harmful claims across multiple languages simultaneously.
Conclusion: Tom: So, we’re wrapping up our discussion on "Multilingual Models for Check-Worthy Social Media Posts Detection," and what's clear is that this isn't just a theoretical exercise; the work has real, practical power.
Jane: It shows us how to build reliable systems that can identify both verifiable claims and harmful content simultaneously, which is a huge step forward in automating the initial phase of fact-checking.
Lu: I think we should be really proud of the XLM-RoBERTa architecture's ability to handle these complex, dual tasks across different languages, too.
Meng: And it’ does that without requiring specialized hardware; the design is practical enough that an AI startup could implement this kind of detection using standard GPU setups.
Lalam: The core message I see is that this framework gives us a standardized way to evaluate truth and harm in diverse cultures, helping to foster a more informed global conversation.
Tom: That's the long-term goal—moving from a fragmented world of misinformation toward a system where verifiable facts are accessible in all our languages.
Jane: It’s truly impressive that the authors managed to achieve these high scores while handling low-resource languages, which is often where AI models struggle the most.
Lu: The ability to see performance metrics for Bulgarian and Polish—languages with huge shares of data in the underlying training corpora—suggest we are seeing a reliable scaling of knowledge here.
Meng: And it's not just high scores, Jane; that also excellent inference time means we can deploy this at scale without needing expensive, massive servers.
Lalam: By making these models efficient and effective, we're building the infrastructure for a more responsible way to engage with information online.
Tom: We hope this work on "Multilingual Models for Check-Worthy Social Media Posts Detection" has a real impact on how people use social media moving forward.
Jane: It provides a much more sophisticated tool than what we had before, making the work of fact-checkers significantly easier and faster.
Lu: We are certainly seeing a very ambitious blend of engineering and linguistic theory in this paper, which is exciting to see.
Meng: I think every company that needs to handle global content needs to take a look at how practical this solution is for their own pipelines.
Lalam: And it's about creating the foundational tools that allow us all to move toward a more informed and harmonious future together.
Kempelen Institute of Intelligent Technologies · Kazimierz Wielki University of Bydgoszcz, Poland
cs.CL
Submitted: 2024-08-13
Updated: 2026-09-04
Code: https://github.com/flairNLP/flair
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 86/100
The gist: The research analyzes the performance of multi-label XLMRoBERTa-base models for detecting verifiable factual claims and harmful claims within social media posts.
Key concepts
- Multi-label Architecture
- Instead of using separate models to check factual claims and harmful content, this approach integrates both checks into a single architecture. This forces the model to learn complex relationships between data points and significantly streamlines the pipeline by requiring only one inference step.
- Check-Worthy Detection
- The system is designed not just to flag errors but to identify posts that warrant human review. It combines criteria for factual verifiability with potential harm, providing a robust baseline for routing flagged content directly to expert fact-checkers.
- Multilingual Processing
- The models are built to natively process low-resource languages, such as Arabic and Polish, within the same architecture. This allows the system to provide a unified voice and handle global discourse without needing separate language-specific models.
Terminology
Summary
The research analyzes the performance of multi-label XLMRoBERTa-base models for detecting verifiable factual claims and harmful claims within social media posts. The core focus of this analysis is to investigate whether sentence length influences model prediction accuracy, a critical consideration given that automated detection systems must be robust across varied textual structures.
Verifiable Factual Claims Detection Analysis
The study conducted an error analysis on 373 samples faultily classified as not verifiable factual claims (false negatives). The analysis revealed significant differences in performance based on sentence length. To quantify this, the group of 373 samples was divided into four subsets representing increasing statistical lengths:
-
0–25% length range (shortest): Minimum length of 26 characters and a maximum length of 95 characters.
-
25–50% length range: Sentences longer than 95 characters and shorter than 145 characters.
-
50–75% length range: Sentences longer than 145 characters and shorter than 223 characters.
-
75–100% length range (longest): Sentences with more than 223 characters.
The results presented in Table 17 confirm a strong correlation between sentence length and detection recall, showing that the highest recall was obtained for the group of the longest sentences.
Specifically, there is a significant discrepancy (6 points) between the shortest and the longest group of sentences,
with recall values observed as follows:
-
0–25%: 0.7867
-
25–50%: 0.7465
-
50–75%: 0.8000
-
75–100%: 0.8452
Impact of Sentence Length on Model Prediction
The analysis of the multi-label XLM-RoBERTa-base model predictions for verifiable factual claims demonstrates that there are significant differences in individual groups of sentences.
The findings lead to the conclusion that the length of the sentence has a significant impact on the model prediction score and longer sentences are more likely to contain verifiable factual claims.
This observation aligns with literature suggesting that longer sentences are more likely to contain truth than shorter sentences (Krickl and Kirrane, 2022).
The researchers hypothesize that this relationship is justified because, by definition, the truth is examined and true sentences are usually longer.
Model Improvement Directions
To enhance the model's performance and make it more sentence length independent,
the paper proposes two primary avenues for future research:
-
Increasing the number of samples in the training set that are relatively short sentences but contain verifiable factual claims.
-
Including
the sentence length as additional (next to the text) feature during pre-training process.
Harmful Claims Detection Analysis
A similar analysis was conducted regarding the detection of harmful claims, utilizing Figure 7. However, in contrast to factual claim detection, no similar phenomenon was detected.
The results show that the discrepancies in predictions for sentences of different lengths are relatively small and no clear trend in this respect was spotted.
Therefore, the study concludes that it cannot be concluded that longer sentences are more likely to contain harmful claims.
Improvements for AI systems
The current system relies heavily on a single feature (sentence length) and a fixed transformer architecture (Multi-label XLMRoBERTa-base). The improvements must focus on robustness, contextual understanding, and integrating the observed correlation patterns scientifically.
Improvement: Implement a comprehensive feature vector that goes beyond simple character/token length. This vector should include:
-
Syntactic Complexity Metrics: Measure dependency parse tree depth, ratio of main verbs to total tokens, and density of named entities (NER).
-
Information Density Score (Novel Metric): Calculate the ratio of unique key terms (using TF-IDF or BERT embedding similarity clustering) to the total token count. High information density is hypothesized to correlate with factual claims.
-
Source/Citation Metadata: If available, integrate structured data about the source (e.g., domain authority score, known bias index, publication date proximity).
Improved Capability: The system will move from a text-only classifier to a Feature-Augmented Contextual Classifier. It can now quantify why a sentence is likely factual—not just that it is long—by analyzing its structural complexity and information payload.
Improvement: Instead of treating Factuality and Harmfulness as two independent classification tasks, implement a single Multi-Task Learning (MTL) framework using a shared encoder backbone (e.g., XLM-RoBERTa). The model should predict:
-
Verifiability Score (Factuality): Is the claim verifiable?
-
Harmfulness Score: Does the claim incite harm?
-
Causality/Mechanism Identification: A structured output that identifies the type of potential falsehood (e.g., Misattribution, Exaggeration, Non-existence Claim).
The training objective should be a weighted combination of losses (L total = alpha L fact + beta L harm + gamma L mechanism).
Improvement: To mitigate the dependence on sentence length (the correlation observed in the 75-100% group), integrate a Contrastive Learning (CL) pre-training step.
-
Process: Pair short, verifiable claims with long, fabricated/misleading claims that share similar surface linguistic features. The model is trained to maximize the distance between the embeddings of these two pairs while minimizing the distance between two truly verifiable factual statements.
-
Implementation: Use a Siamese network structure during fine-tuning, explicitly penalizing length-based feature reliance by forcing the embedding space to rely on semantic content rather than token count.
Improvement: Integrate a dedicated module for Temporal Relationship Extraction (e.g., identifying dates, timelines, and sequences of events). Furthermore, structure the input data to allow for NLI-style comparisons:
-
(Premise): The claim sentence.
-
(Hypothesis/Context): Supporting evidence (e.g., scientific papers, official reports) or counter-evidence.
The model must be trained to classify the relationship between Premise and Hypothesis as Contradictory, Neutral, or Supportive.
Feature Current Limitation Improvement Implemented New Capability Provided
:---:---:---:---
Feature Set Relies on token/char length only. (Single dimension) Incorporates Syntax, Information Density, Source Metadata. (Multi-dimensional) Deep Feature Quantification: Determines why a claim is factual or harmful, not just that it is.
Architecture Two independent classifiers (Factuality & Harmfulness). Multi-Task Learning (MTL) with shared encoder and mechanism prediction. Cross-Task Robustness: Predictions are reinforced by multiple related hypotheses, reducing false positives.
Bias Mitigation Length bias significantly impacts recall score. Contrastive Learning (CL) pre-training step. Length Independence: Consistent high performance regardless of sentence length, addressing the major experimental flaw.
Verification Binary output: Factual/Not Factual, Harmful/Not Harmful. NLI & Temporal Relationship Module (Premise vs. Context). Evidence-Based Verification: Provides quantifiable support strength and identifies specific contradictions with external knowledge bases.
Abstract
This work presents an extensive study of transformer-based NLP models application for detection of social media posts that contain verifiable factual claims and harmful claims. The study covers various activities, including dataset collection, dataset pre-processing, architecture selection, setup of settings, model training (fine-tuning), model testing, and implementation. The study includes a comprehensive analysis of different models, with a special focus on multilingual models where the same model is capable of processing social media posts in both English and in low-resource languages such as Arabic, Bulgarian, Dutch, Polish, Czech, Slovak. The results obtained from the study were validated against state-of-the-art models, and the comparison demonstrated the robustness of the proposed models. The novelty of this work lies in the development of multi-label multilingual classification models that can simultaneously detect harmful posts and posts that contain verifiable factual claims in an efficient way.
Sources
- Two-Stage Classifier for COVID-19 Misinformation Detection Using BERT: a Study on Indonesian Tweets
- g2tmn at Constraint@AAAI2021: Exploiting CT-BERT and Ensembling Learning for COVID-19 Fake News Detection
- Autonomation, Not Automation: Activities and Needs of European Fact-checkers as a Basis for Designing Human-Centered AI Systems
- Identification of COVID-19 related Fake News via Neural Stacking
- Exploring Text-transformers in AAAI 2021 Shared Task: COVID-19 Fake News Detection in English
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Some Observations on Fact-Checking Work with Implications for Computational Support
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
- ClimateBert: A Pretrained Language Model for Climate-Related Text
- XLNet: Generalized Autoregressive Pretraining for Language Understanding
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering