Like Article, Like Audience: Enforcing Multimodal Correlations for Disinformation Detection
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "Like Article, Like Audience: Enforcing Multimodal Correlations for Disinformation Detection".
Jane: The paper was written by Liesbeth Allein, Marie-Francine Moens and Domenico Perrotta from European Commission, Joint Research Centre and KU Leuven.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Title and Authors: Tom: Alright, welcome back to the show, everybody! Today we are diving into a paper that's got a fantastic title: "Like Article, Like Audience: Enforcing Multimodal Correlations for Disinformation Detection." Jane, I gotta say, that title alone got me hooked.
Jane: It really is a great one, Tom. And the authors are Liesbeth Allein, Marie-Francine Moens, and Domenico Perrotta. They're working out of the European Commission's Joint Research Centre and KU Leuven. So we've got a mix of policy-adjacent research and academic rigor here.
Tom: And the title basically sets up the whole premise, right? It's saying that if you want to know if an article is fake, you should look at the people who are sharing it. Like, birds of a feather flock together, but in the news world.
Jane: Exactly. The idea is that a user's online identity, which is built from what they post and what they share, should correlate with the articles they choose to spread. If someone's Twitter feed is full of conspiracy theories, and they share an article, that article might be suspect.
Tom: So it's not just about the article text itself, but about the social context around it. That's a really clever angle. And the authors are using this to build a detection system that doesn't need to know anything about the user at the time of prediction.
Jane: That's the key part, Tom. They only use the user information during training. So the model learns to associate certain article features with certain user profiles, but when it's actually classifying a new article, it only looks at the article text. This is huge for privacy and for practical deployment.
Tom: Oh, that's a big deal. You don't have to go out and collect user data every time you want to check a news story. You just train the model once, and then it's self-sufficient. I love that.
Jane: And it's also a clever way to sidestep the ethical concerns of profiling users. The model isn't making decisions based on who you are; it's just using the patterns it learned during training.
Tom: So, the title is catchy, the authors are credible, and the premise is solid. I'm really curious to see how they actually pull this off. What's the technical machinery behind this?
Jane: Well, we're about to get into exactly that. But first, let's just say this paper is trying to answer a really important question: can we use the wisdom—or lack thereof—of the crowd to spot fake news? And the answer seems to be a promising yes.
Tom: I'm sold. Let's keep going and see how they built this thing.
Summary and Implications: Tom: So, Jane, we've got the title and the authors. Now let's talk about what this paper actually does. The abstract lays out a pretty clear mission.
Jane: It does. The core idea is that they're building a multimodal learning algorithm. That just means they're using two different types of data: the news articles themselves and the user-generated content from Twitter, like tweets and profile descriptions.
Tom: And they're not just throwing them together. They're enforcing correlations. The model is trained to pull the article's representation close to the representations of the users who shared it.
Jane: Right. And also to pull the representations of those users close to each other. The assumption is that people who share the same article probably have something in common, and that commonality is reflected in their language.
Tom: So, it's like the model is learning a sort of "social fingerprint" for each article. If the fingerprint matches a group of users who are known to spread disinformation, the article is more likely to be fake.
Jane: Exactly. And the beauty is that they apply this to three different neural classifiers: a CNN, a Hierarchical Attention Network, and DistilBERT. So it's not a one-off trick; it works across different architectures.
Tom: And the results? They tested it on FakeNewsNet and ReCOVery datasets. The performance gains are real, especially on the smaller datasets like Politifact and ReCOVery.
Jane: That's a good point. The improvement is more noticeable when the data is scarce. That suggests the user information is acting as a kind of regularizer, helping the model generalize better when it doesn't have a ton of article examples.
Tom: So, it's not just about making a better fake news detector. It's about making a more efficient one that can work with less labeled data. That has huge implications for real-world deployment, where you might not have a massive dataset of verified articles.
Jane: And there's another implication that I think is really important. By only using user data during training, they're respecting the European Commission's ethics guidelines for trustworthy AI. They're not building a surveillance tool.
Tom: That's a refreshing take. It's like they're saying, "We can learn from you without having to watch you." I'm really curious to see the technical details of how they enforce these correlations.
Jane: Me too. Let's get into the methodology.
Improvements Suggested: Tom: Alright, Jane, we've covered the big picture. Now let's talk about what this paper is actually improving upon. What's the state of the art, and how does this paper push it forward?
Jane: Great question. A lot of previous work in disinformation detection either ignores users entirely or uses them in a pretty shallow way. Some models just use a user ID number, which doesn't capture anything about who they are.
Tom: Right, like just knowing "user twelve thousand three hundred forty-five" shared an article doesn't tell you much.
Jane: Exactly. Other models use a user's comment on a specific article, but that's a single, short text that's directly related to the article. It's not a representation of the user's overall identity.
Tom: So, this paper is saying, "Let's look at the whole person, not just their reaction to one piece of news."
Jane: Precisely. They represent each user by their profile description and a collection of their recent tweets. That's a much richer signal. And they're not just using that signal to make a prediction; they're using it to shape the latent space of the article encoder.
Tom: That's the key improvement, isn't it? They're not just adding user features as extra input. They're using the user representations to constrain how the article representations are learned.
Jane: Exactly. It's a form of constrained representation learning. The article encoder is forced to produce representations that are close to the user representations in the latent space. This way, the article features become more aligned with the social context in which they're shared.
Tom: And they also mention that this is new in multimodal disinformation detection. Coordinating the article latent space with a loss function that incorporates learned user representations. That's a clever twist.
Jane: It is. And it has a nice side effect. Because the user information is only in the loss function, not in the model input, the final model is simpler and doesn't require user data at inference time.
Tom: So, they're improving on two fronts: they're using richer user data, and they're using it in a smarter way. But I'm wondering, does this actually work in practice? What do the numbers look like?
Jane: Well, the performance gains are there, but they're not uniform. For the CNN and DistilBERT models, the improvements are quite noticeable. For the HAN model, the gains are more modest.
Tom: Interesting. So, the architecture matters. Some models are more amenable to this kind of constraint than others.
Jane: It seems that way. And they also found that the choice of user representation matters. Sometimes using just the profile description works best, sometimes using the tweets works better, and sometimes combining both is the way to go.
Tom: So, it's not a one-size-fits-all solution. You have to tune it for your specific model and dataset. That's a practical consideration.
Jane: Definitely. But the overall trend is positive. The paper shows that this approach can consistently improve performance across different settings.
First Page Deep Dive: Tom: So, Jane, we've talked about the abstract and the improvements. Now let's actually look at the first page of the paper itself. There's a lot of interesting stuff packed into that opening.
Jane: There is. The first thing that jumps out is the figure. It's a great overview of the whole algorithm. You can see the article going into an encoder, and a subset of users going into a separate encoder.
Tom: And then there are three training objectives shown in red. Objective A is the standard classification loss, to tell true from fake. Objective B is minimizing the distance between the article and the users who shared it. And Objective C is minimizing the distance between the users themselves.
Jane: Right. And that figure really clarifies the architecture. It's not a complex, tangled network. It's a clean separation between the article path and the user path, with the loss function tying them together.
Tom: The page also has the abstract, which we've already covered, but there's also a section on the ACM Reference Format. It's a standard academic paper, but the content is anything but standard.
Jane: And I love the keywords they chose: "Disinformation, Fake News, Constrained Representation Learning, Multimodal Data Fusion, Natural Language Processing." That really sums up the contributions.
Tom: It does. And the introduction section starts to lay out the motivations. They talk about how authors and users have both audience-oriented and self-oriented motivations for spreading information.
Jane: That's a really insightful framing. Authors want to persuade or deceive their audience, and users want to construct their online identity. Both of those motivations are reflected in the content they create and share.
Tom: So, the paper is grounded in social science theory, not just computer science. That gives it a lot more depth.
Jane: Absolutely. They're not just throwing a bunch of data at a neural network. They're building a model based on an understanding of human behavior.
Tom: And that understanding is what drives the design of the loss functions. The correlations they're enforcing are based on real-world observations about how people interact with information.
Jane: Exactly. It's a nice example of how interdisciplinary research can lead to better AI systems. The first page really sets the stage for a thoughtful and well-motivated approach.
Conclusion: Tom: Well, Jane, we've reached the end of our discussion on "Like Article, Like Audience: Enforcing Multimodal Correlations for Disinformation Detection." What a ride.
Jane: It really has been. We started with the catchy title, and we've ended with a deep appreciation for the clever methodology. The paper shows that you can leverage user information to improve fake news detection without compromising user privacy.
Tom: And the key takeaway for me is that the model learns to associate article content with the social context of its audience. That's a powerful idea that goes beyond just looking at the text.
Jane: And the fact that they tested it on multiple architectures and datasets makes the findings more robust. It's not a fluke; it's a general principle.
Tom: The implications for real-world applications are huge. Imagine a browser extension or a social media platform that can flag potentially fake news articles with just the text, no user tracking needed.
Jane: That would be a game-changer. And it aligns with the ethical guidelines for trustworthy AI, which is becoming increasingly important.
Tom: So, as we say goodbye to this paper, I think we can both agree that it's a significant step forward in the fight against disinformation.
Jane: Absolutely. It's a smart, ethical, and effective approach. We'll be excited to see what these authors do next.
Tom: And with that, we'll wrap up our discussion. Thanks for listening, and we'll see you for the next paper. Goodbye, "Like Article, Like Audience"!
Liesbeth Allein, Marie-Francine Moens, Domenico Perrotta
European Commission, Joint Research Centre · KU Leuven
cs.CL, cs.IR
Submitted: 2021-08-31
Journal ref: Allein, Liesbeth, Marie-Francine Moens, and Domenico Perrotta. "Preventing profiling for ethical fake news detection." Information Processing & Management 60.2 (2023): 103206
DOI: 10.1016/j.ipm.2022.103206
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 60/100
The gist: This paper investigates whether correlations between user-generated and user-shared content can be leveraged for detecting disinformation in online news articles.
Key concepts
- Multimodal Learning
- This technique uses two distinct types of data: the news articles themselves and the user-generated content from social media. The model integrates these diverse inputs—the text and the users' profiles—to achieve a more comprehensive understanding of how information is being shared.
- Constrained Representation Learning
- This is a core methodology where user representations are used to influence or constrain how the article's representation is learned. The system forces the article's features to become closely aligned with the social context of users who share it, creating a 'social fingerprint.'
- Disinformation Detection
- The paper aims to spot fake news by leveraging crowd behavior. It checks if an article matches a group of users known for spreading disinformation. This approach uses the social context around the content to increase the likelihood of accurate detection.
Terminology
Summary
This paper investigates whether correlations between user-generated and user-shared content can be leveraged for detecting disinformation in online news articles. The authors develop a multimodal learning algorithm for disinformation detection, where "the latent representations of news articles and user-generated content allow that during training the model is guided by the profile of users who prefer content similar to the news article that is evaluated, and this effect is reinforced if that content is shared among different users. A key design choice is that
by only leveraging user information during model optimization, the model does not rely on user profiling when predicting an article’s veracity."
The algorithm is applied to three widely used neural classifiers: CNN, HAN, and DistilBERT. The model is optimized on three training objectives: "(1) discriminate between true and false news articles, (2) minimize the distance between article and user latent representations, and (3) minimize the distance between latent representations of users who share the same article." These are combined into a single loss function as a weighted sum: L = λ1 L pred + λ2 L dist(a,U) + λ3 L dist(U), where L pred is the cross-entropy loss, L dist(a,U) is the mean cosine distance between the article representation and all user representations in the subset, and L dist(U) is the mean cosine distance between all user representations within the subset.
The experiments use three datasets: FakeNewsNet (split into Politifact and Gossipcop) and ReCOVery, combined into one dataset and split into train (80%), validation (10%), and test (10%) sets in a label-stratified manner. Four experimental setups are investigated: base (no user information), +u/d (users represented by profile description), +u/t (users represented by tweets), and +u/d+t (users represented by both description and tweets). Results show that leveraging user information in a multimodal learning algorithm positively influences the model performance for all three neural classifiers and datasets.
For the fake class, the +u/d setup has the highest positive impact on Politifact articles (e.g., CNN: +3.78/+9.09/+9.19% P/R/F1), while for GossipCop and ReCOVery, models benefit most when users are partly represented by tweets (e.g., DistilBERT achieves highest GossipCop results with +u/d+t: -1.21/+4.68/+2.87% P/R/F1; CNN achieves highest ReCOVery results with +u/t: +3.19/+3.01/+3.12% P/R/F1). For the true class, similar preferences are observed.
The paper addresses several research questions. First, regarding whether leveraging user information leads to better latent article representations, statistical visualization techniques (multidimensional scaling and robust PCA) show that implementing our multimodal learning algorithm increases the separation between the two prediction classes in the article latent space,
with overlap measures quantifying the separation. Second, regarding user and tweet selection, experiments with late dissemination users reject the hypothesis that early dissemination users lead to higher performance, while using oldest tweets instead of newest tweets improves performance for ReCOVery articles, suggesting leveraging user information from the same time period as the article that is evaluated improves model performance.
Third, regarding whether the model finds real correlations, experiments with randomly assigned user subsets show that "the user-constrained models benefit from leveraging extra user data by having distance constraints between and within the two modalities enforced during model optimization without actually uncovering the assumed real-world correlations constructed by the user modality." A qualitative analysis shows that similar topics in news articles and profile descriptions can guide the model towards correct predictions, but tweets introduce noise that can eclipse valuable information.
The paper concludes that the algorithm provides an elegant way to integrate and correlate multimodal information without requiring its presence at testing time,
and suggests future work could experiment with other modalities such as images and audio. Limitations include the tendency of fake news articles and creators to disappear from the Internet, leading to reduced shares of fake news in datasets, and the fact that a user's identity reflected in their current description and latest tweets might no longer align with their identity at the time of sharing the article.
Improvements for AI systems
Based on the paper, here are the specific improvements I can implement in AI systems:
Improvement: Add two auxiliary loss functions to the training objective of any text classifier (CNN, HAN, BERT-based):
-
Article-User Distance Loss: Minimize cosine distance between the article's latent representation and the latent representations of users who shared it.
-
User-User Distance Loss: Minimize cosine distance between latent representations of users who shared the same article.
Implementation: Modify the loss function to L = λ1·L pred + λ2·L dist(article,user) + λ3·L dist(user,user), with optimal weights found empirically (e.g., [0.8, 0.1, 0.1] for CNN, [0.5, 0.25, 0.25] for HAN, [0.33, 0.33, 0.33] for DistilBERT).
What the improved system can do: Detect fake news with higher F1-scores (up to +9.19% for Politifact fake class with CNN), while only requiring article text at inference time—user data is used only during training.
Abstract
User-generated content (e.g., tweets and profile descriptions) and shared content between users (e.g., news articles) reflect a user's online identity. This paper investigates whether correlations between user-generated and user-shared content can be leveraged for detecting disinformation in online news articles. We develop a multimodal learning algorithm for disinformation detection. The latent representations of news articles and user-generated content allow that during training the model is guided by the profile of users who prefer content similar to the news article that is evaluated, and this effect is reinforced if that content is shared among different users. By only leveraging user information during model optimization, the model does not rely on user profiling when predicting an article's veracity. The algorithm is successfully applied to three widely used neural classifiers, and results are obtained on different datasets. Visualization techniques show that the proposed model learns feature representations of unseen news articles that better discriminate between fake and real news texts.
Sources
- Graph-based Modeling of Online Communities for Fake News Detection
- DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering