ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety

summary

Video file (mp4)

The gist

" * Abstract and Introduction Conspiracy theories are narratives that attribute significant events to a covert, powerful group operating with malicious intent, serving as counter-narratives to

In short

The episode explores the ConspirED dataset, a collection of snippets from online conspiracy articles categorized by specific cognitive traits. The research shows that while large language models can detect these complex signatures, they often struggle with safety, tending to reproduce the original reasoning patterns when asked to generate fact-checked counter-narratives.

Key concepts

ConspirED
ConspirED is a massive dataset composed of snippets between eighty and one hundred twenty words long. These excerpts are pulled from online conspiracy articles and are grouped by the dominant cognitive trait expressed in each snippet, rather than just by topic.
Cognitive Traits
These are specific labels, such as 'Overriding suspicion' or 'Nefarious intent,' that researchers apply to the text. They represent a consistent way of thinking or a mental lens through which conspiracy theories are interpreted, highlighting how information is processed.
LLM Safety/Vulnerability
This refers to the challenge where large language models struggle with these specific arguments. The study found that LLMs are susceptible to the original framing, often reproducing the input’s own reasoning patterns even when they are prompted to provide factual counter-narratives.

Terminology used across episodes

This episode discusses

The paper

ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety · Read on arXiv

Luke Bates, Max Glockner, Preslav Nakov, Iryna Gurevych

Technical University of Darmstadt, Department of Computer Science and Hessian Center for AI (hessian.AI) · National Research Center for Applied Cybersecurity ATHENE, Germany · Mohamed bin Zayed University of Artificial Intelligence · Ubiquitous Knowledge Processing Lab (UKP Lab)

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "ConspirED: A Dataset for Cognitive Traits of Conspiracy Theories and Large Language Model Safety".

Jane: The paper was written by Luke Bates, Max Glockner, Preslav Nakov and Iryna Gurevych from Technical University of Darmstadt, Department of Computer Science and Hessian Center for AI (hessian.AI) and National Research Center for Applied Cybersecurity ATHENE, Germany and Mohamed bin Zayed University of Artificial Intelligence and Ubiquitous Knowledge Processing Lab (UKP Lab).

Tom: Stay tuned as we take you through the paper and discuss its implications.

Summary: Tom: So, Jane, can you explain what this dataset called C ONSPIR ED actually consists of for our listeners?

Jane: It’s a massive collection of excerpts—snippets between eighty and one hundred twenty words long—pulled from online conspiracy articles.

Meng: And the authors didn't just gather random text; they applied the C ONSPIR cognitive framework, which gives us specific labels for things like "Overriding suspicion."

Lu: This is so much more than simple topic modeling; we’re identifying a mindset—a consistent way of thinking—that applies across all these narratives.

Jane: The data comes from two main sources: the LOCO corpus and GlobalResearch, which provided the material needed to capture this range of conspiratorial thinking.

Tom: It’s interesting because, rather than just grouping articles by topic like "plandemic," they are grouped by the dominant cognitive trait expressed in each snippet.

Meng: That distinction is vital for targeting interventions; you're not just trying to correct a fact, you're trying to disrupt the specific logic used to promote the theory.

Lalam: It’s a way of saying that while misinformation is often about facts, C ONSPIR ED is about how those facts are interpreted through a specific mental lens.

Improvements and Findings: Tom: Now, we know this dataset exists, but what did the authors actually *do* with it? They moved into experiments to test detection and safety.

Jane: They tested how well different AI models could recognize these C ONSPIR traits, which is a big challenge because they are multifaceted.

Lu: The results show that while LLMs can detect these specific cognitive signatures, they are not perfect at identifying *all the applicable* traits in the snippet at once.

Meng: That’s where the lightweight LaGoNN classifier comes in, which is a much faster way to perform this trait detection without needing massive computational power.

Tom: But even more concerning than how well models detect it is how they handle it when they were prompted to rewrite the content journalistically.

Jane: That's the core paradox, Tom; the LLMs are capable of identifying these traits, but they also have a tendency to become "misaligned" by them.

Lalam: They essentially reproduce the input’s own reasoning patterns in their output, even when they are asked to provide fact-checked counter-narratives.

Meng: It suggests that for these specific types of conspiratorial arguments, the models are more susceptible to the original framing than they are to verifiable factual data.

Conclusion: Tom: So, after all this research, what’s the final summary of what CONSPIR ED tells us?

Jane: It confirms that conspiracy narratives possess a specific cognitive architecture—traits like "Nefarious intent" and "Overriding suspicion" are highly prevalent in the data.

Lu: The study shows that this isn't just a problem for AI; it highlights a general vulnerability in how we train large language models to handle complex reasoning.

Meng: It’s a real-world safety risk, proving that even state-of-the-art models struggle with adversarial inputs that are more subtle than simple lies.

Lalam: The entire finding suggests that the way we think about these theories—not what they say, how they say it—is the key to understanding how AI needs to be re-evaluated.

Tom: The C ONSPIR ED framework gives us a powerful, scalable tool for spotting and fighting specific rhetorical patterns.

Jane: It’s a lot of work that has gone into this dataset and the subsequent testing, but it’ truly provides a vital resource for developing better detection systems.

Final Thoughts: Tom: We've seen how the C ONSPIR ED dataset was built and what its findings are regarding LLM safety.

Jane: The complexity of these cognitive traits is something we really need to keep in mind as we continue to develop new AI tools that can be used for content moderation.

Lu: We’re moving away from just simple fact-checking toward understanding the entire mindset, which is a massive intellectual shift for the field.

Meng: It raises an immediate engineering challenge, though; how do we build systems that can reliably distinguish between these complex cognitive traits and a simpler lie?

Lalam: The impact of this research will be felt in how AI learns to understand human bias, allowing us to improve not just the tech, but our cultural ability to process information.

Tom: It’s clear that the C ONSPIR ED project is laying groundwork for a much more nuanced approach to tackling misinformation.

Jane: It's definitely something worth keeping an eye on as we transition into discussing the next paper on arXiv.

More episodes

← Home