ThreatCore: A Benchmark for Explicit and Implicit Threat Detection
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "ThreatCore: A Benchmark for Explicit and Implicit Threat Detection".
Jane: The paper was written by Davide Bruni, Carlo Bardazzi and Maurizio Tesconi from University of Pisa, Computer Science Department and Institute of Informatics and Telematics, National Research Council, Italy Institute of Informatics and Telematics, National Research Council, Italy.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Summary of Findings: Tom: So, what did they actually find when they tested existing systems like Perspective API and various large language models against this new benchmark? The findings are quite stark and really highlight where current AI falls short.
Jane: It turns out that systems like Perspective API are extremely good at catching obvious, explicit threats, but the recall on implicit threats is nearly zero, which is a massive problem for detecting subtle harm.
Lu: This suggests that these models are fundamentally unable to grasp what's happening beneath the surface-level tokens; they just aren't sophisticated enough to handle indirect or implied threatening language.
Meng: The data shows a clear bias too, where both zero-shot classifiers and models tend to overpredict threat labels, meaning they frequently flag nonthreatening content as dangerous.
Lalam: That tendency is deeply concerning because it creates a false sense of security; if we can't reliably tell what is actually threatening versus just aggressive, our safety systems fail.
Improvements through Methodology: Tom: To build this robust system, the researchers had to re-annotate and augment existing data, creating a much cleaner resource. The methodology behind ThreatCore: A Benchmark for Explicit and Implicit Threat Detection is what we need to look at next.
Jane: They took publicly available datasets and systematically re-annotated them under a unified definition of threat, which resolves the inconsistencies that were present in previous studies.
Meng: This systematic approach, combined with the introduction of Semantic Role Labeling, allows us to structure the text into clear roles—who is doing what to whom—which is essential for building a robust detector.
Lu: It turns text into a structural map of intent, and by using SRL we are giving the AI an internal reasoning engine that can identify exactly what is causing harm within the sentence structure itself.
Lalam: By leveraging this method, we are designing systems that move beyond merely classifying language to understanding causality, which is a significant step forward for safety.
Conclusion and Final Thoughts: Tom: We’ve seen how challenging threat detection really is, but the evidence from ThreatCore: A Benchmark for Explicit and Implicit Threat Detection suggests there’s a path toward much higher accuracy.
Jane: It's such an important distinction that we can't just treat every angry post as equally threatening; the data clearly shows how much of that content is just aggressive or non-threatening at all.
Lu: I think the biggest intellectual contribution here is forcing researchers to look at intent, not just surface features, which really pushes the boundaries of what we can achieve with current AI models.
Meng: This framework provides a tangible blueprint for how a commercial threat detection system could be designed, moving beyond guesswork and into structured logic based on real data distribution.
Lalam: And I feel that by providing these nuanced examples, we are actively contributing to making digital platforms safer for everyone using them through understanding subtlety.
Conclusion: Tom: So, as we wrap up our deep dive into this complex area of digital safety, the core message remains: tackling implicit threats requires fundamentally smarter AI tools than what currently exist.
Jane: It's been fascinating to see how much work goes into creating a standard that demands precision in a field that has historically been very messy and subjective.
Lu: For me, the biggest takeaway is that this benchmark forces us to treat linguistic structure as a primary variable, which is a massive conceptual leap for AI research.
Meng: It's the groundwork we needed to understand the real difficulty of this problem before any commercial solutions can even be attempted without reliable data foundations.
Lalam: Ultimately, by defining this standard through ThreatCore: A Benchmark for Explicit and Implicit Threat Detection, we are actively contributing to making digital platforms safer for everyone using them.
Tom: Absolutely, Lalam. It provides a common language for discussing risk, moving us away from vague concepts like "toxicity" toward measurable, actionable intent.
Jane: It’s a truly vital resource that sets a high bar for the entire industry as we prepare to move on to our next topic.
University of Pisa, Computer Science Department · Institute of Informatics and Telematics, National Research Council, Italy Institute of Informatics and Telematics, National Research Council, Italy
cs.CL, cs.AI
Submitted: 2026-05-11
Updated: 2026-05-11
Journal ref: Proceedings of the 15th International Conference on Data Science, Technology and Applications (DATA 2026), Volume 2, pp. 1127-1135, SciTePress, 2026
Code: https://github.com/DavideBruni/ThreatCore
License: http://creativecommons.org/licenses/by-nc-nd/4.0/
Importance score: 92/100
The gist: The paper introduces ThreatCore, a novel and comprehensive benchmark dataset designed to advance the field of threat detection by rigorously distinguishing between explicit threats, implicit threats,
Key concepts
- Implicit vs. Explicit Threats
- Explicit threats are obvious and easily caught by current systems like Perspective API. However, the benchmark shows that these systems have nearly zero recall for implicit or subtle harm, suggesting they are not sophisticated enough to handle indirect language.
- Overprediction Bias
- The data reveals a bias where both zero-shot classifiers and large language models tend to overpredict threat labels. This means they frequently flag nonthreatening content as dangerous, creating a false sense of security for safety systems.
- Semantic Role Labeling (SRL)
- This methodology structures text into clear roles—who is doing what to whom. By using SRL, the AI can identify exactly what is causing harm within the sentence structure itself, providing an internal reasoning engine for safety.
Terminology
Summary
The paper introduces ThreatCore, a novel and comprehensive benchmark dataset designed to advance the field of threat detection by rigorously distinguishing between explicit threats, implicit threats, and nonthreat content. The development of such a specialized resource is critically important because existing general-purpose moderation systems often fail to capture the nuanced complexity of threatening language, particularly when intent is veiled or expressed in unfamiliar languages. ThreatCore aims to provide a robust foundation for developing more sophisticated and reliable systems capable of identifying the full spectrum of violent intent in online communication.
Dataset Architecture and Scope
ThreatCore was constructed using a unified annotation framework, significantly expanding upon existing research materials. The dataset’s sheer scale is a major contribution, comprising a total of 21,764 instances. This corpus is built upon:
-
15,691 instances sourced from various existing datasets. These were re-annotated to ensure consistency and comparability across different types of harmful content.
-
An additional approximately 6,000 synthetic examples. The inclusion of synthetic data helps bolster the dataset’s coverage and robustness against diverse linguistic patterns.
Evaluation of Current Detection Limitations
The experimental results presented in the work highlight significant limitations within current state-of-the-art approaches for threat detection. General-purpose moderation systems and zero-shot classifiers were shown to struggle acutely with implicit threats, often exhibiting strong biases. These biases manifest as either over-predicting or under-detecting harmful content,
indicating a fundamental inability to grasp subtle or indirect forms of violent intent.
In contrast, the evaluation demonstrated that large language models (LLMs) possess a more balanced and nuanced understanding of threatening language, even when deployed in zero-shot settings. Furthermore, the paper emphasizes that multilingual datasets are crucial for comprehensive evaluation, enabling the investigation of cross-lingual generalization capabilities
and allowing LLMs to provide natural language explanations tailored to end users,
thereby improving transparency in unfamiliar languages.
Enhancing Semantic Understanding for Intent Detection
To improve performance beyond standard classification methods, the authors propose incorporating structured semantic information into the detection pipeline. They show that utilizing Semantic Role Labeling (SRL) can significantly enhance model capabilities by making interaction patterns explicit. This structural understanding enables models to perform deeper reasoning about the underlying intent of a statement, moving beyond mere keyword matching or surface-level toxicity scoring.
Future Directions and System Robustness
The ultimate goal of ThreatCore is to support future research aimed at developing more robust systems capable of identifying both explicit and implicit manifestations of violent intent.
The framework also points toward specific avenues for enhancement:
-
Multilingual Support: The need for multilingual datasets is stressed to allow for a more comprehensive evaluation across diverse linguistic contexts.
-
Explainability: LLMs can enhance usability by providing detailed explanations, which is particularly valuable in real-world applications where transparency and user trust are paramount.
-
Structured Reasoning: Continued focus on integrating structured semantic information, such as that provided by SRL, will enable models to better interpret the causal relationships and underlying motives within threatening communication.
The comprehensive nature of the work—including the dataset, detailed annotation guidelines, and specific prompts—is made publicly available to encourage broad collaboration and accelerate advancements in this critical field.
Improvements for AI systems
1. Development of a Multi-Lingual Cross-Cultural Threat Detection Architecture (MCLTDA)
-
Improvement: Implement an architecture that moves beyond simple translation or parallel data structures. The system must incorporate modules for Linguistic Variation Modeling (LVM) and Cultural Context Embeddings (CCE). LVM must identify structural, idiomatic, and register-specific deviations in language use (e.g., slang, code-switching) that signal threat intent but are not captured by standard tokenization or cross-lingual embeddings like mBERT/XLM-R alone. CCE modules will map detected threat patterns against a database of known cultural communication norms and historical conflict lexicons (e.g., specific regional slurs, historical political metaphors).
-
System Capability: The resulting system can detect subtle threats across dozens of languages and cultural contexts, identifying why content is threatening based on its interaction with localized social norms. It moves from
Is this threatening?
toHow is this language functioning within a specific socio-cultural context to generate threat?
This dramatically reduces false negatives in non-Western/low-resource language settings.
2. Intent and Agency Reasoning Layer using Structured Semantic Graphs (ISR-SSG)
-
Improvement: Integrate a dedicated reasoning layer that utilizes sophisticated Semantic Role Labeling (SRL) coupled with Causal Inference Modeling (CIM). Instead of merely classifying tokens, the system must parse the statement into a structured graph: Agent to Action to Target to [Condition]. The CIM component then analyzes the causal relationship between these nodes to determine underlying intent (e.g., is the stated threat an immediate action, a conditional consequence, or an expression of generalized malice?). This layer must specifically model pragmatic implicature—the implied meaning that goes beyond literal syntax.
-
System Capability: The system can accurately differentiate between: 1) A hypothetical discussion about violence (low intent); 2) A statement expressing future capability (medium intent); and 3) A concrete, actionable plan with defined timing and target (high, immediate threat). This moves the AI from simple classification to sophisticated risk assessment, providing actionable threat vectors rather than just binary labels.
3. Zero-Shot Explainability and Trust Module (ZSETM)
- Improvement: Develop a post-processing module that generates natural language explanations for its own predictions, tailored to the end user's technical expertise (e.g., simple explanation for a general moderator vs. detailed feature attribution for an intelligence analyst). This requires implementing Attention Visualization Guided Explanation Generation. When a threat is detected, the system must not only provide the score but also highlight:
-
The specific semantic roles that contributed most to the threat score (e.g., "The high risk was driven by the combination of Target and Action, specifically Target= Minority Group and Action= Violence ").
-
A counterfactual example: What minimal change in the text would drop the threat score below a certain threshold, demonstrating why the original text was problematic.
- System Capability: This module drastically improves transparency and trust. For end-users, it provides immediate, verifiable justification for content moderation actions (
We flagged this because...
). For researchers, it allows for rapid debugging of model biases and blind spots by visualizing the exact reasoning path taken by the AI.
4. Adaptive Dataset Augmentation Pipeline (ADAP)
-
Improvement: Instead of relying solely on static synthetic data generation (like simple paraphrasing), implement a continuous learning pipeline that uses Adversarial Data Generation. The system should maintain a
Hard Negative
dataset—examples that are highly ambiguous or benign but structurally similar to threats. Periodically, an adversarial generator (e.g., a fine-tuned GAN or LLM) must attempt to create new examples designed specifically to fool the current threat detection model's weakest points (e.g., using indirect references, sarcasm, or complex metaphors). The system is then retrained on these self-generated, high-difficulty samples. -
System Capability: This ensures the AI system remains perpetually robust against
concept drift
and adversarial attacks. It guarantees that the model does not become brittle and maintains high performance even when malicious actors adapt their language to bypass known detection methods (a critical requirement in real-world threat intelligence).
Abstract
Threat detection in Natural Language Processing lacks consistent definitions and standardized benchmarks, and is often conflated with broader phenomena such as toxicity, hate speech, or offensive language. In this work, we introduce ThreatCore, a public available benchmark dataset for fine-grained threat detection that distinguishes between explicit threats, implicit threats, and non-threats. The dataset is constructed by aggregating multiple publicly available resources and systematically re-annotating them under a unified operational definition of threat, revealing substantial inconsistencies across existing labels. To improve the coverage of underrepresented cases, particularly implicit threats, we further augment the dataset with synthetic examples, which are manually validated using the same annotation protocol adopted for the re-annotation of the public datasets, ensuring consistency across all data sources. We evaluate Perspective API, zero-shot classifiers, and recent language models on ThreatCore, showing that implicit threats remain substantially harder to detect than explicit ones. Our results also indicate that incorporating Semantic Role Labeling as an intermediate representation can improve performance by making the structure of harmful intent more explicit. Overall, ThreatCore provides a more consistent benchmark for studying fine-grained threat detection and highlights the challenges that current models still face in identifying indirect expressions of harmful intent.
Sources
- Mapping the Italian Telegram Ecosystem: Communities, Toxicity, and Hate Speech
- AMAQA: A Metadata-based QA Dataset for RAG Systems
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering