Counter with Evidence! A Multi-Agent Memory Efficient Reasoning Framework for Hate Category Informed Counterspeech Generation
cs.CL
Submitted: 2026-08-24
Updated: 2026-08-30
Comments: Accepted at EMNLP 2026 Main Conference
Code: https://github.com/C0mRD/Counter_with_evidence
License: http://creativecommons.org/licenses/by/4.0/
The gist: Counterspeech effectively neutralizes the impact of online hate.
Terminology
Abstract
Counterspeech effectively neutralizes the impact of online hate. Although prior work explores automated counterspeech generation, it largely emphasizes stylistic control while treating hate speech as homogeneous, overlooking that distinct forms of abuse require fundamentally different counterspeech strategies. To address this gap, we introduce FIRE (Factuality Informed Multi-Agent Reasoning Framework) that first decomposes hate speech into one of the five distinct categories (misinformation, stereotype, conspiracy, dehumanizing, non-factual), and then maps it to a targeted counterspeech style. To facilitate FIRE, we curate FactualCS, a novel dataset of 4,784 instances that provides the annotations regarding hate categories, reasoning traces, and evidence mappings, which are critical elements for grounded generation that are missing in prior work. A comprehensive evaluation across 28 baseline configurations demonstrates that FIRE significantly surpasses existing methods, despite using compact agents (< 2B). FIRE achieves a about 12 % and about 11 % improvements in factual and category-specific accuracy respectively, while simultaneously reducing toxicity by about 11 % relative to the strongest baselines. Further human evaluation confirms that responses generated by FIRE are significantly preferred over the strongest baselines, underscoring its effectiveness for real-world deployment. These findings show that decomposing the underlying intent of hate speech is essential for generating safe, effective, and contextually precise counterspeech.
Sources
- Automated Hate Speech Detection and the Problem of Offensive Language
- CrowdCounter: A benchmark type-specific multi-target counterspeech dataset
- Static Sandboxes Are Inadequate: Modeling Societal Complexity Requires Open-Ended Co-Evolution in LLM-Based Multi-Agent Simulations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks
- Multi-Agent Retrieval-Augmented Framework for Evidence-Based Counterspeech Against Health Misinformation
- Qwen3 Technical Report
- Mistral 7B
- Qwen2.5 Technical Report
- LoRA: Low-Rank Adaptation of Large Language Models
- BERTScore: Evaluating Text Generation with BERT
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- A Multi-Aspect Framework for Counter Narrative Evaluation using Large Language Models
- Gemini: A Family of Highly Capable Multimodal Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering