From Specialization to Generalization: Instruction-tuned LLMs for Robust Harmful Content Mitigation
cs.CL
Submitted: 2026-08-26
Updated: 2026-08-26
Code: https://github.com/TUM-NLP/HIPPO
Project page: https://hasocfire.github.io/hasoc/2019/call_for_participation.html
Terminology
Sources
- Phi-4 Technical Report
- Automated Detection of Cyberbullying Against Women and Immigrants and Cross-domain Adaptability
- No Language Left Behind: Scaling Human-Centered Machine Translation
- Hate Speech Dataset from a White Supremacy Forum
- The Llama 3 Herd of Models
- Right-wing German Hate Speech on Twitter: Analysis and Automatic Detection
- Investigating Data Contamination for Pre-training Language Models
- Measuring a hate speech spectrum with faceted Rasch item response theory and perspective-aware, explainable-by-design deep learning
- Detecting Online Hate Speech Using Context Aware Models
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Multilingual and Multi-Aspect Hate Speech Analysis
- Toxicity Detection: Does Context Really Matter?
- A Benchmark Dataset for Learning to Intervene in Online Hate Speech
- Qwen3 Technical Report
- HateTinyLLM : Hate Speech Detection Using Tiny Large Language Models
- Predicting the Type and Target of Offensive Posts in Social Media
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering