Refusal is Not Safety! Benchmarking Latent Safety Risks of LLM-Driven Content Humorization
Yu Cui, Ruiqing Yue, Tingyu Li, Sicheng Pan, Zhuoyu Sun, Xufeng Zhang, Baohan Huang, Haibin Zhang, Cong Zuo
cs.CR
Submitted: 2026-07-17
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- gpt-oss-120b & gpt-oss-20b Model Card
- Agent Data Injection Attacks are Realistic Threats to AI Agents
- Chumor 1.0: A Truly Funny and Challenging Chinese Humor Understanding Dataset from Ruo Zhi Ba
- "Humans welcome to observe": A First Look at the Agent Social Network Moltbook
- A Safety Report on GPT-5.2, Gemini 3 Pro, Qwen3-VL, Grok 4.1 Fast, Nano Banana Pro, and Seedream 4.5
- Harm or Humor: A Multimodal, Multilingual Benchmark for Overt and Covert Harmful Humor
- TRACE: Task-Aware Adaptive Self-Evolving Agentic Jailbreaking
- Universal and Transferable Adversarial Attacks on Aligned Language Models
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs