GREAT: Generalizable Backdoor Attacks in RLHF via Emotion-Aware Trigger Synthesis
cs.CR, cs.LG
Submitted: 2025-10-10
Updated: 2026-08-28
Terminology
Sources
- A Closer Look at Memorization in Deep Networks
- Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
- Best-of-Venom: Attacking RLHF by Injecting Poisoned Preference Data
- Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering
- GoEmotions: A Dataset of Fine-Grained Emotions
- Does Learning Require Memorization? A Short Tale about a Long Tail
- The Llama 3 Herd of Models
- Here's a Free Lunch: Sanitizing Backdoored Models with Model Merge
- A Survey on LLM-as-a-Judge
- Simulate and Eliminate: Revoke Backdoors for Generative Large Language Models
- Chain-of-Scrutiny: Detecting Backdoor Attacks for Large Language Models
- Fragile Giants: Understanding the Susceptibility of Models to Subpopulation Attacks
- CleanGen: Mitigating Backdoor Attacks for Generation Tasks in Large Language Models
- Excess Capacity and Backdoor Poisoning
- Prometheus: Inducing Fine-grained Evaluation Capability in Language Models
- Is poisoning a real threat to LLM alignment? Maybe more so than you think
- Weight Poisoning Attacks on Pre-trained Models
- Estimating Training Data Influence by Tracing Gradient Descent
- Improving Diversity of Demographic Representation in Large Language Models via Collective-Critiques and Self-Voting
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension
Related papers
- SoK: AI-Augmented Binary Reversing
- Relaxed Sender Anonymity for CBDC Interbank Settlement: A Zero-Knowledge Approach on Permissioned EVM
- Calibration-Family Overfit: Why Trusted Sabotage Monitors Don't Transfer Across Lineages
- Efficient Fuzzy PSI under One-Sided Assumptions
- Sealing the Audit-Runtime Gap for LLM Skills
- Token Composition: A Graph Based on EVM Logs