PersonaTeaming: Supporting Persona-Driven Red-Teaming for Generative AI
cs.HC, cs.AI, cs.CY
Submitted: 2026-05-07
Updated: 2026-09-08
Comments: Accepted to The ACM Symposium on User Interface Software and Technology (UIST) 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: Recent developments in AI safety research have called for red-teaming methods that effectively surface potential risks posed by generative AI models, with growing emphasis on how red-teamers'
Terminology
Abstract
Recent developments in AI safety research have called for red-teaming methods that effectively surface potential risks posed by generative AI models, with growing emphasis on how red-teamers' backgrounds and perspectives shape their strategies and the risks they uncover. While automated red-teaming approaches promise to complement human red-teaming through larger-scale exploration, existing automated approaches do not account for human identities and rarely incorporate human inputs. In this work, we explore persona-driven red-teaming to advance both automated red-teaming and human-AI collaboration. We first develop PersonaTeaming Workflow, which incorporates personas into the adversarial prompt generation process to explore a wider spectrum of adversarial strategies. Compared to RainbowPlus, a state-of-the-art automated red-teaming method, PersonaTeaming Workflow achieves higher attack success rates while maintaining prompt diversity. However, since automated personas only approximate real human perspectives, we further instantiate PersonaTeaming Workflow as PersonaTeaming Playground, a user-facing interface that enables red-teamers to author their own personas and collaborate with AI to mutate and refine prompts. In a user study with 11 industry practitioners, we found that PersonaTeaming Playground enabled diverse red-teaming strategies and outputs that practitioners perceived as useful, and that AI-generated suggestions in the PersonaTeaming Playground encouraged out-of-the-box thinking even when practitioners did not follow them strictly. Together, our work advances both automated and human-in-the-loop approaches to red-teaming, while shedding light on interaction patterns and design insights for supporting human-AI collaboration in generative AI red-teaming.
Sources
- Who's Asking? Evaluating LLM Robustness to Inquiry Personas in Factual Question Answering
- RainbowPlus: Enhancing Adversarial Prompt Generation via Evolutionary Quality-Diversity Search
- Red Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned
- Ruby Teaming: Improving Quality Diversity Search with Memory for Automated Red Teaming
- How Well Can LLM Agents Simulate End-User Security and Privacy Attitudes and Behaviors?
- AI Transparency in the Age of LLMs: A Human-Centered Research Roadmap
- AutoDAN-Turbo: A Lifelong Agent for Strategy Self-Exploration to Jailbreak LLMs
- AutoDAN: Generating Stealthy Jailbreak Prompts on Aligned Large Language Models
- MIRAGE: Multi-model Interface for Reviewing and Auditing Generative Text-to-Image AI
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal
- Ferret: Faster and Effective Automated Red Teaming with Reward-Based Scoring Technique
- LLM Agents Grounded in Self-Reports Enable General-Purpose Simulation of Individuals
- Red Teaming Language Models with Language Models
- Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks
- Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming
- Understanding the Capabilities, Limitations, and Societal Impact of Large Language Models
- A Tale of Two Identities: An Ethical Audit of Human and AI-Crafted Personas
- Ethical and social risks of harm from Language Models
- GPTFUZZER: Red Teaming Large Language Models with Auto-Generated Jailbreak Prompts
Related papers
- EduGage: A Multimodal Dataset and Benchmark for Sensor-Based Momentary Assessment of Engagement in Self-Guided Video Learning
- EvoDesign: Agentic Editable Diagram Creation via Design Expertise Evolution
- HAGI++: Head-Assisted Gaze Imputation and Generation
- Linking Behaviour and Perception to Evaluate Meaningful Human Control over Partially Automated Driving
- Review of Explainable Decision Support and Adaptive Human-Machine Interfaces for Automation Transparency in Maritime Autonomous Surface Ships
- Towards Cognitive Process-Aware Proactive Writing Support