Evaluating Language Model Safety Across Long Adversarial Conversations

arXiv:2609.38357 · cs.CL · Submitted 2026-09-29 · Read on arXiv

cs.CL

Submitted: 2026-09-29

Updated: 2026-09-29

Code: https://github.com/meta-llama/PurpleLlama

Terminology

Sources

Related papers