LLM Persona Unlearning
cs.CL, cs.LG
Submitted: 2026-09-30
Updated: 2026-10-08
Code: https://github.com/meta-llama/llama-models
Terminology
Sources
- Constitutional AI: Harmlessness from AI Feedback
- Where is the Mind? Persona Vectors and LLM Individuation
- Jailbreaking Black Box Large Language Models in Twenty Queries
- Persona Vectors: Monitoring and Controlling Character Traits in Language Models
- Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
- Training Verifiers to Solve Math Word Problems
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- OpenUnlearning: Accelerating LLM Unlearning via Unified Benchmarking of Methods and Metrics
- Who's Harry Potter? Approximate Unlearning in LLMs
- Gemma 4 Technical Report
- The Llama 3 Herd of Models
- Rethinking Machine Unlearning for Large Language Models
- The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models
- Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models
- Emotion Concepts and their Function in a Large Language Model
- Steering Language Models With Activation Engineering
- Persona Features Control Emergent Misalignment
- Instruction-Following Evaluation for Large Language Models
- Representation Engineering: A Top-Down Approach to AI Transparency
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering