Less Sycophancy, Stronger Refusal? Lessons for AI Safety from Mechanistic Interpretability

arXiv:2609.35544 · cs.CL, cs.AI, cs.LG · Submitted 2026-09-28 · Read on arXiv

cs.CL, cs.AI, cs.LG

Submitted: 2026-09-28

Updated: 2026-09-28

Code: https://github.com/Xu0615/Sycophancy_Safety_via_SAE

Terminology

Sources

Related papers