Intrinsic Self-Correction in LLMs: Towards Explainable Prompting via Mechanistic Interpretability

arXiv:2505.11924 · cs.CL, cs.AI, cs.LG · Submitted 2026-08-21 · Read on arXiv

Yu-Ting Lee, Fu-Chieh Chang, Yu-En Shu, Hui-Ying Shih, Pei-Yuan Wu

cs.CL, cs.AI, cs.LG

Submitted: 2026-08-21

Updated: 2026-08-24

Code: https://github.com/Yu-TingLee/llm-self-correction-mech-interp

Terminology

Sources

Related papers