ReBeCA: Unveiling Interpretable Behavior Hierarchy behind the Iterative Self-Reflection of Language Models with Causal Analysis
cs.CL
Submitted: 2026-02-06
Updated: 2026-09-06
Comments: Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Terminology
Sources
- Invariant Risk Minimization
- Qwen Technical Report
- Teaching Large Language Models to Self-Debug
- CRITIC: Large Language Models Can Self-Correct with Tool-Interactive Critiquing
- Measuring Mathematical Problem Solving With the MATH Dataset
- Large Language Models Cannot Self-Correct Reasoning Yet
- Why Language Models Hallucinate
- SelfReflect: Can LLMs Communicate Their Internal Answer Distribution?
- LLMs Get Lost In Multi-Turn Conversation
- When Hindsight is Not 20/20: Testing Limits on Reflective Thinking in Large Language Models
- Auditing Meta-Cognitive Hallucinations in Reasoning Large Language Models
- REFINER: Reasoning Feedback on Intermediate Representations
- GPT-4 Technical Report
- Shepherd: A Critic for Language Model Generation
- Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement
- The Clever Hans Mirage: A Comprehensive Survey on Spurious Correlations in Machine Learning
- Why Does ChatGPT Fall Short in Providing Truthful Answers?
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering