How the Audit Rule Shapes Faithful Factor Explanations in LLMs
cs.CL
Submitted: 2026-10-01
Updated: 2026-10-01
Terminology
Sources
- Learning to Give Checkable Answers with Prover-Verifier Games
- Measuring Progress on Scalable Oversight for Large Language Models
- Training Verifiers to Solve Math Word Problems
- DeepSeek-V4: Towards Highly Efficient Million-Token Context Intelligence
- RFEval: Benchmarking Reasoning Faithfulness under Counterfactual Reasoning Intervention in Large Reasoning Models
- AI safety via debate
- Prover-Verifier Games improve legibility of LLM outputs
- Measuring Faithfulness in Chain-of-Thought Reasoning
- A Positive Case for Faithfulness: LLM Self-Explanations Help Predict Model Behavior
- Truthfulness Despite Weak Supervision: Evaluating and Training LLMs Using Peer Prediction
- Question Decomposition Improves the Faithfulness of Model-Generated Reasoning
- ElicitationGPT: Text Elicitation Mechanisms via Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering