Explanations that reveal all through the definition of encoding
Aahlad Puli, Nhi Nguyen, Rajesh Ranganath
cs.LG, cs.AI, stat.ML
Submitted: 2024-12-18
Updated: 2026-08-24
Terminology
Sources
- New-Onset Diabetes Assessment Using Artificial Intelligence-Enhanced Electrocardiography
- RISE: Randomized Input Sampling for Explanation of Black-box Models
- Learning to Faithfully Rationalize by Construction
- Goodhart's Law Applies to NLP's Explanation Benchmarks
- Explanations from Large Language Models Make Small Reasoners Better
- Leakage-Adjusted Simulatability: Can Models Generate Non-Trivial Explanations of Their Behavior in Natural Language?
- "Will You Find These Shortcuts?" A Protocol for Evaluating the Faithfulness of Input Salience Methods for Text Classification
- Logic Traps in Evaluating Attribution Scores
- Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness?
- Faithfulness Tests for Natural Language Explanations
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks