A Unifying Framework of Concept-based Explainable AI with Completeness Guarantees
cs.LG, cs.AI, cs.CV
Submitted: 2026-09-28
Updated: 2026-09-28
Terminology
Sources
- Network Dissection: Quantifying Interpretability of Deep Visual Representations
- Subspace-Aware Sparse Autoencoders for Effective Mechanistic Interpretability
- Concept Embedding Models: Beyond the Accuracy-Explainability Trade-Off
- CRAFT: Concept Recursive Activation FacTorization for Explainability
- Net2Vec: Quantifying and Explaining how Concepts are Encoded by Filters in Deep Neural Networks
- Towards Human-Understandable Multi-Dimensional Concept Discovery
- Neural Prototype Trees for Interpretable Fine-grained Image Recognition
- Label-Free Concept Bottleneck Models
- Explainable Artificial Intelligence: Understanding, Visualizing and Interpreting Deep Learning Models
- Language in a Bottle: Language Model Guided Concept Bottlenecks for Interpretable Image Classification
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks