Generative Interpretability via Scalable Neuro-Symbolic Models
cs.LG, cs.AI, cs.CL, cs.SC
Submitted: 2026-09-11
Updated: 2026-10-01
Comments: ACM AI Summit 2026
License: http://creativecommons.org/licenses/by/4.0/
The gist: As the use of Large Language Models moves from chatbots into agentic systems, where outputs become actions with irreversible consequences on reality, the existing paradigm on AI Interpretability
Terminology
Abstract
As the use of Large Language Models moves from chatbots into agentic systems, where outputs become actions with irreversible consequences on reality, the existing paradigm on AI Interpretability research, post-hoc interpretability, is structurally inadequate for safe and trustworthy model deployment: it explains behavior after the fact but cannot audit or intervene in an inference computation before it commits to an output. We therefore argue for a shift toward generative interpretability, an architectural property under which a model's inference pass natively exposes semantically meaningful checkpoints that are human-understandable and amenable to causal intervention. We show the merits of generative interpretability as comparison to other interpretability research paradigms, and propose Neuro-Symbolic Models as a concrete instantiation.
Sources
- Chain-of-Thought Reasoning In The Wild Is Not Always Faithful
- Interpretability without actionability: mechanistic methods cannot correct language model errors despite near-perfect internal representations
- Probing Classifiers: Promises, Shortcomings, and Advances
- International AI Safety Report 2026
- BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding
- Causal Abstraction: A Theoretical Foundation for Mechanistic Interpretability
- Optimal Sparse Decision Trees
- Towards Faithfully Interpretable NLP Systems: How should we define and evaluate faithfulness?
- Concept Bottleneck Models
- The Mythos of Model Interpretability
- Locating and Editing Factual Associations in GPT
- Explanation in Artificial Intelligence: Insights from the Social Sciences
- The Quest for the Right Mediator: Surveying Mechanistic Interpretability Through the Lens of Causal Mediation Analysis
- LLM Explainability with Counterfactual Chains and Causal Graphs
- Direct and Indirect Effects
- On efficiently computable functions, deep networks and sparse compositionality
- Stop Explaining Black Box Machine Learning Models for High Stakes Decisions and Use Interpretable Models Instead
- Beyond the Commitment Boundary: Probing Epiphenomenal Chain-of-Thought in Large Reasoning Models
- Concept Bottleneck Large Language Models
- The many Shapley values for model explanation
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks