Towards Fast and Disentangled Counterfactuals for Visual Foundation Models
cs.LG, cs.CV
Submitted: 2026-10-01
Updated: 2026-10-01
Terminology
Sources
- DINOv3
- Interpreting CLIP with Hierarchical Sparse Autoencoders
- Debiasing Vision-Language Models via Biased Prompts
- Distilling Lightweight Domain Experts from Large ML Models by Identifying Relevant Subspaces
- Investigating the Robustness of Subtask Distillation under Spurious Correlation
- Imbalanced Classification through the Lens of Spurious Correlations
- Reproducibility study on how to find Spurious Correlations, Shortcut Learning, Clever Hans or Group-Distributional non-robustness and how to fix them
- Faithful, Sufficient and Understandable: Rethinking Graph Counterfactual Explanations via Discrete Diffusion Inversion
- Protein Counterfactuals via Diffusion-Guided Latent Optimization
- SCE-LITE-HQ: Smooth visual counterfactual explanations with generative foundation models
- Hierarchical Text-Conditional Image Generation with CLIP Latents
- OrtSAE: Orthogonal Sparse Autoencoders Uncover Atomic Features
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks