Caption Bottleneck Models

arXiv:2607.00578 · cs.CV · Submitted 2026-07-01 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Today's paper: "Caption Bottleneck Models".

Jane: Concept Bottleneck Models (CBMs) provide an interpretable-by-design framework for deep neural networks by routing predictions through human-understandable concepts,

Tom: First, who's behind it and why it matters.

Paper summary: Jane: So, to recap what we've discussed, the Caption Bottleneck Models paper lays out a proposal where they replace the traditional concept bottleneck with natural language descriptions generated by an LMM Jane. They claim this approach allows for interpretability while avoiding the problems associated with predefined concept sets and information leakage Jane.

Tom: Exactly. The core thesis is that instead of relying on expensive expert annotations or static lists of concepts, they use free-form language to drive the classification process directly Tom. They are training a classifier strictly on these textual descriptions derived from images Tom.

Lu: It really hinges on using a structured prompt with the LMM to generate diverse, attribute-centric captions for each image, which seems like the crucial first step in capturing rich visual information Lu.

Meng: I'm still processing how they manage that process without relying on those fixed concept vectors that we're used to seeing in other CBMs Meng. That transition from structured concepts to open-vocabulary language is a big architectural move for any AI system, isn't it?

Lalam: It means the system learns concepts organically from the descriptive text rather than being told what those concepts are upfront, which feels much more aligned with how human cognition works Lalam.

Tom: That organic discovery is exactly what they’re aiming for; they suggest that by letting the language guide the process, we can achieve a more natural form of concept routing through text Tom. They also focus on ensuring this setup is leakage-free by construction Tom.

Jane: And to build on that, they propose a deterministic censoring function to replace class names with generic referents like "the object" during caption generation Jane. This step seems designed specifically to prevent the downstream classification from relying on those specific label-name cues Jane.

Lu: That censoring function is clever because it forces the system to focus purely on descriptive visual attributes rather than memorizing label strings, which tackles a known issue directly Lu.

Meng: So, they are essentially creating a text-only recognition step that operates independently of the image pixels for its core prediction Meng. That separation is what makes it so appealing for practical deployment in many scenarios Meng.

Lalam: If the model is only reading descriptive attributes, it feels like we are getting a much purer signal about the visual content, which should lead to more robust and less brittle understanding over time Lalam.

Tom: That's the core mechanism they present in Caption Bottleneck Models: using text generation as a filter and training the classifier on that filtered language to achieve interpretable results Tom. It’s a solid foundation for our discussion of their overall approach Tom.

Conclusion: Jane: So, wrapping up what we've seen with Caption Bottleneck Models, the authors are proposing this framework as a way to move interpretation away from rigid concept layers toward free-form natural language Jane. They are showing that this method can effectively manage the challenges of defining concept sets by using LMMs and text classifiers trained on generated captions Jane.

Tom: And they really push the idea that this architecture is built to be leakage-free, which is a key feature for building trustworthy AI systems Tom. It’s about ensuring that visual features don't bypass the bottleneck to compromise the integrity of our explanations Tom.

Lu: The implication I see from this work is that we can finally achieve a level of concept discovery that is dataset-specific and truly open-vocabulary without needing constant manual intervention for every new application Lu. That flexibility opens up so many possibilities for future research directions Lu.

Meng: From an engineering viewpoint, if this method delivers competitive accuracy while simplifying the training pipeline by removing the need to define those concept layers upfront, it suggests a path toward more scalable and robust model deployment Meng. It sounds like we could see this applied across various domains quite quickly.

Lalam: I really think the biggest impact is that it makes AI interpretation accessible to everyone, moving us toward a system where understanding comes from descriptive language rather than some hidden internal mechanism Lalam. It changes how we interact with AI systems fundamentally Lalam.

Tom: So, in short, Caption Bottleneck Models is about leveraging natural language to build a more flexible and inherently interpretable AI architecture that discovers relevant concepts through text generation itself Tom. It’s a significant step forward for making our models more transparent Tom.

Seref Baris Cagliyan, Umut Ozdemir, Merve Tapli, Emre Akbas

Department of Computer Engineering, Middle East Technical University (METU) · Robotics & AI Center (ROMER), METU

cs.CV

Submitted: 2026-07-01

Updated: 2026-10-04

Comments: ECCV 2026

Journal ref: Computer Vision - ECCV 2026, LNCS 17036, pp. 445-461, Springer, 2026

DOI: 10.1007/978-3-032-37132-4_25

Code: https://github.com/bariscagliyan/CaptionBottleneckModels

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 92/100

The gist: Concept Bottleneck Models (CBMs) provide an interpretable-by-design framework for deep neural networks by routing predictions through human-understandable concepts, but they face challenges in

Key concepts

Taxonomy Censoring
This is a deterministic function applied to captions that replaces specific class names or higher-level groups with a generic referent like 'the object.' This ensures the final classification relies on descriptive visual attributes rather than relying on exact label names, preventing leakage from the original labels.
Gradient × Embedding Attribution
This technique calculates token-level saliency scores by multiplying the gradients of the output logits with their corresponding embedding vectors. This measures how much each specific word contributes to the final prediction, helping identify important descriptive phrases in a caption.
Span-Masking (Erasure) Scoring
This method tests the importance of a candidate phrase by measuring how much masking that span reduces the target class's logit. A large drop in logit indicates that the masked phrase is crucial for identifying the correct image class, thus scoring its relevance.
Semantic Clustering via Encoder Embeddings
Extracted phrases are mapped into a semantic space using the model's own encoder representations. HDBSCAN then groups similar phrases together, and these clusters are summarized by representative phrases. This process helps distill redundant concepts into meaningful, distinct categories.

Terminology

Summary

Concept Bottleneck Models (CBMs) provide an interpretable-by-design framework for deep neural networks by routing predictions through human-understandable concepts, but they face challenges in defining optimal concept sets and mitigating information leakage. This paper proposes Caption Bottleneck Models (CaBM), a novel framework that circumvents the need for predefined concept sets by replacing rigid layers with free-form natural language, ensuring a leakage-free architecture by construction while autonomously discovering high-quality, dataset-specific concepts.

The gist

CaBM is a framework that replaces rigid concept layers with free-form natural language by representing images via LMM-generated captions and training a classifier strictly on this text, ensuring a leakage-free architecture by construction.

How it works

The CaBM pipeline is structured in three distinct stages:

  1. Multi-Caption Generation with Taxonomy Censoring: Given an image dataset, the framework generates a set of diverse textual descriptions for each image using a Large Multimodal Model (LMM) guided by a fixed structured prompt. This process involves:

  2. Structured Prompting: An LLM proposes an attribute checklist based on the dataset domain and class granularity, resulting in a frozen prompt that explicitly discourages taxonomic naming and encourages grounded descriptions.

  3. Temperature-Driven Diversity: Captions are generated by varying the sampling temperature across multiple decoding runs to obtain multiple textual 'views' of the same image.

  4. Taxonomy Censoring: A deterministic censoring function, denoted as τ, is applied to each caption to replace any matched class string and higher-level taxonomic/group terms with a dataset-generic referent (e.g., “the object”), ensuring downstream classification relies on descriptive visual attributes rather than label-name cues.

Text Classification

The core recognition step involves training a text classifier on the resulting set of censored captions. This stage is designed to be strictly text-only, preventing information leakage from pixels. The process utilizes:

  1. Tokenization and Encoding: A pretrained transformer encoder (Encϕ) processes the censored captions, producing an output representation h.

  2. Linear Classifier: A linear classifier is applied to the final-layer [CLS] representation to predict image labels using the formula:

(3)

(3)

Open-Vocabulary Concept Discovery and Extraction

A central design choice of CaBM is that classification occurs entirely in the text space, allowing for evidence decomposition. The framework proposes an automated pipeline to extract concepts post-hoc:

  1. Correctness Filtering: Extraction is restricted to correctly classified training images.

  2. Gradient × Embedding Attribution: Token-level saliency scores are computed using gradient × embedding [1], defined as the raw saliency of token t, which is then normalized to obtain word-level scores.

  3. Candidate Phrase Proposal (Saliency N-grams): Candidate phrases are proposed using variable-length n-grams based on their sum of word saliencies, removing generic filler terms and selecting top candidates per caption with a maximum 50% span overlap. This yields up to KP candidate phrases per image.

  4. Span-Masking (Erasure) Scoring: Phrase importance is computed via an erasure test that measures how much masking a candidate span decreases the target-class logit, defined as:

(7)

  1. Semantic Clustering via Encoder Embeddings: Extracted phrases are mapped to a semantic embedding space using the trained CaBM encoder’s [CLS]-style pooled representation. HDBSCAN is then used within each class to cluster phrase embeddings, and redundancy is reduced through post-processing merge based on centroid similarity and string-level normalization. Each cluster G is summarized by a representative phrase p∗(G), and a concept importance score S(G) is assigned by aggregating the erasure evidence:

(8)

(9)

Evaluation and Results

The framework's effectiveness is evaluated across six benchmarks (CIFAR-10, CIFAR-100, CUB-200-2011, Food-101, Flowers-102, ImageNet-1K). Key results demonstrate:

(i)

CaBM achieves competitive accuracy while preserving interpretability without the constraints of external dictionaries or manual labeling. Specifically, in Table 4, CaBM achieves Top-1 test accuracy of 96.42% on CIFAR-10 and 81.14% on CIFAR-100, and 76.92% on CUB-200.

(ii)

The semantic quality of the discovered concept sets is quantified using metrics from HybridCBM: Purity, Separation, and Semantics. CaBM achieves high semantic validity (e.g., 0.92 for CUB-200) while matching purity (e.g.

Improvements for AI systems

Here are the specific improvements to AI systems that can be made by implementing Caption Bottleneck Models (CaBM), based on this research:

  1. Replacement of rigid, externally defined concept sets with free-form natural language descriptions for classification.

  2. Guaranteed leakage-free architecture by construction, ensuring that unmodeled visual features bypass the bottleneck layer and compromise explanation integrity are prevented because the classifier only observes discrete text produced upstream (captions), not raw pixels.

  3. Automated discovery of high-quality, dataset-specific concepts directly from the trained classifier's decision process by analyzing post-training text classification behavior, eliminating reliance on expensive expert annotations or LLM-generated lists based solely on class names.

  4. Creation of open-vocabulary concept vocabularies that are discovered through gradient attribution, span masking (erasure), and semantic clustering of phrases found in the captions, yielding concepts that are both human-meaningful and grounded in visual evidence (literal spans).

  5. Enabling robust test-time human interventions where modifying a single top-ranked concept injected into the caption systematically improves classification accuracy on misclassified instances, providing clear, actionable evidence for model debugging and verification.

  6. Development of a flexible, multi-view ensemble inference strategy where image predictions are derived by averaging logits from multiple diverse captions generated at varying sampling temperatures, reducing sensitivity to any single caption's bias.

The improved AI system (CaBM) can:

  1. Perform highly accurate image classification across various benchmarks (CIFAR-10/100, CUB-200, ImageNet-1K) while maintaining high interpretability.

  2. Generate a precise, human-readable explanation for every prediction by listing the exact phrases (concepts) extracted from the model's internal reasoning that led to the classification.

  3. Adapt to new datasets without requiring costly manual concept annotation; it discovers its own relevant visual concepts based on what it learns from the image captions.

  4. Be used in high-stakes environments (e.g., medical imaging, autonomous systems) where accountability is crucial, as its explanations are grounded in literal visual cues rather than opaque latent variables or external dictionaries.

  5. Serve as a reliable tool for concept validation, allowing researchers to verify if the model is relying on correct visual evidence by testing the system's sensitivity to injected concepts at inference time.

Sources

Related papers