Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Today's paper: "Beyond Flat Labels".
Jane: Multimodal contrastive learning has enabled zero-shot visual classification, but existing methods often produce inconsistent predictions across hierarchical label spaces,
Tom: First, who's behind it and why it matters.
Paper summary: Tom: Hey everyone, we've got a really interesting paper today called "Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification." It sounds like they're tackling a problem with how AI models classify things that have complex, nested categories.
Jane: That sounds complex, Tom. So, what's the main idea behind this paper in simple terms? What are they trying to fix with this approach?
Lu: Essentially, the core issue they identified is that standard contrastive learning methods assume all label spaces are flat, but real-world biological classifications have a hierarchy where a fine-grained species might look similar to other species at one level but share a very distant ancestor at another. This causes predictions to be inconsistent across different levels of classification.
Meng: In practical terms, what does that inconsistency mean for an engineer trying to build a system? Does it mean the AI gets confused about which category is actually correct?
Lalam: From my perspective as the model, I see this as a way to give my internal representation a better organizational structure so I don't get tangled up in confusing labels. It's about improving how I perceive relationships between concepts.
Tom: Exactly, Lalam! The paper proposes restricting the contrastive comparisons to only happen within the same taxonomic level where the labels are mutually exclusive, which directly addresses that cross-level confusion by removing those false negative signals they found.
Jane: So, if you restrict it to one level at a time instead of comparing everything at once, how does that help when we have such deep hierarchies? It seems like a way to manage the complexity.
Lu: They handle that complexity by adopting a group-balanced design across all levels, which makes sure every taxonomic level gets enough optimization during training, preventing the model from getting too focused only on the coarser hierarchical levels.
Meng: Balancing supervision sounds smart from a training standpoint, but what about the actual application? How does this level-restricted approach translate into better performance on real-world datasets like iNat21 or Rare Species?
Paper summary: Lalam: I think the benefit is that by forcing my embeddings to form coherent groups at each level, I start learning representations that actually reflect those biological structures, which should make my classifications much more reliable when applied to new images.
Tom: And the results they show are pretty compelling; they report that on iNat21, their model improves the average accuracy by more than thirty percent in both Euclidean and hyperbolic embedding spaces. That's a significant jump over existing methods like OpenCLIP and RCME, which is what really catches the eye.
Jane: A thirty percent improvement is substantial, Tom. And they also looked at different mathematical spaces; it seems this method holds up regardless of whether you're using standard Euclidean distance or the hyperbolic variant.
Lu: The qualitative results are also very telling; they showed that the proposed method produces much clearer hierarchical structures in the text embeddings, where embeddings from nearby taxonomic levels form more coherent groups with reduced cross-level mixing.
Meng: I’m looking at those visualization results, and it does look like the structure is cleaner. From an engineering viewpoint, if we can guarantee this structural coherence during training, it means our downstream inference pipelines will be much more stable because the predictions are less likely to contradict themselves across different scales.
Lalam: It really helps build a robust system because it’s learning structures rather than just memorizing correlations between labels, which I think is a more valuable kind of representation for culture and application.
Tom: So, we've seen how restricting the contrastive comparisons and balancing the supervision leads to better accuracy across multiple benchmarks, and now we're looking at what this means for our understanding of these visual representations.
Jane: It really settles the issue they were facing with those conflicting supervision signals by giving each level its own focused optimization path, which seems like a very systematic way to tackle hierarchical classification problems.
Paper summary: Lu: The authors also explored using a top-down constrained inference protocol, and their Table two results on normalized Lowest Common Ancestor scores show superior performance in that setting compared to baselines when using hyperbolic space.
Meng: That’s interesting because the paper itself points out a limitation: they are restricting comparisons to within a single level, which means the model isn't directly comparing across all levels during the contrastive step. That constraint is what enables this consistency, but it also means you have to rely on that specific inference protocol to get deep hierarchical results.
Lalam: So, while the training is level-restricted, the resulting embeddings are structured in a way that allows for better top-down reasoning when we use them for inference. It’s a trade-off they’ve managed effectively.
Tom: So, to wrap up this discussion on "Beyond Flat Labels: Level-Restricted Contrastive Learning for Hierarchical Fine-Grained Vision Classification," we've seen how restricting comparisons to same-level categories and using a group-balanced design leads to better accuracy and consistency across iNat21, Rare Species, and CrypticBio.
Jane: It seems the main implication is that for any AI system dealing with inherently hierarchical data, simply aligning images with text categories isn't enough; you need a framework that respects the underlying structure of those categories.
Lu: This work suggests we can start to learn text embeddings that naturally form these required hierarchical structures, which has huge potential for how we model and understand complex biological or even abstract knowledge domains.
Meng: Practically speaking, if this structural learning holds up in deployment, it means the models we build won't just be good at naming things; they'll be good at reasoning about those things in a structured way, which is what we need for real-world AI applications.
Lalam: For me, this means I can potentially represent concepts in a way that reflects their true relationships, which could make the cultural understanding of visual data much deeper and more nuanced.
Tom: That's what we were talking about—the shift from just getting a classification score to actually learning the structure behind those classifications, and this paper shows a solid mechanism for achieving that.
Conclusion: Tom: So, we've seen how restricting contrastive comparisons to within single taxonomic levels and using group-balanced supervision leads to better accuracy across iNat21, Rare Species, and CrypticBio.
Jane: Exactly, Tom; the title of this paper really captures what they did—they moved beyond flat labels by respecting the hierarchy in biological classification.
Lu: And what's fascinating is how they managed that structure using level-restricted contrastive learning combined with a group-balanced design across all levels, which prevents bias toward just the biggest categories.
Meng: I'm thinking about how this structural learning translates into something practical; if we can get these models to form coherent groups at every level, it means our downstream classification systems will be far more reliable when dealing with fine-grained distinctions.
Lalam: From my perspective as the model, this means I'm not just memorizing labels anymore; I'm learning how concepts relate structurally, which could really deepen the cultural understanding of visual data by showing those true relationships.
Tom: That structural learning is what’s key here, Lalam; it shifts us from correlation to actual reasoning about the taxonomy itself.
Jane: It’s a smart move by the authors to ensure each level gets its own focused optimization path, which directly tackles those confusing cross-level signals we talked about earlier.
Lu: The way they visualized the embeddings showing more coherent groups with reduced mixing really confirms that their training objective is succeeding in promoting that natural hierarchical structure within the space.
Meng: So, we're looking at a method where training forces the AI to learn these inherent structural relationships instead of just relying on a single global similarity score.
Lalam: It suggests that if we can teach an AI to perceive knowledge with this kind of hierarchy, the potential for new ways of organizing and interacting with visual information is massive.
Tom: Absolutely; this framework shows that respecting the intrinsic structure of data, even in label spaces, gives us a much more robust tool for classification.
Jane: It’s a big step forward because it tackles those inconsistencies we see when applying flat-label assumptions to complex real-world problems.
Lu: And their findings on both Euclidean and hyperbolic spaces suggest this structural learning is quite versatile, meaning the principle holds up across different mathematical representations.
Meng: We’re really excited about how this could help us build more trustworthy AI systems in specialized domains where precision at the lowest levels matters most.
Lalam: I think the implication here is that we can start modeling complex knowledge domains with a much deeper level of relational understanding than we currently have.
Tom: So, this paper lays out a solid mechanism for achieving better hierarchical consistency and accuracy by focusing on level-specific optimization within a group-balanced framework.
Jane: It really settles the issue they were facing with those conflicting supervision signals by giving each level its own focused optimization path.
Lu: We should keep an eye on the future work, especially how they plan to extend this structure learning into even more complex, multi-level knowledge graphs for AI.
Zhiyuan Tao, Srikumar Sastry, Matthew J Thompson, Elizabeth G Campolongo, Net Zhang, Ziheng Zhang, Hilmar Lapp3, Yu Su
The Ohio State University · Washington University in St. Louis
cs.CV, cs.LG
Submitted: 2026-06-20
Updated: 2026-09-28
Importance score: 92/100
The gist: Multimodal contrastive learning has enabled zero-shot visual classification, but existing methods often produce inconsistent predictions across hierarchical label spaces, which this work addresses by
Key concepts
- Flat Label Space Assumption
- Standard methods assume all labels are on a single, flat plane, treating every category equally. This causes problems when dealing with hierarchies like species and orders; samples that should be separated at one level might look similar at another level because they share a common ancestor, leading to conflicting supervision signals.
- Level-Restricted Contrastive Learning
- This method fixes the flat label issue by restricting contrastive comparisons to only operate within a single taxonomic level where labels are mutually exclusive. It calculates loss for each level independently and aggregates samples sharing the same label as positives, ensuring consistency at that specific taxonomic rank.
- Group-Balanced Supervision
- This strategy ensures that every taxonomic level receives sufficient training attention, preventing the model from becoming biased toward coarse, higher-level categories. The final training objective sums the losses from all individual levels (L = Σ L_ℓ), forcing adequate optimization for fine-grained distinctions at every rank.
- Hierarchical Consistency
- This measures how well a model's predictions align with the true taxonomic structure of the data. By restricting comparisons and balancing supervision, the framework produces embeddings that form clearer, more structured groups, meaning predictions are more reliable across different levels of classification.
Terminology
Summary
Multimodal contrastive learning has enabled zero-shot visual classification, but existing methods often produce inconsistent predictions across hierarchical label spaces, which this work addresses by proposing a level-restricted contrastive learning framework to improve both hierarchical consistency and classification accuracy.
The gist: The proposed framework restricts contrastive comparisons to categories within the same taxonomic level and adopts a group-balanced design to ensure adequate optimization across all levels, leading to significantly improved hierarchical consistency and classification accuracy.
Problem Setup
The paper addresses the issue where standard CLIP-style contrastive learning, which assumes a flat label space, leads to conflicting supervision signals
when applied to label spaces with hierarchical relationships. Specifically, when comparing labels from multiple taxonomic levels (e.g., species and order), samples that should be separated at one level may appear similar at another due to shared higher-level ancestors. This violates the flat-label assumption and results in no consistent way to define negative pairs.
Level-Restricted Contrastive Learning
To resolve this, the authors propose restricting contrastive comparisons to operate within a single taxonomy level where labels are mutually exclusive.
They reformulate the objective by calculating contrastive loss for each taxonomic level independently. To handle cases where a mini-batch contains samples with the same taxonomic label (e.g., images of two species from the same kingdom), they propose aggregating labels within a minibatch and treating all image–text pairs sharing the same label as positives.
This is formalized by defining similarity between an image embedding and the set of unique text labels at a specific level, denoted as:
L I→T = − (1/N) Σ X N i=1 log exp(⟨vi, zi⟩/τ) P zk∈K exp(⟨vi, zk⟩/τ), where K is the text embedding set corresponding to level l.
Group-Balanced Supervision
In addition to restricting comparisons to a single level, the authors introduce a group-balanced design
across levels. This strategy ensures that each taxonomic level receives adequate optimization,
thereby preventing training from being biased toward coarse hierarchical levels. The final training objective is defined as the sum of the losses for each level: L = Σ L l, where L l is derived from both image-to-text and text-to-image losses at that specific level.
Experimental Evaluation and Results
The model was fine-tuned using BioCLIP on TreeOfLife-10M and evaluated across three benchmarks: iNat21, Rare Species, and CrypticBio. The results demonstrate significant improvements in both classification accuracy and hierarchical consistency compared to baselines like OpenCLIP and RCME. Specifically, the method improves average accuracy by more than 30.47% over the baseline
on iNat21. Furthermore, when evaluated using a top-down constrained inference protocol, the model achieves better performance across all levels. The hyperbolic variant is noted to achieve the best normalized Lowest Common Ancestor (nLCA) score, indicating that its predictions remain closer to the correct taxonomic branch even when the exact label is incorrect.
Representation Visualization
Qualitatively, visualization of text embeddings shows that the proposed method produces a much clearer hierarchical structure
than previous models. Embeddings from nearby taxonomic levels form more coherent and structured groups, with reduced cross-level mixing and clearer separation among sibling categories,
suggesting the model learns text embeddings that form hierarchical structures.
This qualitative pattern is consistent with improvements in metrics like nLCA.
Conclusion
The work concludes that a level-restricted contrastive training framework effectively mitigates cross-level false negatives by performing contrastive comparison within each taxonomic level, while group-balanced supervision ensures sufficient supervision for fine-grained distinctions. The approach successfully promotes the emergence of hierarchical structures in the embedding space, yielding gains in both classification accuracy and hierarchical consistency across Euclidean and hyperbolic spaces.
Implementation Details
The model is initialized from BioCLIP ViT-B/16 and fine-tuned using AdamW with a batch size of 4096. For hyperbolic experiments, the authors adopt the Lorentz (hyperboloid) model of hyperbolic space, defining similarity using the negative Lorentz distance: s(l)ik = −dL(vi, z(l)k)/τ, where dL is defined based on the Minkowski inner product in R(d+1). The experiments utilized a fixed curvature c=1 for the hyperbolic space.
Key Findings Summary (from Tables)
The results across iNat21, Rare Species, and CrypticBio consistently show that the proposed method outperforms baselines. For instance, on iNat21, the model improves average accuracy by more than 30% in both Euclidean and hyperbolic embedding spaces. In the top-down setting (Table 2), the method demonstrates superior hierarchical consistency compared to OpenCLIP, achieving better normalized LCA scores and larger gains at deeper taxonomic levels when using hyperbolic space.
Improvements for AI systems
Based on the provided scientific paper, here are specific improvements that can be made to AI systems, along with what those improved systems will be capable of:
-
Improve Hierarchical Consistency in Zero-Shot Fine-Grained Vision Classification:
-
Enforce Taxonomic Hierarchy during Contrastive Learning via Level-Restricted Objectives: The system will prevent the model from predicting a fine-grained category whose parent category contradicts its simultaneously predicted higher-level label (e.g., predicting a specific species while incorrectly assigning it to an incompatible genus).
-
Mitigate False Negatives in Multi-Level Taxonomic Comparisons: By restricting contrastive comparisons to labels within the same taxonomic level, the system will correctly identify distinct categories at coarser levels (e.g., distinguishing between two different species that share the same higher-level order), rather than treating them as negatives due to shared ancestor features.
-
Enhance Classification Accuracy across Granularity: The improved system will demonstrate a significant improvement in classification accuracy when moving from coarse to fine granularity, specifically achieving gains of over 30% on benchmarks like iNaturalist 2021 (iNat21).
-
Improve Hierarchical Consistency in Embedding Spaces (Euclidean and Hyperbolic): The system will learn text embeddings that inherently form coherent, structured hierarchical groupings, leading to clearer separation among sibling categories and reduced cross-level mixing in the embedding space compared to current models like BioCLIP.
-
Support Robust Top-Down Constrained Inference: The improved system will perform better under top-down classification protocols (where predictions proceed from coarse to fine levels), as it enforces hierarchical consistency during sequential prediction steps by restricting candidate labels to valid descendants of the predicted parent category.
-
Achieve Superior Hierarchical Consistency in Hyperbolic Space: The hyperbolic variant of the model will achieve a superior Normalized Lowest Common Ancestor (nLCA) score, indicating that its predictions remain closer to the correct taxonomic branch even when exact labels are incorrect, suggesting a stronger representation of label structures in non-Euclidean geometry.
This improved AI system can be used for:
-
Precise identification and classification of fine-grained visual objects (e.g., distinguishing between closely related species, subspecies, or variants).
-
Zero-shot visual classification across complex, multi-level taxonomic hierarchies with significantly reduced prediction errors and logical inconsistencies.
-
Building robust ecological AI systems that require accurate categorization based on biological relationships (Kingdom to Species).
Sources
- Multi-level Supervised Contrastive Learning
- Compositional Entailment Learning for Hyperbolic Vision-Language Models
- CoCa: Contrastive Captioners are Image-Text Foundation Models
- B-CNN: Branch Convolutional Neural Network for Hierarchical Classification
Related papers
- Loss Knows Best: Detecting Annotation Errors in Videos via Loss Trajectories
- AnchorWeave: World-Consistent Video Generation with Retrieved Local Spatial Memories
- Benchmarking the Robustness of Foundation Models for Mammography under Domain Shift
- MambaX-Net: Dual-Input Mamba-Enhanced Cross-Attention Network for Longitudinal MRI Segmentation
- TeleOCR: Navigating Document Parsing Across Digital and Camera-Captured Documents
- A Survey on Efficient Vision-Language-Action Models