Hyperbolic Multimodal Continual Learning
Jiahong Liu, Ming Shen, Xiaohao Liu, Rex Ying, Menglin Yang, Tat-Seng Chua, Irwin King
The Chinese University of Hong Kong · National University of Singapore · Yale University · The Hong Kong University of Science and Technology (Guangzhou)
cs.LG
Submitted: 2026-08-10
Updated: 2026-08-11
Comments: ICML 2026. 33 pages, 11 figures. Code: https://github.com/HUBERILT/HMCL_ICML
Code: https://github.com/HUBERILT/HMCL_ICML
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 95/100
The gist: This paper presents the first systematic study of continual learning for hyperbolic multimodal representations.
Terminology
Summary
This paper presents the first systematic study of continual learning for hyperbolic multimodal representations. The authors establish a theoretical foundation showing that preventing forgetting requires cross-modal invariance under a shared hyperbolic isometry
and that forgetting in hyperbolic continual learning involves both semantic relation drift and hierarchy-related distortion.
They derive a principled continual learning framework called HMCL (Hyperbolic Multimodal Continual Learning) that preserves essential geometric structure while allowing effective adaptation to new tasks.
The paper addresses a gap in existing research: existing hyperbolic multimodal models are almost exclusively developed under static learning assumptions, where the full data distribution is available upfront.
In realistic multimodal systems, however, data distributions evolve over time as new concepts, domains, and tasks emerge,
leaving continual learning in hyperbolic representation spaces largely unexplored.
The authors make three main contributions:
-
Theoretical foundation: They establish "a theoretical foundation for representation preservation in hyperbolic space, showing that it requires cross-modal invariance under a shared hyperbolic isometry (e.g., the same rotation), yielding necessary and sufficient geometric conditions."
-
Principled framework: They derive
a principled continual learning framework for hyperbolic multimodal representations, which preserves essential geometric structure while allowing effective adaptation to new tasks.
-
Empirical validation:
Across diverse continual multimodal benchmarks, HMCL consistently reduces catastrophic forgetting while adapting to new tasks.
The paper defines three preservation conditions for hyperbolic multimodal continual learning:
-
(P1) Intra-modal preservation:
Intra-modal preservation ensures that domain-specific knowledge within each modality remains intact
-
(P2) Inter-modal preservation:
Inter-modal preservation maintains the learned cross-modal associations that enable effective multimodal understanding
-
(P3) Hierarchical preservation:
Hierarchical preservation protects the conceptual hierarchies captured by the hyperbolic cone structure
The central challenge is that the stability conditions (P1)–(P3) are principled but cumbersome to enforce or analyze directly, motivating a simpler equivalent characterization.
The paper proves that "the preservation conditions (P1), (P2), and (P3) hold for task t−1 representations between stages t−1 and t if and only if there exists a shared transformation of the form R = [[1, 0 T], [0, R̃]] with R̃ ∈ SO(d), such that for all modalities m ∈ 1,..., M, Z m,t t−1 = Z m,t−1 t−1 R."
This means that preserving old-task knowledge requires updates to act as a shared hyperbolic isometry, maintaining both properties
of semantic similarity and hierarchy.
The paper shows that preserving the hierarchy of old-task embeddings enforces that the spatial update is orthogonal to the current embedding's spatial direction,
specifically requiring that z T[1:d] Δz[1:d] = 0. This means admissible updates are restricted within the tangent space so as to suppress first-order drift along the time-like dimension.
The paper derives that preserving intramodal, inter-modal, and hierarchical structure in turn induces a strict first-order constraint on admissible parameter updates.
The canonical admissible update is given by:
ΔW m s = δ m s (I − P t−1)
where P t−1 is a projection matrix onto the old-task spatial subspace. This provides a closed-form and geometry-consistent parameter update rule.
The HMCL framework works as follows:
-
Compute
a low-rank orthonormal basis V of Z m,* t−1,s via SVD (or PCA)
-
Project the unconstrained update: ΔW m s = δ m s (I − VV T)
-
Set the time-like update to zero: Δw m 0 = 0
-
Update parameters accordingly
The method introduces no additional learnable parameters
and the memory overhead is constant with respect to the number of tasks.
The authors evaluate on a continual multimodal learning setting with both classification and retrieval tasks
using up to 15 datasets, including image–text classification benchmarks (e.g., CIFAR-10/100, Caltech-101, Food-101) and cross-modal retrieval benchmarks (COCO and Flickr30k).
They use three hyperbolic backbones: MERU-L, MERU-B, and HyCoCLIP-B, with all representations modeled in the Lorentz (hyperboloid) space with fixed curvature K = 0.1.
Baselines include Vanilla (sequential fine-tuning), EWC, GEM, and C-FLAT, all adapted to the hyperbolic setting.
HMCL consistently improves overall performance across backbones.
Key results include:
-
On MERU-L:
HMCL improves the Overall metric from 38.22 to 41.36, corresponding to a relative gain of +8.2% over the Vanilla baseline
-
Classification BWT improves
from −6.46 to −0.74, corresponding to an 88.5% relative reduction in forgetting
-
On HyCoCLIP-B:
HMCL increases image-to-text R@1 by +26.9%
The paper notes that conventional continual learning methods struggle in hyperbolic multimodal settings
and that geometry-aware constraints are key to stable multimodal continual learning.
The paper measures four types of drift: radial, angular, cross-modal, and paired-distance. Results show HMCL reduces:
-
Radial drift by 61.0%
-
Angular drift by 57.8%
-
Cross-modal drift by 73.8%
-
Paired-distance drift by 59.5%
These reductions are statistically significant across old tasks (p ≤ 3.1 × 10−4).
Using image traversals on Flickr30K, the paper shows that HMCL yields coherent hierarchies (e.g., hat → fashion → style), whereas Vanilla drifts toward irrelevant or inconsistent concepts after continual training.
The traversal paths demonstrate that HMCL preserves not only nearest-neighbor retrieval quality, but also the abstraction ordering encoded by the hyperbolic radial geometry.
The paper concludes that "preserving previously learned knowledge requires cross-modal invariance under a shared hyperbolic isometry, and that forgetting is primarily driven by distortions along the radial dimension that encodes semantic hierarchy. The HMCL framework
restricts parameter updates to geometry-preserving directions and
consistently improves both performance and stability, substantially reducing catastrophic forgetting compared to existing continual learning methods."
The authors note limitations and future directions: Combining this route with replay-style mechanisms, such as exemplar/generative replay or hyperbolic memory selection, and extending it to full-backbone adaptation remain promising directions.
Improvements for AI systems
Improvements to AI Systems Based on This Paper:
- Geometry-Aware Continual Learning for Multimodal Models
-
Improvement: Integrate HMCL’s projection-based parameter updates (using SVD-derived orthonormal bases and zeroing time-like components) into existing multimodal transformers (e.g., CLIP, MERU, HyCoCLIP) to enforce shared hyperbolic isometry during sequential task learning.
-
Capability: The AI system can learn new tasks (e.g., new image-text domains) without catastrophic forgetting, maintaining stable cross-modal retrieval and classification accuracy across up to 15 datasets, with up to 88.5% reduction in backward transfer forgetting.
- Hierarchical-Semantics Preservation in Dynamic Data Streams
-
Improvement: Apply the theoretical constraint (P3) that spatial updates must be orthogonal to current embedding directions, preserving radial hierarchy in hyperbolic space.
-
Capability: The system retains abstract-to-specific concept ordering (e.g.,
hat → fashion → style
) even after fine-tuning on new data, enabling reliable zero-shot hierarchical reasoning and concept navigation in evolving knowledge bases.
- Low-Overhead, Memory-Efficient Adaptation
-
Improvement: Use HMCL’s constant-memory update rule (no extra learnable parameters, only a fixed orthonormal basis per task) in edge or resource-constrained deployments.
-
Capability: The AI can continuously adapt to new multimodal tasks on-device (e.g., mobile vision-language assistants) with negligible memory growth, while still outperforming replay-free baselines and matching or exceeding methods that require stored exemplars.
- Drift-Aware Representation Stability for Retrieval Systems
-
Improvement: Implement HMCL’s drift reduction (radial −61%, angular −57.8%, cross-modal −73.8%, paired-distance −59.5%) as a regularizer in production retrieval engines that update embeddings incrementally.
-
Capability: The system maintains high-quality cross-modal search (image↔text) over time, with statistically significant stability (p ≤ 3.1×10−4), preventing performance collapse when new items are added without full retraining.
- Theoretically Grounded Fine-Tuning for Foundation Models
-
Improvement: Use Theorem 1’s necessary-and-sufficient condition (shared hyperbolic rotation) to design fine-tuning schedules for large multimodal models on sequential downstream tasks.
-
Capability: The AI can be safely fine-tuned on new benchmarks (e.g., COCO after CIFAR) without distorting its original semantic geometry, improving overall accuracy (e.g., +8.2% relative gain on MERU-L) while preserving generalization to unseen tasks.
- Plug-and-Play Module for Existing Continual Learning Frameworks
-
Improvement: Replace or augment standard regularizers (EWC, GEM, C-FLAT) with HMCL’s closed-form update projection, which is compatible with any hyperbolic embedding backbone.
-
Capability: The improved system can be deployed as a drop-in upgrade to current continual learning pipelines, yielding consistent gains across different backbones (MERU-L, MERU-B, HyCoCLIP-B) and task types (classification and retrieval) without architectural changes.
Abstract
Hyperbolic geometry has recently emerged as a powerful representation space for multimodal learning, as it naturally captures hierarchical semantic structure across modalities. Despite this progress, how such representations behave under continual learning poses fundamentally different challenges that remain underexplored. This work provides a geometric perspective on this problem and establishes a theoretical foundation for representation preservation in hyperbolic space, showing that preventing forgetting requires cross-modal invariance under a shared hyperbolic isometry. We further show that forgetting in hyperbolic continual learning involves both semantic relation drift and hierarchy-related distortion, motivating preservation of both cross-modal relational structure and hierarchical geometry. Guided by these insights, a principled continual learning framework is derived that preserves essential geometric structure while allowing effective adaptation to new tasks. Experiments on continual multimodal benchmarks corroborate the effectiveness of the proposed approach.
Sources
- Describing Textures in the Wild
- A streamlined Approach to Multimodal Few-Shot Class Incremental Learning for Fine-Grained Datasets
- What to align in multimodal contrastive learning?
- HC-GLAD: Dual Hyperbolic Contrastive Learning for Unsupervised Graph-Level Anomaly Detection
- ImageBind-LLM: Multi-modality Instruction Tuning
- Position: Beyond Euclidean -- Foundation Models Should Embrace Non-Euclidean Geometries
- EuroSAT: A Novel Dataset and Deep Learning Benchmark for Land Use and Land Cover Classification
- CLEVR: A Diagnostic Dataset for Compositional Language and Elementary Visual Reasoning
- Continual Hyperbolic Learning of Instances and Classes
- CLIP-Powered Domain Generalization and Domain Adaptation: A Comprehensive Survey
- Enhancing Hyperbolic Graph Embeddings via Contrastive Learning
- Continual Multimodal Contrastive Learning
- Principled Multimodal Representation Learning
- Decoupled Weight Decay Regularization
- Fine-Grained Visual Classification of Aircraft
- The interplay of the polar decomposition theorem and the Lorentz group
- Progressive Neural Networks
- Hyperbolic Neural Networks++
- Hyperbolic Graph Neural Networks: A Review of Methods and Applications
- Recent Advances of Multimodal Continual Learning: A Comprehensive Survey
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks