FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data

arXiv:2608.10857 · cs.LG · Submitted 2026-08-11 · Read on arXiv

Viktoria Schuster, Sana Tonekaboni, Caroline Uhler

Massachusetts Institute of Technology · Broad Institute of MIT and Harvard · Technical University of Denmark · Vector Institute

cs.LG

Submitted: 2026-08-11

Updated: 2026-08-12

Code: https://github.com/viktoriaschuster/FiGuRO

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data Viktoria Schuster, Sana Tonekaboni, Caroline Uhler Summary This paper introduces Fidelity-Guided Rank Optimization (FiGuRO), a novel

Terminology

Summary

FiGuRO: Intrinsic Dimension Estimation for Multi-Modal Data

Viktoria Schuster, Sana Tonekaboni, Caroline Uhler

Summary

This paper introduces Fidelity-Guided Rank Optimization (FiGuRO), a novel framework for estimating the intrinsic dimension (ID) of uni- and multi-modal data. The ID is defined as the minimum number of variables needed to describe the data without significant information loss. The authors address a critical gap in existing techniques, which are often static, uni-modal, or in the case of contrastive methods, adapt only to the shared ID implicitly. FiGuRO is the first method to achieve multi-modal ID estimation and information decoupling in a single training pass, explicitly estimating the IDs of both shared and modality-specific (private) subspaces.

The method combines two powerful principles: (1) learning latent spaces via low-rank decomposable layers using truncated Singular Value Decomposition (SVD), inspired by adaptive rank reduction (ARR) and LoRA, and (2) an algorithm based on Rate-Distortion Theory that uses relative reconstruction fidelities to decide when to increase or decrease subspace ranks. The core of FiGuRO is an algorithm that optimizes the dimension of latent subspaces under a user-defined acceptable level of lossy compression. The authors frame this as a greedy algorithm for finding an efficient operating point on the rate-distortion curve, minimizing reconstruction loss given model parameters and rank under a distortion budget λ.

The architecture for multi-modal data decomposes uni-modal embeddings into shared (zs) and private (zm) representations via pruned weight matrices. The authors provide theoretical guarantees, including convergence in finite time (Theorem A.1.6) and that the returned rank approximates the true ID (Theorem A.1.7), under assumptions of sufficient model capacity and a suitable fidelity proxy. They also prove that disentanglement emerges as the most rank-efficient solution, as a non-disentangled allocation would require redundant encoding of shared information, and an all-shared solution is optimizationally unstable.

In experiments, FiGuRO demonstrated robustness across a comprehensive hyperparameter sweep (N=1080 runs), arriving at an average estimate of 4.89 ± 0.01 for a true ID of 5. It proved robust to diverse data characteristics, though high dropout or noise could lead to overestimation. In uni-modal ID estimation benchmarks on complex simulated data, FiGuRO was the only framework to provide consistent, close-to-accurate estimates across both noisy and highly non-linear regimes, outperforming classical methods (PCA, MLE, TwoNN) and neural network approaches (ARD-VAE, RRA, FLIPD).

For multi-modal estimation, FiGuRO showed a strong ability to recover the underlying dimensional structure, with an average deviation from ground truth of 2.9 (SEM 0.7) across four simulated datasets, compared to 9.5-36.3 for baselines like JIVE, AJIVE, SLIDE, and ShIndICA. The method also successfully disentangled shared and private information, with predictability per label highest in its corresponding subspace. An ablation study confirmed that a naive rank-reduction-only approach collapsed to minimum ranks, demonstrating the necessity of the fidelity-guided algorithm.

On real-world datasets, FiGuRO was applied to Audio MNIST, So2Sat, and NYU Depth V2. It reduced total combined ranks by factors of 5-36 with only marginal increases in test loss. ID estimates aligned with domain expectations, such as allocating more dimensions to image modalities than radar or depth. For MNIST, FiGuRO converged to an average rank of 11.4 ± 0.6, fitting into the previously reported range of 7-25. The decomposed subspaces proved effective for downstream tasks, with shared subspaces achieving higher classification accuracies than uni-modal embeddings (e.g., 0.94 for Audio MNIST digit recognition). On NYU Depth V2, FiGuRO demonstrated its potential as a scalable latent probe for frozen pretrained models, decoupling the decomposition mechanism from feature extraction.

The authors conclude that FiGuRO provides a critical missing piece to multi-modal learning, offering a principled method to quantify informational complexity. They note that disentanglement occurs without the need for auxiliary orthogonality or sparsity constraints, and highlight the framework's potential for continual learning scenarios. They recommend using FiGuRO to estimate bounds under low and high distortion budgets for robust results.

Improvements for AI systems

Improvements to AI Systems:

  1. Adaptive Latent Dimensionality in Multi-Modal Models
  • Replace fixed latent dimensions in VAEs, autoencoders, or multi-modal fusion networks with FiGuRO’s rank-optimization loop. The system automatically grows or shrinks per-modality and shared subspaces during training, based on reconstruction fidelity under a user-set distortion budget.

  • Result: No manual tuning of latent size; models avoid over- or under-parameterization, reducing memory and compute while preserving information.

  1. Explicit Shared/Private Subspace Disentanglement Without Auxiliary Losses
  • Integrate FiGuRO’s SVD-based pruned weight matrices into encoders to force separation of shared (cross-modal) and private (modality-specific) representations. No need for orthogonality penalties or adversarial training—disentanglement emerges from rank-efficiency.

  • Result: Downstream tasks (e.g., classification, retrieval) can use only the shared subspace for cross-modal transfer, or only private subspaces for modality-specific features, with measurable gains in accuracy and interpretability.

  1. Rate-Distortion-Guided Model Compression
  • Use FiGuRO’s fidelity-guided rank reduction as a principled compression technique for pretrained models (e.g., frozen transformers or CNNs). The system probes each layer’s intrinsic dimension and prunes low-rank components until the reconstruction loss exceeds a predefined threshold.

  • Result: Up to 36× reduction in combined rank with minimal test loss increase, enabling deployment on edge devices without retraining.

  1. Continual Learning via Dynamic Subspace Expansion
  • Apply FiGuRO’s rank-increase mechanism when new tasks introduce novel information. The system detects when the current latent capacity is insufficient (via rising reconstruction loss) and allocates new private dimensions, while preserving existing shared dimensions.

  • Result: Avoids catastrophic forgetting; the model expands its representational capacity only when needed, maintaining a compact and interpretable latent structure over time.

  1. Robust Intrinsic Dimension Estimation for High-Dimensional, Noisy Data
  • Use FiGuRO as a preprocessing step to estimate the true ID of any dataset (e.g., genomic, imaging, sensor data) before applying downstream algorithms like clustering, visualization, or anomaly detection. It outperforms PCA/MLE/TwoNN in non-linear and noisy regimes.

  • Result: More accurate dimensionality reduction and better generalization in downstream tasks, especially when data is corrupted or highly non-linear.

  1. Latent Probe for Frozen Pretrained Models
  • Attach FiGuRO’s rank-optimizing decoder to any frozen feature extractor (e.g., CLIP, BERT) to estimate the intrinsic dimensionality of its embeddings for a given dataset. This reveals how much redundancy exists and which dimensions are shared across modalities.

  • Result: Enables principled model selection, feature pruning, and cross-modal alignment without modifying the pretrained weights.

  1. Multi-Modal Fusion with Balanced Information Allocation
  • In fusion models (e.g., audio-visual speech recognition, multimodal sentiment), use FiGuRO to automatically allocate rank to each modality based on its information content relative to the task. The shared subspace is sized to capture only task-relevant cross-modal redundancy, preventing one modality from dominating.

  • Result: Improved robustness to missing or noisy modalities, and better performance on tasks where modalities have unequal information value.

  1. Hyperparameter-Free Dimensionality Selection in Generative Models
  • For GANs or diffusion models, replace fixed latent codes with FiGuRO’s adaptive rank mechanism. The generator learns to use only the necessary dimensions to produce realistic samples, reducing mode collapse and improving sample diversity.

  • Result: More stable training and higher-quality generations, with the latent space automatically matching the true data manifold.

Abstract

Determining the complexity, or Intrinsic Dimension (ID), of data is fundamental to efficient and interpretable representation learning. This is particularly challenging in multi-modal settings when trying to learn disentangled representations for shared and private information. Existing techniques leave a critical gap: they are often static, uni-modal, or in the case of contrastive methods, adapt only to the shared ID implicitly. We introduce Fidelity-Guided Rank Optimization (FiGuRO), a framework for approximating the ID of uni- and multi-modal data under constraints of model capacity and hyperparameters. FiGuRO learns the dimensions of low-rank projections using truncated singular value decomposition and an algorithm that determines when to reduce or increase dimension and in which latent space. Disentanglement of shared and private information arises as an emergent property of this optimization, eliminating the need for complex auxiliary loss functions. We demonstrate that FiGuRO outperforms existing ID estimation techniques and is more robust to hyperparameter changes. Across simulations and real-world data, FiGuRO captures distinct ID scales and varying subspace ratios, and decomposes shared and private information successfully. Furthermore, we show that FiGuRO can be applied to modern uni-modal pretrained models, enabling efficient, post-hoc disentanglement of multi-modal representations.

Related papers