FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning

arXiv:2605.09355 · cs.LG · Submitted 2026-05-10 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: I'm Tom, and with me are Jane, Lu, senior AI researcher at Tsinghua, Meng, lead engineer at a mysterious AI startup and Lalam, the in-house Large Language Model.

Jane: Today's paper: "FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning".

Tom: Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: multi-task pretraining, where related tasks borrow representational strength, and continual adaptation,

Jane: First, who's behind it and why it matters.

Title and authors: Tom: So, we’re looking at FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning, and the authors are Xing Han, Shravan Chaudhari, Tanvi Ranade from Johns Hopkins University. They’ve framed this as addressing the limitations where pretraining tasks can't be exhaustive while skipping joint training loses the gains we get from related tasks.

Jane: It seems like they’re proposing a solution that combines a sparse Mixture-of-Experts architecture with a mechanism for continual adaptation, which is pretty smart because it tackles those two separate problems in one framework.

Lu: The title itself points directly to the core innovations: the adaptive nature of the MoE and how it handles both multi-task pretraining and continual learning across multimodal settings.

Meng: I’m curious about how this architecture actually manages that "adaptive" part when a new modality combination enters the system; does it just throw a new expert at it, or is there a smarter way?

Lalam: The authors explain that sparse activation enables modular capacity expansion as new tasks arrive, and routing decouples modality-level computation from task-level composition, which is what makes it adaptable.

The paper's summary: Tom: They summarize FLAME by focusing on its two main goals: first, jointly training a unified model across Flexi-Modal multitasks with heterogeneous modality combinations, and second, handling arbitrary train-test modality combination shifts while keeping the continual learning capacity when new tasks with different compositions arrive.

Jane: To put that in simpler terms, they’re showing how you can train a single AI system to handle many related prediction goals using different kinds of inputs together from the start, and then letting it keep learning new, unexpected data combinations without forgetting what it already knows.

Lu: They achieve this by introducing a per-modality routing mechanism where each modality type gets its own router that sends sequences to a shared expert pool, which handles the decoupling of modality computation from task composition.

Meng: I see how that structure helps manage complexity, but the paper also mentions a specific way they handle temporal irregularities in the input data, which is important since real-world sensor data isn't always perfectly timed.

Lalam: The paper details that each expert processes its input using a length-preserving 1D convolution along the time axis with a kernel size κ to capture local dynamics, followed by a position-wise two-layer MLP applied at every step, which allows sequences of any length and any modality to traverse the same expert.

The paper's improvements: Tom: Now we’re looking at the actual improvements they propose, and one big thing is that they introduce an efficient continual learning scheme that combines spectral compression with lightweight router expansion to handle new tasks without starting from scratch.

Jane: That sounds like a smart way to save resources; instead of retraining everything for every new task, they compress the knowledge from previous tasks into low-rank memory subspaces using a specific process defined in Equation (three).

Lu: The insight here is that the functional energy of trained experts concentrates in a sharply low-rank subspace of its routed input distribution, which allows for structural no-forgetting guarantees through cursor-based inference.

Meng: From an engineering standpoint, that low-rank compression sounds promising because it means we can potentially store and retrieve knowledge much more efficiently than methods like EWC or LoRA, achieving parameter counts five to fifteen times smaller with a structural no-forgetting property.

Lalam: They enforce this structural no-forgetting through parameter isolation, freezing components reserved for prior tasks during training, and then using a cursor τ = k(t) pointing to the stage where task t was first trained to calculate the effective expert weight.

Conclusion: Tom: So, wrapping up FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning, it seems like they successfully addressed training unified models across heterogeneous tasks and handling shifts in modality composition by using their sparse MoE routing and the spectral compression technique for continual learning.

Jane: It really suggests that we can deploy multimodal AI systems in environments where new data streams keep appearing, as long as the underlying architecture is flexible enough to manage both pretraining on diverse data sets and subsequent adaptation without losing prior performance.

Lu: The paper highlights how this framework allows for flexible modality combinations and helps discover task pairings by analyzing the routing fingerprints, which opens up avenues for more structured AI collaboration between different tasks.

Meng: Practically speaking, this means we can deploy these complex models in real-world clinical settings with less operational overhead because the knowledge absorption is so efficient via that low-rank memory approach.

Lalam: I think the biggest cultural impact is on how we view model development; it shows that instead of a monolithic AI, we can build modular systems where new capabilities are added incrementally through routing and compression, which feels much more scalable for long-term research.

Department of Computer Science, Johns Hopkins University · Department of ECE, Johns Hopkins University

cs.LG

Submitted: 2026-05-10

Updated: 2026-09-28

Code: https://github.com/aaronhan223/FLAME

Importance score: 92/100

The gist: Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: multi-task pretraining, where related tasks borrow representational

Key concepts

Mixture-of-Experts (MoE)
A sparse architecture where a single router directs input sequences to a shared pool of specialized experts. This decouples task composition from modality computation, enabling flexible combinations of inputs and tasks.
Spectral Compression
A technique that compresses the weights of trained experts into low-rank memory subspaces. This captures the essential knowledge efficiently, allowing for fast inference with significantly fewer parameters than traditional fine-tuning methods.
Structural No-Forgetting Guarantee
A mechanism ensuring new tasks do not degrade performance on old tasks. It works by isolating expert weights reserved for prior knowledge and using a 'cursor' to select only relevant weights during inference, preventing interference.
Per-Modality Routing
Each modality type has its own dedicated router that processes inputs regardless of the task they belong to. This allows the framework to handle sequences from any modality structure while maintaining structural sharing among experts.

Terminology

Summary

Real-world model deployment across multiple domains requires multimodal models to operate under two complementary regimes: multi-task pretraining, where related tasks borrow representational strength, and continual adaptation, where new tasks emerge with previously unseen modality combinations. FLAME proposes a scalable Mixture-of-Experts (MoE) framework designed to simultaneously support joint multitask pretraining across heterogeneous modalities and continual learning over sequential multimodal tasks by leveraging per-modality routing and low-rank memory compression.

How it works

FLAME is structured around a sparse MoE architecture where modality-specific routers dispatch sequences from different tasks into a shared expert pool, decoupling modality-level computation from task composition. This design allows the framework to support flexible modality combinations by instantiating one router Gm per modality type, which processes embeddings from any task regardless of its specific input structure. The routing mechanism is performed at the sample level using a learnable query that attends over time, resulting in a fixed-size summary vector, denoted as the sample-level summary z¯m, whose size is independent of the number of modalities or tasks involved.

Each expert Ei processes a full single-modality sequence at its native temporal resolution by applying a length-preserving 1D convolution along the time axis with kernel size κ to capture local temporal dynamics, followed by a position-wise two-layer MLP applied identically at every step. This ensures that sequences of any length and any modality can traverse the same expert, providing structural inter-task sharing. The routing distribution divergence loss, L(k)div, is used to regulate cross-modal expert overlap between tasks within a single task, ensuring modalities specialize on disjoint experts or fuse representations when beneficial.

Continual Adaptation via Spectral Compression

To handle continual learning over new tasks without retraining from scratch, FLAME employs an efficient continual learning scheme that combines spectral compression with lightweight router expansion. The framework leverages the observation that the functional energy of trained experts concentrates in a sharply low-rank subspace despite near-full-rank weights. This motivates the approach where accumulated expert knowledge is compressed into low-rank memory subspaces using a process defined by Equation (3):

Wf(t)i = U(t)i Σ(t)i V(t)⊤i, W(0)i ← U(0)i, 1:r0 Σ(0)i, 1:r0 V(0)⊤i, 1:r0, Π j ← (Wf(t)) (Eq. 3).

This compression is applied to the expert weights and the variable-length attention layers inside each modality encoder. The key insight is that the functional energy of Wf(t)i concentrates in a low-rank subspace of its routed input distribution, which allows for a structural no-forgetting guarantee via cursor-based inference at 5–15× fewer parameters than fine-tuning methods like EWC or LoRA.

Structural No-Forgetting Guarantee

The framework achieves a structural no-forgetting property through two mechanisms: parameter isolation and cursor-based inference. Parameter isolation is enforced during training by freezing all components reserved for prior tasks, ensuring that the optimizer cannot modify blocks reserved for earlier knowledge. Inference is isolated by using a cursor τ = k(t) pointing to the stage at which task t was first trained; the effective expert weight is calculated as Weff i(τ) = P τj=0 W(j) i (Eq. 4). This construction ensures that "any W(j) i with j > k(t) is excluded from task t’s forward pass by construction," meaning new tasks cannot perturb the predictions of earlier tasks regardless of what they learn.

Task-Specific Router and Head Expansion

For continual learning, FLAME extends the per-modality router Gm into a task-indexed family of routers, G(t) m t. At stage t, a new lightweight router head G(t) m is added for every modality m in the set of tasks Mt at that stage, while earlier heads are frozen. A new task head h(t) is also added in parallel. This expansion ensures that the per-stage router overhead is O(Mt d N), negligible relative to the experts. The inference step then selects the appropriate G(t) m, the cursor τ = k(t) on every stack Π i, and the task head h(t) for prediction.

Key Contributions and Validation

FLAME addresses two core challenges: jointly training a unified model across Flexi-Modal multitasks with heterogeneous modality combinations and handling arbitrary train-test modality combination shift. The framework is validated on multiple healthcare multimodal benchmarks, demonstrating "competitive multitask pretraining performance while alleviating catastrophic forgetting and improving parameter efficiency.

Improvements for AI systems

Based on the scientific paper FLAME: Adaptive Mixture-of-Experts for Continual Multimodal Multi-Task Learning, here are specific improvements that can be implemented in AI systems, along with what those improved systems can achieve:


) The core improvement involves replacing monolithic models with a highly modular, sparse Mixture-of-Experts (MoE) architecture. This is specifically designed to handle the dual requirements of massive pretraining and continuous adaptation across diverse data streams.

  1. The system will employ a per-modality routing mechanism where each modality (e.g., clinical text, chest X-ray, time series) has its own dedicated router that dispatches tokens to a shared pool of experts.

  2. This architecture is augmented with an efficient continual learning scheme that utilizes low-rank memory subspaces for expert knowledge compression. Specifically, when a new task arrives, instead of retraining the entire model or fine-tuning all parameters, only the lightweight routers and new task heads are expanded/trained; accumulated knowledge from prior tasks is compressed into rank-limited additive slices of the fixed expert pool.

) The improved AI system can perform:

  1. Joint Multimodal Pretraining across Heterogeneous Tasks: The model can be trained simultaneously on a wide variety of related tasks (e.g., predicting Length-of-Stay, Mortality Rate, and Breast Disease Prediction) that rely on different combinations of modalities (e.g., vital signs and clinical notes). This allows the system to learn cross-modal representations that benefit all related tasks simultaneously, overcoming the limitations where existing models assume a fixed input structure.

  2. Continual Adaptation with Structural No-Forgetting: The model can be continuously deployed in dynamic environments where new prediction tasks emerge with previously unseen modality combinations (e.g., a new sensor modality or institutional protocol). Crucially, it retains performance on all previously learned tasks without catastrophic forgetting because the system uses cursor-based inference to selectively sum only the knowledge reserved for prior tasks.

  3. Efficient Parameter Scaling: The model can remain computationally affordable even as the number of tasks grows substantially (e.g., moving from 9 to dozens). This is achieved by reusing functionally idle capacity through spectral compression, allowing the system to absorb new task complexity with only a fixed-size expert pool and rank-limited updates, resulting in a parameter footprint that is 5–15 times smaller than traditional continual learning baselines (Simple FT, EWC, LoRA) while maintaining high retention.

) The improved AI system can specifically achieve:

  1. Enhanced Robustness to Modality Gaps: It can effectively bridge the gap between tasks that share some modalities but require others, by allowing the per-modality routers to selectively route tokens to shared experts or complementary experts, leading to more nuanced and accurate predictions than models that must use fixed cross-attention fusion methods.

  2. Data-Driven Task Pairing Recommendations: The system can provide interpretable insights into which task combinations benefit most from joint training by analyzing the routing fingerprints. This allows researchers to discover new high-value task pairings without manual trial-and-error, acting as a data-driven recommender for model collaboration.

  3. Scalable Knowledge Absorption: It can absorb an infinite stream of new clinical or sensor tasks in a single deployment without requiring expensive full retraining or validation cycles for every new requirement, drastically lowering the operational overhead of maintaining complex multimodal AI systems in real-world medical settings.

Sources

Related papers