TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation
Pengyu Zhang, Yangqin Jiang, Klim Zaporojets, Congfeng Cao, Paul Groth
University of Amsterdam · University of Hong Kong · Aarhus University
cs.IR, cs.AI
Submitted: 2026-08-24
Updated: 2026-08-25
Code: https://github.com/HKUDS/DiffMM
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 95/100
The gist: TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation Summary This paper introduces TimeRoute, a diffusion-based multi-modal recommender system designed to address the
Terminology
Summary
TimeRoute: Time-Aware Modality Routing and Diffusion for Multi-Modal Recommendation
Summary
This paper introduces TimeRoute, a diffusion-based multi-modal recommender system designed to address the problem of modality time-scale mismatch,
where the relevance of different item modalities (e.g., text, images, audio) shifts over time at different rates for the same user. The authors identify two coupled challenges arising from this mismatch: (1) users require different modality fusion proportions in different temporal contexts, and (2) modalities that are less relevant in the current context are more likely to introduce outdated or misleading signals into the recommender.
To address these challenges, TimeRoute consists of two complementary components:
-
Temporal-Aware Modal Router: This component addresses the first problem by generating personalized modality fusion weights for each user based on their temporal behavior profile. The router constructs a 16-dimensional temporal profile from each user's interaction timestamps, including fields such as interaction recency, calendar context, inter-event gaps, and weekday/hour patterns. A two-layer MLP with softmax output maps this profile to a per-user modality distribution. The router includes a weight floor (ε = 0.05) to prevent global modality collapse and a diversity regularizer to encourage meaningful per-user differentiation. User-side fusion uses these personalized weights, while item-side fusion retains globally learned modality weights due to the lack of item-level temporal context.
-
Time-Conditioned Diffusion Reconstructor: This component addresses the second problem by conditioning the diffusion-based graph reconstruction on temporal signals. The denoiser is modulated through Feature-wise Linear Modulation (FiLM) with dual-stream long- and short-term heads. The long-term stream captures slowly evolving signals (e.g., long-term usage history, seasonal patterns), while the short-term stream models rapidly changing interaction dynamics (e.g., inter-event timing, time-of-day effects). Each stream produces scale and shift parameters that modulate the user context embedding, and a per-user gate combines the two reconstructions. Additionally, a time-reweighted reconstruction loss biases the denoiser to prioritize accurate reconstruction near the temporal frontier, and a modality-aware alignment regularizer ties reconstructed user-item scores to the modality feature space.
The model is trained with a three-phase per-epoch schedule: (1) diffusion training, (2) graph reconstruction, and (3) GCN and recommendation training. The joint loss combines BPR ranking loss, cross-modal contrastive learning, per-modality diffusion losses, alignment regularization, diversity regularization, and weight regularization.
Experiments on TikTok, Amazon-Baby, and Amazon-Sports with 10-seed paired tests show consistent improvements over strong baselines, with gains up to 9.8% in Recall@20, Precision@20, and NDCG@20. Under a chronological split, TimeRoute shows particularly large gains on NDCG@20 (+16.11% on TikTok, +11.02% on Amazon-Baby, +12.59% on Amazon-Sports). Controlled ablations confirm that the gains come from temporal signals rather than extra router parameters (noise-input router performs statistically indistinguishably from no-router baseline, p = 0.50). Quartile analysis shows systematic per-user routing differentiation, with more active users exhibiting more confident, image-dominant routing. Component ablations show that all four temporal mechanisms (modal router, dual-stream denoising, FiLM conditioning, time-weighted loss) contribute independently, with relative drops ranging from −2.19% to −4.06% when any single component is removed. Noise-robustness experiments demonstrate that temporal routing preserves its benefits under uneven modality corruption, with the performance gap between TimeRoute and a no-temporal variant remaining stable in the range 0.0045–0.0053 across clean, image-noisy, text-noisy, and mixed-noise settings.
The paper concludes that TimeRoute effectively addresses modality time-scale mismatch through complementary fusion-level and reconstruction-level temporal awareness, though it notes limitations including dependence on timestamp coverage and potential convergence toward near-single-modality weights for highly active users on datasets with a dominant modality.
Improvements for AI systems
Improvements to AI Systems:
-
Dynamic Modality Weighting via Temporal User Profiling: Integrate a temporal-aware router that constructs a compact 16-dimensional profile from interaction timestamps (recency, calendar context, inter-event gaps, weekday/hour patterns) and maps it via a lightweight MLP to per-user modality fusion weights. This allows the AI to automatically adjust how much it relies on text, images, or audio for each user at each moment, preventing stale or irrelevant modalities from dominating.
-
Time-Conditioned Denoising for Sequential Reconstruction: Add a dual-stream (long-term and short-term) FiLM-modulated denoiser to any diffusion-based generative or recommender system. The long-term stream captures slow trends (e.g., seasonal shifts), while the short-term stream models rapid changes (e.g., time-of-day effects). A per-user gate blends these streams, enabling the AI to reconstruct user-item relationships that are accurate at the current temporal frontier rather than averaged over all history.
-
Time-Reweighted Loss for Recency Bias: Modify the reconstruction loss to assign higher weight to errors near the most recent interaction timestamps. This forces the AI to prioritize predicting the user’s immediate next behavior, improving performance on chronological splits and real-time recommendation tasks.
-
Modality-Aware Alignment Regularizer: Add a regularizer that ties the reconstructed user-item scores back to the modality feature space (e.g., aligning score gradients with text/image embeddings). This prevents the diffusion process from drifting into purely latent, uninterpretable spaces and keeps the model grounded in observable item attributes.
-
Robustness to Noisy Modalities: Use temporal routing to maintain performance when some modalities are corrupted or missing (e.g., broken images or garbled text). The system automatically down-weights unreliable modalities based on temporal context, preserving recommendation quality under uneven data degradation.
What the Improved AI System Can Do:
-
Personalized Real-Time Recommendations: For each user, it can instantly rebalance which item features (visual, textual, acoustic) matter most based on their recent activity patterns—e.g., favoring images for a user browsing at night but text for the same user during work hours.
-
Accurate Next-Item Prediction: It can predict the user’s very next interaction with higher precision (up to +16% NDCG@20) by prioritizing recent signals over long-term averages, making it suitable for streaming, e-commerce, or social media feeds.
-
Graceful Degradation Under Data Corruption: If an item’s image is corrupted or its text is truncated, the system can still recommend effectively by shifting reliance to other modalities, without retraining.
-
Explainable Modality Choices: The router’s per-user weights provide interpretable insight into why a recommendation was made (e.g., “this user currently prefers visual content”), aiding transparency and user control.
-
Scalable Adaptation to New Users: With only a few timestamps, the system can construct a temporal profile and immediately generate sensible modality weights, reducing cold-start issues compared to static fusion models.
Sources
- Time Series Analysis in Frequency Domain: A Survey of Open Challenges, Opportunities and Benchmarks
- Retrieval and Distill: A Temporal Data Shift-Free Paradigm for Online Recommendation System
- A Probabilistic Framework for Temporal Distribution Generalization in Industry-Scale Recommender Systems
Related papers
- The Price of Isolation: Estimating the Ecosystem Cost of Symmetric Two-Sided A/B Testing
- SCAR: Semantic Continuity-Aware Retrieval for Efficient Context Expansion in RAG
- MixLoRA-DSI: Dynamically Expandable Mixture-of-LoRA Experts for Rehearsal-Free Generative Retrieval over Dynamic Corpora
- RRCM: Ranking-Driven Retrieval over Collaborative and Meta Memories for LLM Recommendation
- Right Family, Wrong Skill: Evaluating Risk Exposure in Agent Skill Retrieval
- UltRAG: a Universal Simple Scalable Recipe for Knowledge Graph RAG