M3TR: Temporal Retrieval Enhanced Multi-Modal Micro-video Popularity Prediction
Jiacheng Lu, Weijian Wang, Mingyuan Xiao, Yang Hua, Tao Song, Bo Peng, Cheng Hua, Haibing Guan
cs.MM, cs.AI
Submitted: 2026-08-18
Updated: 2026-08-19
Terminology
Sources
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale
- The Faiss library
- Transformer Feed-Forward Layers Are Key-Value Memories
- Contextual LSTM (CLSTM) models for Large scale NLP tasks
- AST: Audio Spectrogram Transformer
- Multi-modal and Metadata Capture Model for Micro Video Popularity Prediction
- A Content-Driven Micro-Video Recommendation Dataset at Scale
Related papers
- ControlFoley: Unified and Controllable Video-to-Audio Generation with Cross-Modal Conflict Handling
- Zero-shot Video Moment Retrieval via Off-the-shelf Multimodal Large Language Models
- Mitigating GenAI-powered Evidence Pollution for Out-of-Context Multimodal Misinformation Detection
- A Rate-Distortion-Classification Approach for Lossy Image Compression