Empowering VLMs for Few-Shot Multimodal Time Series Classification via Tailored Agentic Reasoning
cs.AI, cs.LG, cs.MA, cs.MM
Submitted: 2026-05-10
Updated: 2026-09-05
Comments: 17 pages, 12 figures, 8 tables. Accepted by ACM MM 2026
Code: https://github.com/HuangJW0821/MarsTSC
License: http://creativecommons.org/licenses/by/4.0/
The gist: In this paper, we propose the first VL gentic easoning framework for few- hot multimodal ime eries lassification (MarsTSC), which introduces a self-evolving knowledge bank as a dynamic context
Terminology
Abstract
In this paper, we propose the first VL gentic easoning framework for few- hot multimodal ime eries lassification (MarsTSC), which introduces a self-evolving knowledge bank as a dynamic context iteratively refined via reflective agentic reasoning. The framework comprises three collaborative roles: i) Generator conducts reliable classification via reasoning; ii) Reflector diagnoses the root causes of reasoning errors to yield discriminative insights targeting the temporal features overlooked by Generator; iii) Modifier applies verified updates to the knowledge bank to prevent context collapse. We further introduce a test-time update strategy to enable cautious, continuous knowledge bank refinement to mitigate few-shot bias and distribution shift. Extensive experiments across 12 mainstream time series benchmark datasets demonstrate that delivers substantial and consistent performance gains across 5 VLM backbones, outperforming both classical and foundation model-based time series baselines under few-shot conditions, while producing interpretable rationales that ground each classification decision in human-readable feature evidence. Code is available at https://github.com/HuangJW0821/MarsTSC.
Sources
- Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts
- Can Time-Series Foundation Models Perform Building Energy Management Tasks?
- Harnessing Vision Models for Time Series Analysis: A Survey
- A Time Series is Worth 64 Words: Long-term Forecasting with Transformers
- Language Model Self-improvement by Reinforcement Learning Contemplation
- TRACE: Contrastive learning for multi-trial time-series data in neuroscience
- Multi-Modal View Enhanced Large Vision Models for Long-Term Time Series Forecasting
- A-MEM: Agentic Memory for LLM Agents
- A Closer Look at Bearing Fault Classification Approaches
- Structured Agentic Workflows for Financial Time-Series Modeling with LLMs and Reflective Feedback
- Chronos: Learning the Language of Time Series
- The UEA multivariate time series classification archive, 2018
- Reflect, Retry, Reward: Self-Improving LLMs via Reinforcement Learning
- Explainable-AI powered stock price prediction using time series transformers: A Case Study on BIST100
- Plots Unlock Time-Series Understanding in Multimodal Models
- TimeXL: Explainable Multi-modal Time Series Prediction with LLM-in-the-Loop
- Enhanced Fault Detection and Cause Identification Using Integrated Attention Mechanism
- Qwen3 Technical Report
- HeartLLM: Discretized ECG Tokenization for LLM-Based Diagnostic Reasoning
- Recurrent Neural Network Regularization
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection