Unlocking Pretrained Vision Transformers for Time Series Classification
cs.LG, cs.AI, cs.CV
Submitted: 2025-06-10
Updated: 2026-09-21
Comments: GCPR 2026 Oral
Code: https://github.com/ExplainableML/TiViT
License: http://creativecommons.org/licenses/by/4.0/
The gist: Adapting vision models for time series analysis is compelling, yet all existing approaches are falling short of dedicated time series foundation models (TSFMs) in classification.
Terminology
Abstract
Adapting vision models for time series analysis is compelling, yet all existing approaches are falling short of dedicated time series foundation models (TSFMs) in classification. In this work, we propose Time Vision Transformer (TiViT), the first framework that successfully unlocks the representational power of frozen Vision Transformers (ViTs) pretrained on large-scale image datasets for time series classification. TiViT achieves state-of-the-art performance without any finetuning by utilizing the hidden representations of OpenCLIP models. We explore the structure of TiViT representations and find that intermediate ViT layers with high intrinsic dimension are the most effective for time series classification. Furthermore, we assess the alignment between TiViT and TSFM representation spaces and identify a strong complementarity, with additional performance gains achieved through feature concatenation. Finally, we unfreeze the ViT backbone of TiViT for continual pretraining and contrastive alignment with TSFMs on time series, enhancing the performance of lightweight TiViT variants. Our findings reveal a new direction for the domain and task adaptation of vision foundation models. Code is available at https://github.com/ExplainableML/TiViT.
Sources
- The UEA multivariate time series classification archive, 2018
- Mamba4Cast: Efficient Zero-Shot Time Series Forecasting with State Space Models
- VisionTS: Visual Masked Autoencoders Are Free-Lunch Zero-Shot Time Series Forecasters
- From Colors to Classes: Emergence of Concepts in Vision Transformers
- Mantis: Lightweight Foundation Model for Time Series Classification
- NuTime: Numerically Multi-Scaled Embedding for Large-Scale Time-Series Pretraining
- Moirai-MoE: Empowering Time Series Foundation Models with Sparse Mixture of Experts
- Harnessing Vision Models for Time Series Analysis: A Survey
- Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting
- SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
- Towards a General Time Series Forecasting Model with Unified Representation and Adaptive Transfer
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks