Fidel-TS: A High-Fidelity Multimodal Benchmark for Time Series Forecasting
cs.LG, stat.ML
Submitted: 2025-09-29
Updated: 2026-09-18
Comments: new version
Code: https://github.com/VEWOXIC/Universal-Cross-Modal-Time-Series-ForecastingPipeline
License: http://creativecommons.org/licenses/by/4.0/
The gist: The evaluation of time series forecasting models is hindered by a lack of high-quality benchmarks, leading to overestimated assessments of progress.
Terminology
Abstract
The evaluation of time series forecasting models is hindered by a lack of high-quality benchmarks, leading to overestimated assessments of progress. Existing datasets suffer from issues ranging from small-scale, low-frequency, pre-training data contamination in unimodal designs to the temporal and description leakage prevalent in early multimodal designs. To address this, we formalize the core principles of high-fidelity benchmarking, focusing on data sourcing integrity, leak-free design, and structural clarity. We introduce Fidel-TS, a new large-scale benchmark built from these principles. Our experiments reveal the limitations of prior benchmarks and the potential discrepancies in model evaluation, providing new insights into multiple existing unimodal and multimodal forecasting models and LLMs across various evaluation tasks.
Sources
- DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
- TimeGPT-1
- jina-embeddings-v3: Multilingual Embeddings With Task LoRA
- Llama 2: Open Foundation and Fine-Tuned Chat Models
- Deep Time Series Models: A Comprehensive Survey and Benchmark
- Context is Key: A Benchmark for Forecasting with Essential Textual Information
- Intervention-Aware Forecasting: Breaking Historical Limits from a System Perspective
- Qwen3 Technical Report
- Qwen2.5 Technical Report
- Qwen2.5-1M Technical Report
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks