Distributed JEPA: A Self-Supervised Framework for Energy Forecasting
cs.LG, cs.AI
Submitted: 2026-09-15
Updated: 2026-09-15
License: http://creativecommons.org/licenses/by/4.0/
The gist: Traditional energy forecasting solutions rely on task-specific supervision and energy asset representations, limiting transferability and the ability to capture general temporal dynamics across
Terminology
Abstract
Traditional energy forecasting solutions rely on task-specific supervision and energy asset representations, limiting transferability and the ability to capture general temporal dynamics across heterogeneous assets. We address this by proposing a distributed Joint Embedding Predictive Architecture (JEPA) for self-supervised learning from heterogeneous energy time-series. The framework predicts latent representations of masked temporal segments while integrating temporal observations and contextual information within a shared embedding space. To prevent representation collapse, training combines a latent-space predictive objective with covariance and temporal variance regularization. The evaluation was conducted on energy consumption and generation datasets under data-degradation scenarios and compared with a Transformer forecasting baseline. The learned representations remained stable (cosine similarity about 0.98; effective rank 185-235). JEPA achieved performance comparable to a Transformer on building energy data, higher R squared in 3/5 consumer clusters, and outperformed the baseline on 9/10 unseen PVs (R squared =0.73-0.88 vs. <0.45), while showing greater robustness to missing data.
Sources
- Autoformer: Decomposition Transformers with Auto-Correlation for Long-Term Series Forecasting
- A Time Series is Worth 64 Words: Long-term Forecasting with Transformers
- iTransformer: Inverted Transformers Are Effective for Time Series Forecasting
- TimesNet: Temporal 2D-Variation Modeling for General Time Series Analysis
- Joint Embeddings Go Temporal
- Unsupervised Scalable Representation Learning for Multivariate Time Series
- Unsupervised Representation Learning for Time Series with Temporal Neighborhood Coding
- CoST: Contrastive Learning of Disentangled Seasonal-Trend Representations for Time Series Forecasting
- Self-Supervised Contrastive Pre-Training For Time Series via Time-Frequency Consistency
- How JEPA Avoids Noisy Features: The Implicit Bias of Deep Linear Self Distillation Networks
- V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning
- A-JEPA: Joint-Embedding Predictive Architecture Can Listen
- JEPA as a Neural Tokenizer: Learning Robust Speech Representations with Density Adaptive Attention
- LLM-JEPA: Large Language Models Meet Joint Embedding Predictive Architectures
- SC-JEPA: Stabilizing Latent Predictive Learning for Time-Series Anomaly Prediction
- LeJEPA: Provable and Scalable Self-Supervised Learning Without the Heuristics
- Connecting Joint-Embedding Predictive Architecture with Contrastive Self-supervised Learning
- Attention Is All You Need
- RankMe: Assessing the downstream performance of pretrained self-supervised representations by their rank
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the Hypersphere
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks