WaveTLM: Reliable Time-Series Language Modeling through Task Compilation
cs.LG
Submitted: 2026-09-16
Updated: 2026-09-16
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Time-series language models provide a shared natural-language interface across temporal tasks, but plausible text does not guarantee reliable task outputs.
Terminology
Abstract
Time-series language models provide a shared natural-language interface across temporal tasks, but plausible text does not guarantee reliable task outputs. Responses may appear reasonable while hallucinating the required object: numerical sequences can violate shape, scale, channel order, or temporal alignment, and textual decisions can fall outside the legal label space. We formulate reliable time-series language modeling, separating task-object reliability from predictive quality. We introduce ExecTS-QA, a contract-grounded benchmark spanning forecasting, imputation, classification, anomaly detection, and waveform analysis. We further propose WaveTLM, a unified compiler-executor model whose task compiler transforms user requests, visible arguments, and wave-grounded evidence into typed task states, while task-native executors construct numerical tensors, legal decisions, or structured records. On ExecTS-QA, a single WaveTLM checkpoint achieves 99.40% contract-valid coverage, compared with 37.83% for the strongest evaluated string-first baseline, while retaining balanced predictive performance across all five task families. Evaluations on SciTS, TSQA, IRTS-ToolBench, and ARFBench provide additional evidence of transfer. The code, construction scripts, and ExecTS-QA dataset will be publicly released upon publication. These results show that task compilation can convert plausible language generation into reliable time-series outputs.
Sources
- An Empirical Evaluation of Generic Convolutional and Recurrent Networks for Sequence Modeling
- TimeGPT-1
- Large Language Models Are Zero-Shot Time Series Forecasters
- TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language Models
- Time-LLM: Time Series Forecasting by Reprogramming Large Language Models
- iTransformer: Inverted Transformers Are Effective for Time Series Forecasting
- AutoTimes: Autoregressive Time Series Forecasters via Large Language Models
- Lag-Llama: Towards Foundation Models for Probabilistic Time Series Forecasting
- TEST: Text Prototype Aligned Embedding to Activate LLM's Ability for Time Series
- SciTS: Scientific Time Series Understanding and Generation with LLMs
- TimeFound: A Foundation Model for Time Series Forecasting
- ARFBench: Benchmarking Time Series Question Answering Ability for Software Incident Response
- TimeMaster: Training Time-Series Multimodal LLMs to Reason via Reinforcement Learning
- Insight Miner: A Time Series Analysis Dataset for Cross-Domain Alignment with Natural Language
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks