CodeTS: Verifiable Text-to-Time Series Generation via Executable Code
cs.LG, cs.AI
Submitted: 2026-09-14
Updated: 2026-09-14
Comments: Preprint
Code: https://github.com/alibaba/clusterdatahttps:
License: http://creativecommons.org/licenses/by/4.0/
The gist: Text-to-Time Series Generation (Text-to-TS) provides a promising paradigm for synthesizing time series from natural language, enabling scenario-specific generation when real observations are scarce
Terminology
Abstract
Text-to-Time Series Generation (Text-to-TS) provides a promising paradigm for synthesizing time series from natural language, enabling scenario-specific generation when real observations are scarce or costly to acquire. However, existing methods typically lack an explicit mechanism for deriving generation logic from textual descriptions to guide time series synthesis. In this paper, we propose CodeTS, a verifiable framework that uses code as an intermediate generation interface, reformulating Text-to-TS generation as a Text-to-Code-to-TS process. CodeTS first maps textual temporal descriptions into an explicit code space, where executable code specifies how textual requirements shape target temporal patterns, and then obtains the time series through code execution. To learn this code generation process reliably without real code annotations, CodeTS constructs aligned Text-Code-TS triplets from structured temporal attributes for supervised initialization. More importantly, we further design multi-stage execution-based rewards that verify format validity, code executability, and time series quality, enabling real Text-TS pairs to provide training signals for Reinforcement Learning with Verifiable Rewards (RLVR). Extensive experiments on eight benchmarks across short, medium, and long generation lengths demonstrate that CodeTS provides a strong zero-shot solution for Text-to-TS generation, outperforming LLM-based baselines and achieving better averaged results than supervised generative baselines trained on the target datasets.
Sources
- Chronos: Learning the Language of Time Series
- Evaluating Large Language Models Trained on Code
- TimeOmni-1: Incentivizing Complex Reasoning with Time Series in Large Language Models
- Qwen2.5-Coder Technical Report
- GPT-4o System Card
- Tulu 3: Pushing Frontiers in Open Language Model Post-Training
- BRIDGE: Bootstrapping Text to Control Time-Series Generation via Multi-Agent Iterative Optimization and Diffusion Modeling
- Thoth: Mid-Training Bridges LLMs to Time Series Understanding
- Conditional Sig-Wasserstein GANs for Time Series Generation
- Devstral: Fine-tuning Language Models for Coding Agent Applications
- Seed-Coder: Let the Code Model Curate Data for Itself
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- Qwen3.5-Omni Technical Report
- Transformers in Time Series: A Survey
- SciTS: Scientific Time Series Understanding and Generation with LLMs
- IQuest-Coder-V1 Technical Report
- Spectral-Aware Text-to-Time Series Generation with Billion-Scale Multimodal Meteorological Data
Related papers
- Polynomial-Augmented Neural Networks (PANNs) with Weak Orthogonality Constraints for Enhanced Function and PDE Approximation
- AIRL-S: Unifying Reinforcement Learning and Search-Based Test-Time Scaling via Adversarial Inverse Reinforcement Learning
- Transformers as Bayesian In-Context Experimenters: Smoothness-Adaptive Efficient ATE Estimation
- Convergence issues in Relational Concept Analysis based on AOC-posets
- Beliefs Beyond Posteriors: Local-Consistency Optimisation for Bayesian Neural Networks
- Understanding Diffusion Models via Ratio-Based Function Approximation with SignReLU Networks