OpenFinGym: A Verifiable Multi-Task Gym Environment for Evaluating Quant Agents
cs.AI, cs.LG
Submitted: 2026-06-24
Updated: 2026-09-10
Comments: EMNLP 2026 Findings
License: http://creativecommons.org/licenses/by/4.0/
The gist: Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tasks, while the financial relevance of benchmark
Terminology
Abstract
Although large language model agents are increasingly applied to quantitative-finance workflows, their evaluation remains fragmented across isolated tasks, while the financial relevance of benchmark tasks is often overlooked. Yet financial workflows are inherently multi-stage, spanning interdependent tasks such as forecasting, strategy construction, risk management, and trading. Existing platforms typically focus on a single task, and can therefore overstate agent competence and fail to reveal weaknesses in generalization, real-market interaction, and financially meaningful decision-making. We introduce OpenFinGym, a unified gym environment for quantitative-finance agent development that covers forecasting, market generation, real-time trading, and fraud detection under a single execution and verification interface. OpenFinGym additionally provides an automated task-construction pipeline that turns quantitative finance publications into executable task packages; a containerised runtime with a host-side verifier service that supports scalable agent rollouts and prevents runtime train-test leakage; a paper trading engine with a low-latency data-stream design; deferred-resolution support for long-horizon and event-market forecasts; and integration for SFT and RL post-training
Sources
- MOSAIC: Modular Orchestration for Structured Agentic Intelligence and Composition
- Look-Ahead-Bench: a Standardized Benchmark of Look-ahead Bias in Point-in-Time LLMs for Finance
- TimeSeriesGym: A Scalable Benchmark for (Time Series) Machine Learning Engineering Agents
- StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
- Gated recurrent neural network with TPE Bayesian optimization for enhancing stock index prediction accuracy
- Real-valued (Medical) Time Series Generation with Recurrent Conditional GANs
- Deep Generative Models for Synthetic Financial Data: Applications to Portfolio and Risk Modeling
- Profit Mirage: Revisiting Information Leakage in LLM-based Financial Agents
- FinanceBench: A New Benchmark for Financial Question Answering
- DSGym: A Holistic Framework for Evaluating and Training Data Science Agents
- When Agents Trade: Live Multi-Market Trading Benchmark for LLM Agents
- Re(Visiting) Time Series Foundation Models in Finance
- Dynamic graph neural networks for enhanced volatility prediction in financial markets
- DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
- LiveTradeBench: Seeking Real-World Alpha with Large Language Models
- PIXIU: A Large Language Model, Instruction Data and Evaluation Benchmark for Finance
- Benchmarking Benchmark Leakage in Large Language Models
- Qwen3 Technical Report
- Qlib: An AI-oriented Quantitative Investment Platform
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection