Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems
cs.AI, cs.CE, q-fin.CP, q-fin.TR
Submitted: 2026-06-06
Updated: 2026-09-22
License: http://creativecommons.org/licenses/by-nc-sa/4.0/
The gist: Large language models (LLMs) and agentic systems are increasingly proposed for financial trading, yet their reported performance remains difficult to compare because studies vary in data provenance,
Terminology
Abstract
Large language models (LLMs) and agentic systems are increasingly proposed for financial trading, yet their reported performance remains difficult to compare because studies vary in data provenance, temporal split discipline, execution timing, turnover treatment, and transaction-cost modeling. This article presents a targeted topical review and reproducibility audit of execution realism in LLM-based trading research. A coded evidence matrix covering 30 trade-relevant primary studies is used to assess point-in-time controls, split transparency, held-out evaluation, cost and turnover treatment, execution semantics, universe definition, and artifact release. Across the audited sample, architecture reporting is generally clearer than the evaluation assumptions needed to judge whether a trading result is economically interpretable or reproducible. A 10-equity worked example is included only as a methodological scaffold to illustrate how explicit friction and timing choices can materially compress active-strategy results. The main conclusion is that the next useful step for LLM trading research is not only better agent design, but also clearer reporting standards for execution realism, reproducibility, and evaluation comparability.
Sources
- Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models
- Alpha-GPT: Human-AI Interactive Alpha Mining for Quantitative Investment
- TradingGPT: Multi-Agent System with Layered Memory and Distinct Characters for Enhanced Financial Trading Performance
- FinMem: A Performance-Enhanced LLM Trading Agent with Layered Memory and Character Design
- QuantAgent: Seeking Holy Grail in Trading by Self-Improving Large Language Model
- A Multimodal Foundation Agent for Financial Trading: Tool-Augmented, Diversified, and Generalist
- Learning to Generate Explainable Stock Predictions using Self-Reflective Large Language Models
- FinLlama: Financial Sentiment Classification for Algorithmic Trading Applications
- StockGPT: A GenAI Model for Stock Prediction and Trading
- When AI Meets Finance (StockAgent): Large Language Model-based Stock Trading in Simulated Real-world Environments
- A Reflective LLM-based Agent to Guide Zero-shot Cryptocurrency Trading
- FinCon: A Synthesized LLM Multi-Agent System with Conceptual Verbal Reinforcement for Enhanced Financial Decision Making
- TradingAgents: Multi-Agents LLM Financial Trading Framework
- INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent
- Sentiment trading with large language models
- AlphaAgents: Large Language Model based Multi-Agents for Equity Portfolio Constructions
- MM-ARC: Multimodal Adaptive Routing of Capital with Robustness-Audited Strategy Pools
- FinRL-DeepSeek: LLM-Infused Risk-Sensitive Reinforcement Learning for Trading Agents
- AlphaAgent: LLM-Driven Alpha Mining with Regularized Exploration to Counteract Alpha Decay
- Can Large Language Models Trade? Testing Financial Theories with LLM Agents in Market Simulations
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection