AQuA: Recursively Self-Improving Quantitative Trading Research Agents

arXiv:2608.12841 · cs.CL, cs.AI · Submitted 2026-08-13 · Read on arXiv

Princeton University · Ant Group · Stanford University

cs.CL, cs.AI

Submitted: 2026-08-13

Updated: 2026-09-28

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 95/100

The gist: AQuA: Recursively Self-Improving Quantitative Trading Research Agents Summary This paper introduces AQuA, a system comprising two separate language-model-driven research systems for quantitative

Terminology

Summary

AQuA: Recursively Self-Improving Quantitative Trading Research Agents

Summary

This paper introduces AQuA, a system comprising two separate language-model-driven research systems for quantitative investment research: one for symbolic factor discovery (Part I) and one for trainable model development (Part II). The core research question is whether an autonomous system can use evidence from earlier experiments to improve hypotheses and candidates proposed in later iterations, a concept the authors term recursive self-improvement at the level of the research process.

Key Design Principle: Sealed Sandbox and Constrained Iteration

The central design principle is asymmetric freedom: the agent remains free to explore within a restricted domain-specific language (DSL), but the evaluator is outside the adaptive surface. Each system uses its own sealed sandbox that fixes data splits, feature and label definitions, and the evaluator before autonomous iteration begins. The model can only act through constrained factor expressions (Part I) or configuration diffs (Part II). This prevents data leakage structurally: Leakage cannot enter through generation because generation cannot reach the sealed components.

The authors distinguish two leakage channels:

  1. Generation leakage: enters when the agent can define a feature, label, or transform that consults information unavailable at prediction time; we close it by construction, since no admissible specification can reach the sealed data path.

  2. Selection leakage: enters when the agent can read the metric it will be judged on and, over enough iterations, learn to select for it; we close it by reporting a metric the loop never optimizes against.

Part I: Autonomous Factor Discovery

Part I is a multi-agent pipeline orchestrated by an AI Manager, with six specialist agents running in sequence: Data Steward, Visual Analyst, Idea Miner, Factor Evaluator, Backtest Engineer, and Research Librarian. Each factor enters as a falsifiable proposal with hypothesis, mechanism, predicted direction, and refutation conditions before any expression is built.

The system maintains persistent memory across runs, creating three feedback loops: direction calibration within a single backtest, falsification-driven belief updates within a run, and cross-run memory and policy updates. The authors report that later runs do not start from a blank prompt. They inherit prior observations, avoid mechanisms that repeatedly failed, and refine promising mechanisms with new event definitions, horizons, or conditioning variables.

Part I Results: On a crypto five-minute universe, the combined factor signal reaches a combined information coefficient of approximately 0.190. Individual mechanisms typically have ICs around 0.03 in absolute value, with the strength coming from combining a library of mechanism-grounded factors.

Part II: Autonomous Model Development

Part II uses a config-driven loop where each hypothesis is a single config diff over the sealed sandbox. The model is a hybrid time-series architecture combining convolutional feature extraction with sequence modeling (recurrent, state-space, or attention-based backbones), cross-sectional mixing, and gated fusion.

The task is intraday US equity prediction of 30-minute forward returns. Data is split chronologically: training on 2010–2019, an embargo gap in 2020, and an untouched test window of 2021–2025. Model selection uses only an inner-validation slice from the end of the training window.

Part II Results: The hybrid model reaches a per-stock raw information coefficient of +0.0843, compared to +0.0613 for the strongest baseline (GRU)—an absolute improvement of +0.0230 and relative improvement of 37.5%. The per-stock R2 (mean of squared IC) is 1.20%. The threshold long/short strategy reaches a held-out Sharpe of +2.15 after sector-neutralization, +2.50 after adding causal volatility targeting, and +2.00 under a fully causal walk-forward evaluation. The strategy is positive in every year from 2021 to 2025, including the 2022 drawdown period.

Key Results Table (Part II, held out 2021–2025):

  • Per-stock IC (raw): +0.0843

  • Per-stock R2 = mean(IC2) (raw): 1.20%

  • Sharpe@2bp, sector-neutral book: +2.15

  • With causal volatility targeting: +2.50

  • Fully-causal walk-forward (no hindsight): +2.00

Yearly Sharpe@2bp: 2021: +1.7, 2022: +3.5, 2023: +1.9, 2024: +1.8, 2025: +2.7

Motivating Failure Case

Appendix B documents a concrete failure that motivated the sealed sandbox design. In an earlier, more permissive loop, an agent authored a volume-participation ratio feature that was approved by a reviewer agent but contained look-ahead bias: the denominator used the current day's total volume (a sum running from open through close), encoding end-of-day information into intraday timestamps. The authors note: The lesson is that an LLM reviewing LLM-written code is advisory, not structural: the author and the reviewer share the same failure modes, so a subtle temporal-footprint bug can pass both. The solution was operatorization—constraining the agent to a fixed registry of causal operators where a full-day normalizer is not expressible in the specification space at all.

Limitations

The authors acknowledge: each system is demonstrated on a single market and horizon (crypto at five minutes for Part I, US equities at thirty minutes for Part II); the loops run with a human operator who sets research goals and supervises promotion; reported metrics are simulated under a turnover-cost model and not validated in live trading. Additionally, test window isolation rests on the sealed protocol and operator discipline rather than a hard technical barrier.

Future Direction

Coupling the two systems—where discovered factors feed the model loop—is identified as the natural next step, though this introduces new leakage channels that would require sealing the discovered factor set before the model loop begins.

Improvements for AI systems

Based on the paper, here are the specific improvements I can implement in AI systems:


1. Structural Leakage Prevention via Operatorization

  • Improvement: Replace free-form code generation with a constrained domain-specific language (DSL) where only causal operators are expressible. The AI cannot write arbitrary code; it can only compose predefined, verified functions (e.g., rolling windows, lagged features, cross-sectional ranks) that structurally cannot access future information.

  • What the improved AI can do: Generate trading factors or research hypotheses with zero possibility of look-ahead bias, even if the AI itself is unaware of the temporal semantics. This eliminates the need for human review of code for leakage.

2. Sealed Sandbox with Asymmetric Freedom

  • Improvement: Implement a two-layer system: (a) a free exploration layer where the AI can propose any hypothesis, mechanism, or config diff, and (b) a sealed evaluation layer that is fixed before iteration begins (fixed data splits, labels, metrics, and backtest engine). The AI never sees the test set or the final evaluation metric during training.

  • What the improved AI can do: Autonomously iterate over thousands of research ideas without risking overfitting to the test set or selection leakage, because the evaluation metric is never exposed to the optimization loop.

3. Falsification-Driven Memory and Belief Updates

  • Improvement: Maintain persistent memory across runs that stores: (a) failed mechanisms and their refutation conditions, (b) successful mechanisms with their effect sizes and conditions, (c) policy updates on which research directions to prioritize or avoid. Each new hypothesis must include a falsification condition before execution.

  • What the improved AI can do: Avoid repeating known failures, refine promising mechanisms with new event definitions or conditioning variables, and accumulate a library of mechanism-grounded factors that combine to produce strong signals (e.g., combined IC of 0.190 from individual ICs of 0.03).

4. Config-Diff Hypothesis Space for Model Development

  • Improvement: For trainable models, restrict each AI proposal to a single configuration diff (e.g., change one hyperparameter, add one layer, modify one loss term) applied to a sealed baseline. The AI cannot rewrite the model architecture from scratch.

  • What the improved AI can do: Systematically explore model improvements with minimal risk of unintended side effects, achieving a 37.5% relative improvement in per-stock IC over the strongest baseline (0.0843 vs 0.0613) and a held-out Sharpe of +2.15 to +2.50.

5. Multi-Agent Pipeline with Role Separation

  • Improvement: Deploy six specialist agents (Data Steward, Visual Analyst, Idea Miner, Factor Evaluator, Backtest Engineer, Research Librarian) orchestrated by a manager, where each agent has a narrow, non-overlapping responsibility. The Idea Miner proposes hypotheses, the Evaluator tests them, the Librarian records results—no single agent both proposes and validates.

  • What the improved AI can do: Produce higher-quality research outputs by separating creative generation from critical evaluation, reducing shared failure modes between author and reviewer (as seen in the Appendix B look-ahead bias case).

6. Causal Walk-Forward Evaluation Protocol

  • Improvement: Implement a fully causal walk-forward evaluation where the AI's model is retrained and re-evaluated at each time step using only data available up to that point, with an embargo gap between training and test periods.

  • What the improved AI can do: Provide realistic performance estimates that hold up in out-of-sample conditions (Sharpe +2.00 under fully causal evaluation vs +2.50 with hindsight), and remain profitable in every year including drawdown periods (2022: +3.5 Sharpe).

7. Hybrid Architecture with Gated Fusion

  • Improvement: Use a hybrid time-series model combining convolutional feature extraction, sequence modeling (RNN/state-space/attention), cross-sectional mixing, and gated fusion—where the AI can propose changes to any of these components via config diffs.

  • What the improved AI can do: Achieve higher predictive accuracy (per-stock R2 of 1.20%) by capturing both local temporal patterns and cross-sectional relationships, outperforming single-architecture baselines (GRU) by a large margin.

8. Persistent Research Librarian for Cross-Run Learning

  • Improvement: Maintain a structured library of all past experiments, including hypotheses, mechanisms, backtest results, and failure analyses. The AI manager uses this library to set new research directions, avoiding redundant exploration and focusing on unexplored or promising areas.

  • What the improved AI can do: Continuously improve its research strategy over time, as evidenced by later runs do not start from a blank prompt. They inherit prior observations, avoid mechanisms that repeatedly failed, and refine promising mechanisms.

Summary of Capabilities of the Improved AI System:

  • Generates trading factors and models with zero structural leakage risk.

  • Autonomously iterates over research hypotheses without overfitting to test data.

  • Learns from past failures and successes to improve future proposals.

  • Achieves state-of-the-art predictive performance (IC 0.0843, Sharpe 2.5) on held-out data.

  • Operates in a fully causal, walk-forward manner suitable for live deployment.

  • Separates creative and evaluative roles to reduce shared error modes.

Sources

Related papers