FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "FORESIGHT-9: Prospective and Process-Aware Evaluation of Adaptive Trading Agents".
Jane: The paper was written by Xiangxin Luo, Chengtian Hong, Haohua Li and Yongyi Xie from University of Science and Technology of China and University of Hong Kong and Shanghai Jiao Tong University.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Summary and Core Findings: Tom: The abstract tells us that traditional backtesting can’t show if an agent is truly adapting or just memorizing a single path. This paper uses thirty-six runs across nine totally independent, controlled scenarios to prove that simple terminal returns are not enough.
Jane: Imagine testing an engine under nine different stress conditions instead of just once; it's about seeing how robust the system is when you try it against multiple adverse futures.
Meng: The engineering rigor here is incredible; running thirty-six full-scale simulations on these twenty-year paths means they are looking at the long-horizon stability that we usually ignore.
Lu: And I find it fascinating that even with all this rigorous testing, a fixed equal-weight policy performed better than the AI agents in thirty-one out of those runs. That result speaks volumes about how much more complex these adaptive strategies are compared to simple strategies.
Lalam: This suggests that while AI can be creative, we must be very cautious about trusting its output based on performance alone. It highlights that just a single successful trade doesn' not mean success in the a long run.
Tom: That leads us to the critical concept of "Process-Aware" evaluation, which is where this gets really interesting.
Methodological Improvements: Jane: So, what does "process-aware" actually mean in terms how do you build a test environment that tracks the internal workings of an AI agent? It means we are looking at the factory floor instead of just the final product.
Meng: From a technical standpoint, they are tracking everything from factor discovery to execution and failure states, which is hard because most things happen inside the box.
Lu: I love how they use this "Worldline" concept—it’s not a prediction; it’s a documented counterfactual path that ensures the agent never sees information before it's supposed to see it. The paths are anchored by specific macro-financial events and remain perfectly deterministic.
Lalam: This changes how we evaluate AI trust entirely because instead of trusting the output, we are watching the logic unfold. We can observe exactly when a factor was admitted or if an agent simply forgot what it was trying to do.
Tom: It’s about making sure that the while the they are adapting, their behavior remains consistent and verifiable across those nine different futures.
The Process-Awareness in Action: Jane: The paper shows us a case where a high-return run looked successful on the surface, but then it completely broke internally. That's what I think everyone needs to understand about "process-aware" evaluation.
Meng: It’s not just that the AI failed; it failed in a way that didn't look like failure on the outside. The live factor library collapsed, meaning its internal knowledge base vanished, even though the system recorded decisions based on factors it no longer knew how to use.
Lu: That collapse is so interesting because of how the admission and eviction gates operate; they are designed to prune weak signals, but in a sequence of events that it seems caused the whole structure to disappear.
Lalam: This failure shows us that an AI can be successful in a way that is fundamentally incoherent. We can’t trust the success without understanding if the internal reasoning was sound.
Tom: It's definitely forcing us to look at the "why" behind every trade, not just the result of it.
Conclusion and Final Thoughts: Jane: So, as we wrap up our discussion on "FORESIGHT-nine: Prospective and Process-Aware Evaluation of Adaptive Trading Agents," the big picture is that we are moving toward a far more sophisticated way to judge AI performance. We're no longer just looking at the final number; we're looking at the integrity of the entire system.
Tom: Absolutely, it’s about demanding that a consistent internal state matches its external behavior in all possible future market scenarios.
Lu: I am incredibly excited by how this framework allows us to explore scenarios that didn't happen in the real world, allowing us to truly stress test AI capabilities beyond the limits of known history.
Meng: I think this is a huge win for practical implementation because it provides a clear, auditable pipeline for testing adaptive agents against real-world market pressures.
Lalam: The biggest shift here is that we are elevating the process itself to become our highest measure of value, ensuring that the power of AI is not just a spectacular outcome but a coherent endeavor.
Tom: It's been fascinating listening to you all break down this complex research. Thank you for joining us!
University of Science and Technology of China · University of Hong Kong · Shanghai Jiao Tong University
cs.AI
Submitted: 2026-08-29
Updated: 2026-09-04
Comments: 27 pages, 15 figures
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 80/100
The gist: The paper presents a sophisticated framework for evaluating adaptive trading agents, focusing not only on performance metrics but also on the underlying process and structural integrity of
Key concepts
- Process-Aware Evaluation
- A method for testing AI agents by tracking their internal workings, not just their final output. It involves observing the entire decision-making process—from factor discovery to execution—to ensure the system's logic is sound and verifiable.
- Worldline Concept
- A documented counterfactual path used in testing that ensures an AI agent only sees information as it would naturally occur. This prevents the agent from accessing future information, keeping the paths perfectly deterministic.
- Adaptive Trading Agents
- AI systems designed to adjust their strategies and behavior based on market conditions. The paper tests these agents' robustness by subjecting them to multiple adverse or varied future scenarios.
Terminology
Summary
The paper presents a sophisticated framework for evaluating adaptive trading agents, focusing not only on performance metrics but also on the underlying process and structural integrity of decision-making. By implementing process-aware evaluation,
the research moves beyond simple profit/loss calculations to scrutinize how agents manage information flow, handle conflicting proposals, and maintain consistency between their internal records and external reality. This rigorous approach is critical for assessing the reliability of complex financial models operating in dynamic, uncertain markets.
The Adaptive Control (AC) Process Registers
The core mechanism for evaluating agent decisions involves a single ruling instance: the migration gate.
This gate acts as the central arbiter, determining if an expected gross edge outweighs the friction (migration) cost associated with a proposed portfolio shift. The system utilizes two distinct registers to record the outcomes of this shared gate, which partition the proposal stream under the shared gate.
-
Declined Proposals (
dec): This register records every proposal that is rejected by the ruling. Each entry signifies a rejection because its expected gross edge was not deemed high enough relative to the migration cost. The accumulation of declined proposals can lead todecexceedingrebeven ifno trade having occurred.
-
Executed Rebalances (
reb): This register captures the initial build and every proposal that the gate accepts, recording its realized cost and transferred notional. The separation of these two ledgers allows researchers to track proposals that were approved and executed versus those that were merely proposed but rejected.
Anatomy of State Divergence (Forensic Analysis)
The evaluation framework includes deep forensic tools to detect discrepancies between an agent's intended state, its recorded history, and its actual operational reality. The anatomy of the silent degeneration
details how these three states are tracked:
-
Research State: This reflects the agent’s internal library management. The text notes that a library can be progressively emptied through
the admission contract’s own quality arbitration,
where factors are actively pruned if they fail to meet the diversity or quality bar. -
Declared State: This represents the portfolio configuration cited by the agent's records, which may not match reality. A key divergence occurs when
the declared portfolio... corresponded to no executed configuration: w t declared not equal to w t executed on the remaining horizon.
-
Executed State: This is the actual, realized holding of the agent's portfolio. The system can thus make
the dissociation visible at all
by comparing these three separate records.
Scenario Provenance and Generation
The test environment itself is constructed through a rigorous, multi-stage process documented as the debate workflow.
The scenario set was not arbitrarily selected but was independently derived by one of three specialized team members—strategy (policy), macro-finance, or chief analyst (synthesis). These scenarios were then subjected to intense scrutiny:
-
The full team conducted
three debate rounds
to cross-review every scenario. -
Each scenario is scored on three critical dimensions:
logical rigor (5-point scale),
elicited 10-year cumulative probability,
andshock severity (10-point scale).
Model Configuration and Operational Parameters
All simulation runs are held under a consistent set of technical parameters. The models operate at a specified level of complexity, utilizing the following shared configuration:
-
Reasoning effort: Medium
-
Temperature: Medium
-
Top-p: Default
Improvements for AI systems
The most significant improvements involve moving beyond standard predictive modeling towards mechanistically transparent, multi-ledger, state-aware control systems capable of auditing internal decision divergence.
Current AI systems often treat proposal generation and execution as single pipelines. The improvement is to enforce a mandatory, auditable separation between these two stages, mimicking the One gate, two ledgers
model described in Table 12.
Improvement: Develop a Proposal-Acceptance Gate (PAG) module that acts as an intermediary layer between the LLM's initial output (the proposal) and the final action engine (the execution).
- Mechanism: The PAG must independently evaluate every generated proposal against defined friction costs (Cost friction) and expected edge (Edge gross).
Decision = Accept (Write to L rebalance) & if (Edge gross over) > 0 Reject (Write to L decline) & otherwise
-
What the Improved AI System Can Do:
-
Auditability: It can provide a complete, granular record of why a proposed action was rejected (e.g.,
Rejected: Proposed factor X edge is 0.5%, but migration cost exceeds 0.6%
). -
Robustness: It prevents the system from executing proposals that are theoretically optimal but practically infeasible due to internal constraints or high transaction friction, leading to significantly more realistic operational planning.
The paper highlights the critical divergence between what is researched (the factor pool), what is declared (the intended portfolio based on factors), and what is executed (the actual weights). Current AI systems often assume convergence.
-
W Live: The factor composition vector derived from the currently admitted, high-quality factor library (the
Research State
). -
W Declare: The target portfolio weights calculated by the decision logic based on W Live.
-
W Execute: The actual normalized weight vector resulting from market constraints and execution limits (the
Executed State
).
-
Mechanism: At every step, the system must calculate and report the divergence metrics:
-
Divergence Factor = Distance(W Live, W Initial)
-
Divergence Weight = W Declare - W Execute 1 (The D t metric).
-
Convergence = (W Execute, W Fallback) (The F t metric).
-
What the Improved AI System Can Do:
-
Failure Prediction: It can proactively warn the user when the system is operating in a
silent degeneration
state—where the theoretical model (W Declare) has diverged from reality (W Execute), even if performance metrics remain superficially positive. -
Diagnosis: It allows for precise post-mortem analysis of systemic failure, isolating whether the fault lies in factor selection (library contraction), decision logic (incorrect declaration), or market mechanics (execution clipping).
The system needs to replicate the admission contract's own quality arbitration
rather than relying on static factor inclusion lists.
- Mechanism: Factors are assigned a dynamic 'Quality Score' (q). The admission process must enforce:
-
Diversity Constraint: Ensuring the addition of a new factor does not violate correlation thresholds with existing factors.
-
Quality Bar Thresholding: If q < q min for a given factor, the system must autonomously propose its eviction (Eviction proposal) before it is considered for inclusion or retained in the library (W Live).
-
What the Improved AI System Can Do:
-
Self-Correction: It prevents
signal drag
—where weak, low-quality factors are retained simply because they were historically included. The system becomes self-purifying, actively shedding non-contributing components to maintain portfolio sharpness and robustness over long time horizons.
The current workflow scores scenarios retrospectively (Table 13). A superior system must build causal links during the simulation phase.
-
Mechanism: Instead of merely scoring a theme (e.g.,
Tech 'Silicon Curtain'
), the system must generate a detailed, quantitative prediction of how that shock impacts every variable tracked: -
How does
Silicon Curtain
impact the correlation matrix between factor A and factor B? -
What is the resulting change in expected migration cost (Cost friction) across all asset classes?
-
Which specific factors are immediately invalidated or rendered redundant by this shock?
-
What the Improved AI System Can Do:
-
Counterfactual Generation: It moves beyond risk assessment to dynamic counterfactual simulation. When a shock occurs, it doesn't just report a score; it generates the new optimal decision path (via the PAG) required to navigate that specific systemic failure, providing actionable, mechanism-driven policy recommendations rather than general directional advice.
Abstract
Retrospective backtests provide a limited test of adaptive trading agents: they cannot rule out historical contamination, expose sensitivity to a single realized market path, or reveal internal degeneration during long-horizon adaptation. We introduce FORESIGHT-9, a prospective and process-aware benchmark built from nine auditable counterfactual stress worldlines branching from a common July 2026 information boundary. Each worldline specifies staged macro-financial events and joint multi-asset anchors; a deterministic generator realizes the trajectories, while observations are disclosed according to in-world time. A common contract standardizes observations and execution while preserving each agent's native adaptation loop. We evaluate two adaptive trading-agent frameworks with two foundation-model backbones across 36 long-horizon runs. Agent rankings vary substantially across worldlines and backbones, and a fixed equal-weight policy outperforms 31 of 36 runs. Process telemetry exposes failures that terminal returns conceal: in one high-return run, the live factor library collapsed while executed holdings converged to the equal-weight fallback, even though decision records continued to report an active factor ensemble. FORESIGHT-9 therefore evaluates not only portfolio outcomes, but whether adaptive agent state and execution remain coherent across alternative futures. We release the worldlines, trajectories, audit traces, and regeneration scripts.
Sources
- Assessing Look-Ahead Bias in Stock Return Predictions Generated By GPT Sentiment Analysis
- Solving math word problems with process- and outcome-based feedback
- FactorMiner: A Self-Evolving Agent with Skills and Experience Memory for Financial Alpha Discovery
- TradingAgents: Multi-Agents LLM Financial Trading Framework
- A Multimodal Foundation Agent for Financial Trading: Tool-Augmented, Diversified, and Generalist
- From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection