CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents
Listen
Radio episode about this paper
Transcript
Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.
Tom: Next we'll be talking about the paper "CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents".
Jane: The paper was written by Bo Qu and Mingguang Chen from.
Tom: Stay tuned as we take you through the paper and discuss its implications.
Jane: We also have Lu with us today — senior AI researcher at Tsinghua.
Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.
Jane: We also have Lalam with us today — the in-house Large Language Model.
Tom: Alright, let's get started.
Title: Tom: Jane, we are starting our show with a paper that has a really long name.
Jane: You mean "CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents"?
Tom: That's the one.
Jane: It is written by Bo Qu and Mingguang Chen.
Tom: They are tackling a huge problem in how we judge AI agents that trade stocks.
Jane: Most people just look at the final profit, don't they?
Tom: Exactly, but that can be very misleading if the agent was just lucky with a market trend.
Lu: It's like judging a marathon runner only by their time, without checking if they actually followed the correct route!
Tom: That is a perfect analogy, Lu.
Jane: So instead of ranking them on a leaderboard, they want to diagnose them?
Meng: I wonder if this kind of deep diagnostic testing is even possible to scale in a real-world setting.
Tom: That is the challenge the authors are trying to solve with this benchmark.
Jane: It seems like they want to see if the agent's logic actually holds up over time.
Lalam: If we can move toward this kind of evaluation, we can build AI that behaves in ways humans can truly understand.
Summary: Tom: Now that we have the title out of the way, let's look at how this CLQT system actually works.
Jane: They use a loop where an agent gathers data and moves into synthesis before building a portfolio.
Tom: After they build it, they execute trades and then reflect on those actions to learn.
Jane: And the most important part is the APM-CS scorecard they use to grade that whole process.
Tom: It looks at things like Coherence to see if the agent's reasoning matches its actions.
Jane: They also check Acuity to see if it can ignore noisy data and focus on real signals.
Lu: I think the three-tier memory system they built is incredibly clever for long-term learning!
Tom: It really is, Lu, because it lets the agent actually remember its past mistakes.
Jane: And they even included a way to measure how much an agent's decisions change during volatile times.
Meng: I was impressed by how they modeled transaction costs and financing fees so accurately.
Tom: You mean they didn't just pretend that trading was free?
Meng: Right, because ignoring those costs makes any benchmark for a trader completely useless.
Jane: They also use a Reliability axis to make sure the agent doesn't crash technically while it works.
Lalam: By looking at all these different layers, we get a much better sense of what intelligence really looks like in practice.
Improvements: Tom: We've seen how the system is built, so let's talk about what happened when they actually ran the experiments.
Jane: They did these ablation studies where they intentionally turned off certain modules to see if it mattered.
Tom: And even though the profit sometimes stayed the same, the actual capability scores dropped significantly.
Jane: One of their biggest findings was that "Coherence Gap" we mentioned earlier.
Tom: You mean when an agent says one thing but does another?
Jane: Yes, they found that agents often write a great research report but then make totally different trades.
Lu: That's such a wild finding because it shows the agent is just performing for the prompt!
Tom: It really does, Lu.
Lu: It means we need to stop trusting what an AI says and start looking at its actual behavior.
Meng: I also noticed that using a "structured" committee of agents helps prevent these logic collapses.
Jane: You mean having specialized roles like a Risk Officer instead of just one big agent?
Meng: Exactly, because it provides the guardrails that keep the agent from drifting away from its mandate.
Tom: It's basically adding a layer of institutional discipline to the AI.
Lalam: Bridging that gap between an agent's words and its actions is vital for creating reliable AI partners.
Conclusion: Tom: That brings us to the end of our deep dive into "CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents."
Jane: It's a fascinating look at how we can move past simple leaderboards and toward real capability maps.
Tom: We've learned that being profitable doesn't always mean you are a good agent.
Jane: And that the way an agent thinks is just as important as the money it makes.
Lu: I'm excited to see how this helps us build agents that can actually learn from their own experiences!
Meng: And I'll be watching for these diagnostic scores to become a standard in our engineering pipelines.
Lalam: Ultimately, this work moves us closer to a world where AI decisions are transparent and accountable.
Tom: Thanks for joining us, everyone.
Jane: See you next time!
cs.AI, cs.LG, q-fin.CP, q-fin.PM
Submitted: 2026-06-29
Updated: 2026-09-12
Comments: 53 pages, 14 figures, 10 tables. Expanded benchmark comparison and related work, compressed and copy-edited presentation, references converted to author-year, appendices consolidated, and a blinded human-anchoring study of the held-out LLM judge added
License: http://creativecommons.org/licenses/by/4.0/
Importance score: 85/100
The gist: CLQT is a closed-loop, cost-aware, and strategy-consistent benchmark designed for the diagnostic evaluation of LLM portfolio-management agents.
Key concepts
- APM-CS Scorecard
- A grading system used by the CLQT benchmark to evaluate the entire trading process. It measures Coherence to ensure reasoning matches actions, Acuity to see if agents can distinguish real signals from noisy data, and Reliability to ensure technical stability during operation.
- Coherence Gap
- A phenomenon where an AI agent's written reasoning or research reports do not align with its actual trading actions. This finding suggests that agents may provide impressive text while making contradictory or illogical financial decisions.
- Structured Committee of Agents
- An approach to managing AI agents by assigning them specialized roles, such as a Risk Officer, instead of using one single large agent. This provides institutional discipline and guardrails that prevent the agent from drifting away from its mandate.
Terminology
Summary
CLQT is a closed-loop, cost-aware, and strategy-consistent benchmark designed for the diagnostic evaluation of LLM portfolio-management agents. It addresses critical gaps in existing trading benchmarks—specifically look-ahead bias, unrealistic cost modeling, and the inadequacy of return-only leaderboards—by reframing evaluation as a diagnosis rather than ranking.
By localizing where an agent's process succeeds or fails, CLQT provides a durable, extensible map of agent competencies and limitations
rather than a single scalar score.
System Architecture and Design Pillars
The benchmark operates through a five-stage cycle following the reason–act–reflect tradition
: gather, synthesize, construct a mandate-aware portfolio, execute under realistic costs, then reflect. This process is implemented via a six-agent architecture
where specialized roles—Research, Analyst, PM, Risk Officer, Execution, and Reflection—receive specific tool subsets. To ensure data integrity and auditability, every cycle seals one DecisionRound
into a recompute-verifiable hash chain,
making every metric reconstructable from the trail.
The substrate is built upon six core design pillars:
-
A hard TimeGate that treats leakage as a failed precondition;
-
Institutional transaction and financing-cost modeling;
-
Strategy-consistency scoring across rounds;
-
A three-tier consolidating memory (working, episodic, and semantic);
-
A Model-Context-Protocol tool layer with 19 quantitative tools;
-
A mandate-aware synthesis pipeline.
The Five-Axis Capability Scorecard
Rather than relying on a single performance metric, CLQT utilizes a five-axis capability scorecard (APM-CS)
to map agent limitations. This diagnostic approach separates outcome from capability
across both backtest and live tracks. The axes represent finance-instantiated versions of general agent capabilities:
-
Coherence: Judged by whether the chosen allocation follows from the agent's stated research and signals;
-
Acuity: The ability to attend to
informative signals amid a large noisy universe
; -
Composure: Measuring robustness against
overreaction to volatility
and transient noise; -
Discipline: The capacity for
systematic alignment to a standing mandate
without relying on hard guardrails; -
Reliability: Operational robustness, including protocol adherence and completion rates.
Experimental Validation and Findings
The benchmark was validated through a contamination-controlled multi-model backtest
and a live broker track on unseen, post-cutoff data.
The results demonstrate that the instrument successfully separates decision quality from market outcomes. A primary finding is that the capability leader is not the Sharpe leader,
as nominal winners are often exposed as reliability artifacts
where high returns are achieved through an insufficient number of rounds or unstable processes.
Furthermore, the study reveals that:
-
Module value registers on the capability axes when returns cannot separate it,
meaning specific components like memory or sentiment analysis show impact on reasoning even if they do not move the Sharpe ratio; -
Agents often exhibit a
coherence gap,
where they trade in the right direction but their allocations fail to follow their own stated analysis; -
Reliability deficits [are] hidden by returns,
allowing models with poor schema adherence to appear profitable in a returns-only view.
Improvements for AI systems
1. Implementation of a Three-Tier Memory System (Working, Episodic, Semantic) integrated with a ReAct-Reflexion loop.
- What the improved AI system can do: The system can transform raw historical event-action-outcome records into generalized procedural knowledge. This allows the agent to move beyond simple pattern matching to adapt its strategy across different market regimes and avoid repeating specific, identified historical errors.
2. Decoupling of Proposed Intent
from Final Execution
via a Mandate-Aware Portfolio Construction (MAPC) reconciler.
- What the improved AI system can do: The system ensures that an agent's stated reasoning (declarative knowledge) is strictly aligned with its actual trades (procedural action). It prevents
coherence gaps
where an agent's allocations contradict its own stated analysis, and it ensures all trades remain within strict risk, turnover, and budget constraints without relying on the LLM to self-regulate perfectly.
3. Transition from single-orchestrator autonomy to Structured Skill Modes
using role-gated multi-agent committees.
- What the improved AI system can do: The system prevents
agentic orchestration failures,
such as non-terminating tool-use loops, by enforcing specialized workflows (e.g., a dedicated Risk Officer gate). This provides institutional-grade stability and ensures that the failure of one reasoning step does not cause a total collapse of the decision cycle.
4. Integration of a 3-Dimensional SCOUT stage (Technical Momentum, Core Fundamentals, and Macro Correlation) and a Volatility-Regulated Turnover model.
- What the improved AI system can do: The system can differentiate between high-conviction alpha signals and
noise
(such as RSI oscillators or transient volatility). This enablesComposure,
preventing the agent from overreacting to market turbulence and ensuring it ignores momentum-based distractors in favor of fundamental and macro-correlated drivers.
5. Implementation of a Recompute-Verifiable Hash-Linked Audit Chain and a strict TimeGate data-access protocol.
- What the improved AI system can do: The system guarantees absolute temporal integrity, making look-ahead bias (information leakage) mathematically impossible. It provides a tamper-proof, forensic-grade audit trail where every observation, tool call, and decision is hashed and linked, allowing for independent, third-party verification of every trade for regulatory compliance.
6. Integration of multi-component, calibrated cost-modeling (transaction, financing, and market-impact costs).
- What the improved AI system can do: The system optimizes for net-of-cost profitability rather than gross-return maximization. It can accurately predict the drag of leverage, financing, and slippage, preventing the deployment of
paper-profitable
strategies that would fail in real-world, high-friction execution environments.
Abstract
LLM agents are increasingly cast as autonomous portfolio managers, yet the dominant evaluation idiom, a leaderboard of returns over a fixed window, certifies neither the soundness of an agent's process nor the durability of its edge: one period's return is dominated by the market path, and apparent alpha can dissolve once look-ahead bias and trading costs are controlled. We introduce CLQT, a closed-loop benchmark that reframes LLM trading evaluation as diagnosis rather than ranking. CLQT enforces point-in-time data access through a hard TimeGate, models institutional transaction and financing costs, scores strategy consistency across rounds, and seals every gather-analyze-decide-execute-reflect cycle into a recompute-verifiable audit chain; the same model runs as a constrained investment committee or a single autonomous orchestrator, making scaffolding an experimental variable. From the audit trail CLQT computes a five-axis capability scorecard (Coherence, Acuity, Composure, Discipline, Reliability), with coherence scored partly by a held-out LLM judge to curb self-preference bias. We validate CLQT on a contamination-controlled, year-long multi-model backtest campaign with a 13-configuration ablation grid and a four-week live broker paper-trading track on post-cutoff data. The diagnosis-first read shows that the capability leader is not the Sharpe leader; that agents' allocations systematically fail to follow their own stated analysis, a stating-versus-doing gap stable across both tracks (+0.30 backtest, +0.23 live); and that module value registers on the capability and behavioral axes when returns alone cannot separate it. Net of realistic costs, agents clear defensive baselines but do not cleanly beat the index. Credible evaluation of LLM investment agents must therefore diagnose the process rather than rank a period's return, the standard CLQT operationalizes and makes auditable.
Sources
- StockBench: Can LLM Agents Trade Stocks Profitably In Real-world Markets?
- LiveTradeBench: Seeking Real-World Alpha with Large Language Models
- AI-Trader: Benchmarking Autonomous Agents in Real-Time Financial Markets
- Will LLMs be Professional at Fund Investment? DeepFund: A Live Arena Perspective
- INVESTORBENCH: A Benchmark for Financial Decision-Making Tasks with LLM-based Agent
- Can LLM-based Financial Investing Strategies Outperform the Market in Long Run?
- PortBench: A Correlation-Aware, Full-Pipeline Benchmark for LLM-Driven Portfolio Management
- Profit Mirage: Revisiting Information Leakage in LLM-based Financial Agents
- From Knowing to Doing: A Memory-Controlled Benchmark for LLM Trading Agents on Stock Markets
- TradingAgents: Multi-Agents LLM Financial Trading Framework
- FinMem: A Performance-Enhanced LLM Trading Agent with Layered Memory and Character Design
- Beyond Task Completion: Revealing Corrupt Success in LLM Agents through Procedure-Aware Evaluation
- Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena
- Holistic Evaluation of Language Models
- ReAct: Synergizing Reasoning and Acting in Language Models
- Reflexion: Language Agents with Verbal Reinforcement Learning
- Generative Agents: Interactive Simulacra of Human Behavior
- MemGPT: Towards LLMs as Operating Systems
- Cognitive Architectures for Language Agents
- ToolLLM: Facilitating Large Language Models to Master 16000+ Real-world APIs
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection