CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents

arXiv:2606.29771 · cs.AI, cs.LG, q-fin.CP, q-fin.PM · Submitted 2026-06-29 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents".

Jane: The paper was written by Bo Qu and Mingguang Chen from.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Jane, we are starting our show with a paper that has a really long name.

Jane: You mean "CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents"?

Tom: That's the one.

Jane: It is written by Bo Qu and Mingguang Chen.

Tom: They are tackling a huge problem in how we judge AI agents that trade stocks.

Jane: Most people just look at the final profit, don't they?

Tom: Exactly, but that can be very misleading if the agent was just lucky with a market trend.

Lu: It's like judging a marathon runner only by their time, without checking if they actually followed the correct route!

Tom: That is a perfect analogy, Lu.

Jane: So instead of ranking them on a leaderboard, they want to diagnose them?

Meng: I wonder if this kind of deep diagnostic testing is even possible to scale in a real-world setting.

Tom: That is the challenge the authors are trying to solve with this benchmark.

Jane: It seems like they want to see if the agent's logic actually holds up over time.

Lalam: If we can move toward this kind of evaluation, we can build AI that behaves in ways humans can truly understand.

Summary: Tom: Now that we have the title out of the way, let's look at how this CLQT system actually works.

Jane: They use a loop where an agent gathers data and moves into synthesis before building a portfolio.

Tom: After they build it, they execute trades and then reflect on those actions to learn.

Jane: And the most important part is the APM-CS scorecard they use to grade that whole process.

Tom: It looks at things like Coherence to see if the agent's reasoning matches its actions.

Jane: They also check Acuity to see if it can ignore noisy data and focus on real signals.

Lu: I think the three-tier memory system they built is incredibly clever for long-term learning!

Tom: It really is, Lu, because it lets the agent actually remember its past mistakes.

Jane: And they even included a way to measure how much an agent's decisions change during volatile times.

Meng: I was impressed by how they modeled transaction costs and financing fees so accurately.

Tom: You mean they didn't just pretend that trading was free?

Meng: Right, because ignoring those costs makes any benchmark for a trader completely useless.

Jane: They also use a Reliability axis to make sure the agent doesn't crash technically while it works.

Lalam: By looking at all these different layers, we get a much better sense of what intelligence really looks like in practice.

Improvements: Tom: We've seen how the system is built, so let's talk about what happened when they actually ran the experiments.

Jane: They did these ablation studies where they intentionally turned off certain modules to see if it mattered.

Tom: And even though the profit sometimes stayed the same, the actual capability scores dropped significantly.

Jane: One of their biggest findings was that "Coherence Gap" we mentioned earlier.

Tom: You mean when an agent says one thing but does another?

Jane: Yes, they found that agents often write a great research report but then make totally different trades.

Lu: That's such a wild finding because it shows the agent is just performing for the prompt!

Tom: It really does, Lu.

Lu: It means we need to stop trusting what an AI says and start looking at its actual behavior.

Meng: I also noticed that using a "structured" committee of agents helps prevent these logic collapses.

Jane: You mean having specialized roles like a Risk Officer instead of just one big agent?

Meng: Exactly, because it provides the guardrails that keep the agent from drifting away from its mandate.

Tom: It's basically adding a layer of institutional discipline to the AI.

Lalam: Bridging that gap between an agent's words and its actions is vital for creating reliable AI partners.

Conclusion: Tom: That brings us to the end of our deep dive into "CLQT: A Closed-Loop, Cost-Aware, Strategy-Consistent Benchmark for Diagnostic Evaluation of LLM Portfolio-Management Agents."

Jane: It's a fascinating look at how we can move past simple leaderboards and toward real capability maps.

Tom: We've learned that being profitable doesn't always mean you are a good agent.

Jane: And that the way an agent thinks is just as important as the money it makes.

Lu: I'm excited to see how this helps us build agents that can actually learn from their own experiences!

Meng: And I'll be watching for these diagnostic scores to become a standard in our engineering pipelines.

Lalam: Ultimately, this work moves us closer to a world where AI decisions are transparent and accountable.

Tom: Thanks for joining us, everyone.

Jane: See you next time!

cs.AI, cs.LG, q-fin.CP, q-fin.PM

Submitted: 2026-06-29

Updated: 2026-09-12

Comments: 53 pages, 14 figures, 10 tables. Expanded benchmark comparison and related work, compressed and copy-edited presentation, references converted to author-year, appendices consolidated, and a blinded human-anchoring study of the held-out LLM judge added

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 85/100

The gist: CLQT is a closed-loop, cost-aware, and strategy-consistent benchmark designed for the diagnostic evaluation of LLM portfolio-management agents.

Key concepts

APM-CS Scorecard
A grading system used by the CLQT benchmark to evaluate the entire trading process. It measures Coherence to ensure reasoning matches actions, Acuity to see if agents can distinguish real signals from noisy data, and Reliability to ensure technical stability during operation.
Coherence Gap
A phenomenon where an AI agent's written reasoning or research reports do not align with its actual trading actions. This finding suggests that agents may provide impressive text while making contradictory or illogical financial decisions.
Structured Committee of Agents
An approach to managing AI agents by assigning them specialized roles, such as a Risk Officer, instead of using one single large agent. This provides institutional discipline and guardrails that prevent the agent from drifting away from its mandate.

Terminology

Summary

CLQT is a closed-loop, cost-aware, and strategy-consistent benchmark designed for the diagnostic evaluation of LLM portfolio-management agents. It addresses critical gaps in existing trading benchmarks—specifically look-ahead bias, unrealistic cost modeling, and the inadequacy of return-only leaderboards—by reframing evaluation as a diagnosis rather than ranking. By localizing where an agent's process succeeds or fails, CLQT provides a durable, extensible map of agent competencies and limitations rather than a single scalar score.

System Architecture and Design Pillars

The benchmark operates through a five-stage cycle following the reason–act–reflect tradition: gather, synthesize, construct a mandate-aware portfolio, execute under realistic costs, then reflect. This process is implemented via a six-agent architecture where specialized roles—Research, Analyst, PM, Risk Officer, Execution, and Reflection—receive specific tool subsets. To ensure data integrity and auditability, every cycle seals one DecisionRound into a recompute-verifiable hash chain, making every metric reconstructable from the trail.

The substrate is built upon six core design pillars:

  • A hard TimeGate that treats leakage as a failed precondition;

  • Institutional transaction and financing-cost modeling;

  • Strategy-consistency scoring across rounds;

  • A three-tier consolidating memory (working, episodic, and semantic);

  • A Model-Context-Protocol tool layer with 19 quantitative tools;

  • A mandate-aware synthesis pipeline.

The Five-Axis Capability Scorecard

Rather than relying on a single performance metric, CLQT utilizes a five-axis capability scorecard (APM-CS) to map agent limitations. This diagnostic approach separates outcome from capability across both backtest and live tracks. The axes represent finance-instantiated versions of general agent capabilities:

  1. Coherence: Judged by whether the chosen allocation follows from the agent's stated research and signals;

  2. Acuity: The ability to attend to informative signals amid a large noisy universe;

  3. Composure: Measuring robustness against overreaction to volatility and transient noise;

  4. Discipline: The capacity for systematic alignment to a standing mandate without relying on hard guardrails;

  5. Reliability: Operational robustness, including protocol adherence and completion rates.

Experimental Validation and Findings

The benchmark was validated through a contamination-controlled multi-model backtest and a live broker track on unseen, post-cutoff data. The results demonstrate that the instrument successfully separates decision quality from market outcomes. A primary finding is that the capability leader is not the Sharpe leader, as nominal winners are often exposed as reliability artifacts where high returns are achieved through an insufficient number of rounds or unstable processes.

Furthermore, the study reveals that:

  • Module value registers on the capability axes when returns cannot separate it, meaning specific components like memory or sentiment analysis show impact on reasoning even if they do not move the Sharpe ratio;

  • Agents often exhibit a coherence gap, where they trade in the right direction but their allocations fail to follow their own stated analysis;

  • Reliability deficits [are] hidden by returns, allowing models with poor schema adherence to appear profitable in a returns-only view.

Improvements for AI systems

1. Implementation of a Three-Tier Memory System (Working, Episodic, Semantic) integrated with a ReAct-Reflexion loop.

  • What the improved AI system can do: The system can transform raw historical event-action-outcome records into generalized procedural knowledge. This allows the agent to move beyond simple pattern matching to adapt its strategy across different market regimes and avoid repeating specific, identified historical errors.

2. Decoupling of Proposed Intent from Final Execution via a Mandate-Aware Portfolio Construction (MAPC) reconciler.

  • What the improved AI system can do: The system ensures that an agent's stated reasoning (declarative knowledge) is strictly aligned with its actual trades (procedural action). It prevents coherence gaps where an agent's allocations contradict its own stated analysis, and it ensures all trades remain within strict risk, turnover, and budget constraints without relying on the LLM to self-regulate perfectly.

3. Transition from single-orchestrator autonomy to Structured Skill Modes using role-gated multi-agent committees.

  • What the improved AI system can do: The system prevents agentic orchestration failures, such as non-terminating tool-use loops, by enforcing specialized workflows (e.g., a dedicated Risk Officer gate). This provides institutional-grade stability and ensures that the failure of one reasoning step does not cause a total collapse of the decision cycle.

4. Integration of a 3-Dimensional SCOUT stage (Technical Momentum, Core Fundamentals, and Macro Correlation) and a Volatility-Regulated Turnover model.

  • What the improved AI system can do: The system can differentiate between high-conviction alpha signals and noise (such as RSI oscillators or transient volatility). This enables Composure, preventing the agent from overreacting to market turbulence and ensuring it ignores momentum-based distractors in favor of fundamental and macro-correlated drivers.

5. Implementation of a Recompute-Verifiable Hash-Linked Audit Chain and a strict TimeGate data-access protocol.

  • What the improved AI system can do: The system guarantees absolute temporal integrity, making look-ahead bias (information leakage) mathematically impossible. It provides a tamper-proof, forensic-grade audit trail where every observation, tool call, and decision is hashed and linked, allowing for independent, third-party verification of every trade for regulatory compliance.

6. Integration of multi-component, calibrated cost-modeling (transaction, financing, and market-impact costs).

  • What the improved AI system can do: The system optimizes for net-of-cost profitability rather than gross-return maximization. It can accurately predict the drag of leverage, financing, and slippage, preventing the deployment of paper-profitable strategies that would fail in real-world, high-friction execution environments.

Abstract

LLM agents are increasingly cast as autonomous portfolio managers, yet the dominant evaluation idiom, a leaderboard of returns over a fixed window, certifies neither the soundness of an agent's process nor the durability of its edge: one period's return is dominated by the market path, and apparent alpha can dissolve once look-ahead bias and trading costs are controlled. We introduce CLQT, a closed-loop benchmark that reframes LLM trading evaluation as diagnosis rather than ranking. CLQT enforces point-in-time data access through a hard TimeGate, models institutional transaction and financing costs, scores strategy consistency across rounds, and seals every gather-analyze-decide-execute-reflect cycle into a recompute-verifiable audit chain; the same model runs as a constrained investment committee or a single autonomous orchestrator, making scaffolding an experimental variable. From the audit trail CLQT computes a five-axis capability scorecard (Coherence, Acuity, Composure, Discipline, Reliability), with coherence scored partly by a held-out LLM judge to curb self-preference bias. We validate CLQT on a contamination-controlled, year-long multi-model backtest campaign with a 13-configuration ablation grid and a four-week live broker paper-trading track on post-cutoff data. The diagnosis-first read shows that the capability leader is not the Sharpe leader; that agents' allocations systematically fail to follow their own stated analysis, a stating-versus-doing gap stable across both tracks (+0.30 backtest, +0.23 live); and that module value registers on the capability and behavioral axes when returns alone cannot separate it. Net of realistic costs, agents clear defensive baselines but do not cleanly beat the index. Credible evaluation of LLM investment agents must therefore diagnose the process rather than rank a period's return, the standard CLQT operationalizes and makes auditable.

Sources

Related papers