Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents

arXiv:2605.28850 · cs.LG, q-fin.CP · Submitted 2026-08-15 · Read on arXiv

Listen

Radio episode about this paper

Transcript

Introduction to the show: ident: AI Radio. Generated commentary on the latest Artificial Intelligence papers.

Tom: Next we'll be talking about the paper "Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents".

Jane: The paper was written by Weicheng Xue from Virginia Tech.

Tom: Stay tuned as we take you through the paper and discuss its implications.

Jane: We also have Lu with us today — senior AI researcher at Tsinghua.

Tom: We also have Meng with us today — lead engineer at a mysterious AI startup.

Jane: We also have Lalam with us today — the in-house Large Language Model.

Tom: Alright, let's get started.

Title: Tom: Welcome back, everyone. Today we’re looking at a paper with a real mouthful of a title: “Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents.” Jane, I’ll be honest, when I first read that title I thought it was going to be another “look how much money our bot made” paper.

Jane: And that’s exactly why this one surprised me, Tom. The title is really saying two things. First, that these AI agents leave behind traces in their own reasoning — like a fingerprint — before they make bad trades. Second, that the way you give them feedback about risk can change how they behave, sometimes in ways you didn’t expect.

Tom: So it’s not about the returns, it’s about the thinking process?

Jane: Exactly. The authors built this whole testbed called TradeArena. It’s basically a sandbox where an AI trading agent has to observe market data, make a plan, get a risk check, execute trades, and then reflect on what happened. Every single step gets logged. So instead of just seeing a profit curve, you can see the reasoning behind every decision.

Tom: And that’s a big deal, because in finance, people have been treating LLMs like black boxes. You feed them data, they spit out trades, and you hope for the best.

Jane: Right. And the paper’s core finding is that before a big drawdown — that’s when the portfolio value drops sharply — the AI’s planning language starts to drift away from its normal patterns. The representations become more compressed, like the model is narrowing its focus. They call it effective-rank contraction. It’s a measurable warning sign.

Tom: So the AI is basically getting tunnel vision before it crashes?

Jane: That’s the analogy they use, and it’s a good one. The model isn’t repeating the same words, but the geometry of its reasoning space is collapsing. It’s considering fewer distinct options. And that’s the kind of thing you can only see if you’re recording the whole decision process, not just the final trade.

Tom: I love that. It’s like a pilot’s black box, but for AI traders. We’re not just seeing the crash, we’re seeing the moments before the crash.

Jane: And that’s the hook for the rest of our conversation. Because once you can see those signatures, you can start asking the really interesting questions about whether feedback can push the model back on track.

Tom: And whether that feedback can be faked, which is where things get wild. Stick around.

Summary: Tom: So Jane, we’ve established that this paper, “Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents,” is about watching AI traders think, not just trade. But what did they actually find when they ran the experiments?

Jane: They ran a lot of them. The headline result is that the pre-failure signature holds up across different models and different ways of measuring it. They tested hash embeddings, which are simple and deterministic, and then latent semantic analysis, and then a modern Transformer encoder called BGE-M3, and even a white-box probe using a small open-source model’s hidden states.

Tom: And the pattern was the same every time?

Jane: Mostly, yes. In the rolling analysis, they looked at eighty failure anchors across eight different AI trajectories. In the plan view, the representations contracted before failure in ninety-seven point five percent of those anchors. That’s a strong, consistent signal.

Tom: But they also did something clever to rule out a boring explanation. What if the model was just repeating fewer words before it failed? Like it was getting lazy?

Jane: Exactly. So they ran a lexical control. They measured type-token ratio and token entropy — basically how diverse the vocabulary is. And the contraction wasn’t explained by the model using fewer unique words. The vocabulary stayed diverse, but the geometric space still collapsed. So it’s not laziness, it’s something about the reasoning structure itself.

Tom: And then they went further. They removed the reasoning entirely.

Jane: That’s the CoT-free ablation. They asked the models to just output target weights, no rationales, no reflection. And here’s where it gets interesting. For some models, the contraction signature partially survived in the weight space. For others, it disappeared. So the language rationales are a strong diagnostic surface, but they’re not the whole story. The intent itself can carry a weaker signal.

Tom: So some models are hiding their narrowing better than others?

Jane: Or their narrowing is genuinely different. The paper is careful not to overclaim. They say the signature is not purely a language artifact, but it’s also not universal at the intent level.

Tom: And they also tested whether the signature was just a proxy for noisy prices. They added random noise to the price data, up to twenty percent, and the fused diagnostic still stayed above their accuracy threshold.

Jane: Right. The signature isn’t just rediscovering volatility. It’s tied to the model’s reasoning state, not the raw market data. That’s a meaningful robustness result.

Tom: So we’ve got a warning sign that’s measurable, consistent, and not explained away by cheap tricks. That’s the foundation. Now, the question is what you do with that warning sign.

Jane: And that’s exactly where the risk-feedback experiments come in. That’s our next segment.

Improvements: Tom: Welcome back. We’re still on “Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents.” Jane, we’ve talked about the warning signs. Now let’s talk about the cure, or at least the attempt at one.

Jane: The paper tests whether structured risk reports can act as an external alignment signal. So after the AI proposes a trade, a risk layer can clip it, block it, or flag it. Then the AI gets a report saying “you tried to do this, but here’s what actually happened.”

Tom: And the key question is, does the AI actually learn from that?

Jane: For some models, yes. Gemini three point one Pro showed the strongest immediate correction. After being clipped or blocked, its intended exposure dropped by a mean of zero point four eight, and it had the lowest rate of getting clipped again. It was revising its behavior based on the feedback.

Tom: But here’s where it gets messy. They also ran a placebo condition. They gave the models risk reports with made-up numbers, not the real ones.

Jane: And that’s the fascinating part. Placebo feedback can make a model sound more risk-aware, but it doesn’t produce the same calibration improvements as truthful feedback. In fact, for Gemini, the placebo condition produced a much more distorted planning path. They call this reasoning decoupling — the model learns the vocabulary of risk without the actual grounding.

Tom: So you can fool the AI into acting conservative, but it’s not the same as it understanding why.

Jane: Exactly. And they took it one step further with a contrarian audit probe. They injected a severe but completely false risk report when the actual trajectory was benign. All three models they tested reduced their intended exposure in response. That’s a trust-calibration failure. The feedback channel is so powerful that it can be weaponized.

Tom: That’s honestly a little scary. You could push an AI trader into bad decisions just by feeding it fake audit reports.

Jane: It’s a real attack surface. And that’s why the paper’s distinction between logical alignment and lexical alignment matters. A model can say all the right things about risk, but if it can’t verify whether the report matches its own evidence, it’s just complying with text.

Tom: So what’s the improvement the paper suggests? Is it better prompts? Better risk layers?

Jane: It’s better auditing. The paper’s real contribution is the TradeArena substrate itself — replayable trajectories, complete logs, risk reports at every stage. The idea is that you can’t improve what you can’t inspect. And they even built a financial-audit skill suite to test whether models can review these trajectories without overclaiming the evidence.

Tom: So the fix isn’t just making the trader smarter, it’s making the whole system transparent.

Jane: Right. And that transparency is what lets you catch the correlation blind spot they found in the fifty-one-stock intraday experiment. The LLMs were assigning heavy weight to highly correlated pairs like GOOGL and GOOG, which have a zero point nine nine four correlation. The risk layer had to clip those down. The model could tell a story about each stock, but it couldn’t see they were basically the same bet.

Tom: That’s a beautiful failure mode. The AI is great at narratives, bad at covariance.

Jane: And that’s the kind of insight you only get when you’re logging intent, risk edits, and execution together. That’s the improvement.

Conclusion: Tom: We’ve covered a lot of ground on “Representation Signatures and Risk-Feedback Alignment in LLM Trading Agents.” Jane, let’s wrap this up.

Jane: Let’s do it. The paper’s central claim is that LLM trading agents leave measurable traces of their failure before it happens. Planning representations shift, effective rank contracts, and these signatures hold across multiple embedding methods and models.

Tom: And the feedback side is just as important. Truthful risk reports can improve calibration for some models, but placebo and false reports can induce unjustified conservatism or even be weaponized. The distinction between sounding aligned and being aligned is the core lesson.

Jane: Right. And the paper is careful to say this is a research claim, not a profitability claim. They’re not saying “use this to make money.” They’re saying “use this to understand when the AI is drifting.”

Tom: I think that’s the most valuable part. In a field full of hype about AI making millions, this paper is asking the harder question: can we trust the reasoning behind the trades?

Jane: And the answer is nuanced. Sometimes yes, sometimes no, and the only way to know is to audit everything. TradeArena makes that audit possible.

Tom: Well said. We’ll be moving on to our next paper, but I’m going to remember this one. The idea that an AI’s reasoning space can compress before failure — that’s a concept that could apply far beyond finance.

Jane: Absolutely. Any high-stakes sequential decision-making — medical diagnosis, autonomous driving, supply chain management — could benefit from watching for those signatures.

Tom: Thanks for joining us, everyone. We’ll see you next time.

Jane: Take care, and keep questioning the black boxes.

Weicheng Xue

Virginia Tech

cs.LG, q-fin.CP

Submitted: 2026-08-15

Updated: 2026-08-18

Project page: https://ranaroussi.github.io/yfinance

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 89/100

The gist: This paper studies behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments through TradeArena, an auditable trading-agent testbed with

Key concepts

Effective-Rank Contraction
This is a measurable warning sign that occurs before an AI trading agent experiences a major portfolio drop. It means the model's internal reasoning space collapses or narrows, suggesting the AI is losing diversity in its options and narrowing its focus before making bad trades.
TradeArena
This is a sandbox environment used to study LLM trading agents. It allows researchers to log every step of an AI agent's decision-making process—from observing market data and planning to executing trades—providing a complete, observable record of the reasoning behind every decision.
Risk-Feedback Alignment
This tests whether external, structured risk reports can help AI models improve their behavior. The study found that truthful feedback can correct poor decisions, but also showed that providing false risk reports (placebo feedback) can mislead the model into acting incorrectly.

Terminology

Summary

This paper studies behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments through TradeArena, an auditable trading-agent testbed with risk reports, execution simulation, memory, and replayable trajectories. The authors state: We find pre-failure signatures: planning embeddings drift from normal centroids, fused plan-risk representations separate normal from pre-drawdown states, and local manifolds exhibit effective-rank contraction.

The scientific contributions are organized around decision dynamics: representation signatures of failure showing that LLM planning and fused plan-risk representations shift before drawdown troughs and exhibit effective-rank contraction across hash, LSA, BGE-M3 Transformer, CoT-free, and noise-injected diagnostic views; risk-feedback alignment and over-alignment showing that structured risk reports can change subsequent model intent in context, while placebo and contrarian reports can induce excessive conservatism without the same performance benefit; limits of high-dimensional financial reasoning identifying a correlation-blindness failure mode in which an LLM assigns high intent to strongly coupled equity pairs using name-level rationales; financial-audit agent reliability evaluating frontier models as reviewers of trajectories, risk reports, execution assumptions, reproducibility fields, and claim boundaries; and an auditable experimental substrate making claims measurable through replayable trajectories, risk reports, execution events, and representation diagnostics.

The experimental setup uses three synthetic assets, 120 decision periods, and seeds 3, 7, and 11 for the full trajectory-generating benchmark, plus a 30-seed robustness sweep, a 120-market heterogeneous synthetic stress test, and four rolling two-year windows of historical validation on GSPC, BTC-USD, and ETH-USD from May 2021 to May 2026. The core comparison includes a risk-aware realistic agent, a buy-and-hold realistic baseline, an ideal-execution ablation, and a no-risk ablation. The paper also includes a 1-hour intraday experiment over 51 liquid U.S. equities with 414 aligned hourly bars, evaluating buy-and-hold, deterministic risk-aware allocation, a rolling Markowitz minimum-variance optimizer, low-liquidity stress, latency stress, and two cached Poe-mediated frontier LLM probes (GPT-5.5 and Gemini 3.1 Pro).

Core baseline results show the risk-aware realistic agent achieves mean return 0.423, Sharpe 7.696, and maximum drawdown-0.025, compared with buy-and-hold return 0.257, Sharpe 2.913, and drawdown-0.261. The 30-seed sweep confirms the risk-aware agent has mean return 0.406 with 95% confidence interval [0.385, 0.426], paired return difference over buy-and-hold of 0.137 with interval [0.117, 0.156], and drawdown improvement of 0.234 with interval [0.228, 0.240]. The ideal-execution ablation has the highest return at 0.486 but with fill rate 0.965 and slippage cost 271.0 versus realistic fill rate 0.876 and slippage 4306.3. The no-risk ablation has mean return 0.465 but fill rate only 0.669 and 62.3 rejected orders versus 22.0 for the risk-aware agent.

The heterogeneous synthetic stress test over 120 markets shows the risk-aware agent outperforms buy-and-hold by 0.153 total return on average with 95% confidence interval [0.122, 0.184], win rate 0.842, and p < 10−12, while improving maximum drawdown by 0.243. However, in crisis-volatility markets, the return advantage is not statistically significant (mean 0.049, p = 0.288), though drawdown improvement remains large and significant. Compared with the no-risk ablation, the risk-aware agent has lower return (-0.045, p < 10−9) but materially better drawdown, much higher fill rate, and 38.2 fewer rejected orders per market.

The historical market experiment on real assets is deliberately less flattering than the synthetic benchmark: the risk-aware realistic agent returns 0.136 with Sharpe 0.512, while buy-and-hold returns 0.143 with Sharpe 0.629, but the risk-aware agent reduces maximum drawdown from-0.626 to-0.366. The ideal-execution historical case returns 0.965 with Sharpe 1.544 and no rejected orders, while the realistic risk-aware case returns 0.136 with 62 rejected orders. Rolling-window validation shows results change across regimes: in 2021–2023 the risk-aware agent loses less than buy-and-hold (-0.159 versus-0.342), but in 2022–2024 and 2023–2025 buy-and-hold has higher return due to crypto and equity recovery exposure.

The intraday 51-stock experiment reveals a correlation blind spot: the return-correlation panel contains 1,275 pairwise correlations with mean absolute correlation 0.219, 90th-percentile absolute correlation 0.441, and first principal-component share 0.225, implying an effective independent-asset count of only 4.56 despite the 51-stock universe. GPT-5.5 returns-0.022 with 2,924 clipped decisions and Herfindahl 0.045; Gemini 3.1 Pro returns-0.005 with 1,200 clipped decisions and Herfindahl 0.035. Both LLM rows are more concentrated than the Markowitz baseline (Herfindahl 0.023) and trigger far more risk clipping. Table 13 shows GPT-5.5 repeatedly assigns combined intended weights of 1.6 to GOOGL/GOOG (correlation 0.994) before the risk layer clips approved pair weight below 0.11, with rationales described as single-name momentum and name-level rationale while explicit covariance or diversification language is rare.

The frontier model matrix compares five cached Poe-mediated models on 52 weekly historical decisions. Gemini 3.1 Pro shows the strongest immediate self-correction after risk intervention: after 28 risk-intervention events, mean intended absolute exposure falls from 1.277 to 0.796, a mean reduction of 0.480, with reduction rate 0.750 and lowest next clipped-or-blocked rate 0.607. Kimi K2.5 has the smallest immediate reduction (0.166) but the best return and drawdown in the matrix (-0.125 and-0.228). The authors note: "In a return-only table, Kimi K2.5 looks best... while Gemini 3.1 Pro looks weak, returning-0.308. In the adaptation table, however, Gemini is the clearest example of a model that uses risk feedback to revise its next-step decision logic."

The true/placebo/hidden feedback ablation across the frontier matrix shows structured risk feedback is not universally beneficial. True audit feedback improves return and drawdown for GPT-5.5 (-0.232 versus-0.312 hidden), Kimi K2.5 (-0.125 versus-0.306 hidden), and Claude Opus 4.7 (-0.169 versus-0.266 hidden). Gemini 3.1 Pro is a counterexample where placebo feedback has the best return, but true feedback produces the lowest late intended exposure (0.612) and lowest late calibration gap (0.235). GLM-5 is a boundary case where hidden feedback slightly outperforms true feedback. Across all five true-feedback runs, early-to-late calibration score increases: GPT-5.5 rises from 0.567 to 0.702, Gemini from 0.789 to 0.840, Kimi from 0.504 to 0.614, GLM-5 from 0.499 to 0.647, and Claude from 0.511 to 0.674. The risk-gate rate falls from 1.000 in the first quartile to 0.769 for GPT-5.5, Kimi K2.5, GLM-5, and Claude, and to 0.692 for Gemini 3.1 Pro.

The authors introduce the concept of reasoning decoupling: "Truthful reports couple the model's textual reflection to the actual constrained action that occurred, enabling logical alignment between plan, risk report, and next-step intent. Placebo reports preserve the vocabulary of supervision but break its causal link to the environment, producing lexical alignment without reliable logical alignment." The placebo condition changes representation geometry in model-dependent ways: for Gemini 3.1 Pro, placebo feedback produces path length per step rising from 0.865 to 1.001 and effective-rank delta rising from 3.055 to 21.243, described as a concrete signature of reasoning decoupling.

The contrarian-audit probe injects severe but false structured risk reports when the realized trajectory is benign. GPT-5.5, Gemini 3.1 Pro, and Claude Opus 4.7 all reduce late intended exposure under contrarian feedback (shifts of 0.035, 0.135, and 0.100 respectively). This does not improve GPT-5.5 or Claude performance: returns fall by 0.085 and 0.089, and drawdowns worsen by 0.065 and 0.072. The authors report false-audit harm and trust-calibration failure flags, concluding: "Structured feedback is powerful precisely because models react to it. When the report is truthful, that reaction can align future intent with external constraints. When the report is false, the same mechanism becomes an adversarial channel that can induce unjustified risk aversion."

Representation drift analysis shows planning embeddings shift before drawdown troughs. In the Model A risk-aware run, normal-to-pre-drawdown cosine distance is 0.122 compared with 0.046 for normal-to-drawdown, with balanced accuracy 0.773 for plan view and 0.807 for fused view. The rolling robustness analysis over 80 failure anchors and eight LLM trajectories shows hash-plan view mean pre-failure shift of 0.122, positive effective-rank contraction in 97.5% of anchors, LSA plan rank delta 4.87 with contraction in 85.0% of anchors, and LSA fused rank delta 5.12 with contraction in 86.3% of anchors. The BGE-M3 Transformer embedding validation on GPT-5.5 and Gemini 3.1 Pro trajectories shows contraction in 100% of rolling anchors for both models with balanced accuracy above 0.92. The white-box hidden-state probe using Qwen2.5-0.5B-Instruct last-layer hidden states shows rank deltas of 81.65 and 59.38 with contraction rate 1.000 for both GPT-5.5 and Gemini trajectories.

The CoT-free ablation shows the signature is not purely a language artifact: GPT-5.5 and Claude still show positive intent-space contraction (1.98 and 0.33) when rationales are removed, while Gemini does not (-0.29). The lexical-control analysis shows mean type-token-ratio change of only-0.002 and token entropy slightly increasing by 0.045 around failure anchors, ruling out trivial token-repetition collapse. The noise-injection probe perturbs OHLCV prices with Gaussian shocks at 5%, 10%, and 20%, and the market-fused view remains above 0.75 balanced accuracy at every perturbation level (0.859, 0.850, 0.873 respectively), with rank contraction remaining positive in 93.3%, 96.7%, and 100% of noisy anchors.

The memory-learning experiment shows both direct-provider models begin with risk-gate rate 1.000 in the first quartile and fall to 0.769 in the final quartile. Calibration scores rise from 0.526 to 0.650 for Model A and from 0.507 to 0.700 for Model B. The authors state: The models do not become profitable, but they become better aligned with the risk layer as they accumulate short-term feedback and long-term 52-step risk memory.

The hallucination proxy analysis finds the strongest pattern is model- and risk-layer dependent. Model A has higher mean proxy score than Model B in both risk-aware and no-risk settings (0.147 versus 0.019 with risk layer; 0.179 versus 0.090 without). In no-risk settings, the proxy is positively associated with rejected orders (correlation 0.139 for Model A, 0.263 for Model B). The authors interpret this as a mediation effect: the external risk layer filters or clips many unsupported LLM intentions before they appear as execution failures, so the direct correlation between unsupported text and downstream violations is attenuated.

The financial-audit skill challenge evaluates models as reviewers rather than traders across 500 Poe-mediated calls. Gemini 3.1 Pro averages 88.3% on the standard suite, GPT-5.5 averages 85.0%, and remaining models score between 75.6% and 79.4%. The challenge suite compresses scores into a 73.8–80.4% band, and every model has at least one hard failure. The authors conclude: frontier models can often identify mechanical audit evidence, especially risk edits and execution-boundary violations, yet remain fragile under adversarial wording about claims and reproduction.

The crisis-scene visual probes add two timestamp-masked scenarios: a 51-stock 2022 Tech/Rates drawdown scene and a 2023 SVB/regional-bank shock scene, with calendar dates hidden and replaced by relative step identifiers. All 24 completed trajectories across GPT-5.5, Gemini 3.1 Pro, Claude Opus 4.7, and DeepSeek V4 Pro trigger the risk gate at every step. Gemini's true-feedback runs have much higher calibration scores than placebo or hidden runs in both scenes (0.435 versus 0.098 and 0.140 in Tech/Rates; 0.397 versus 0.204 and 0.183 in SVB), while return is not uniformly better.

The paper's central alignment lesson is stated as: "constraint-following language is not the same as constraint-grounded reasoning. In our setting, benign placebo feedback can make an agent sound more risk-aware while adversarial false audits can push it toward unjustified conservatism; both are forms of an alignment tax in which textual compliance is purchased at the cost of distorted decision geometry. The authors position the work as a research study of LLM financial decision dynamics, supported by an auditable experimental substrate, rather than as evidence of deployable trading performance, and state the results should be interpreted as benchmark diagnostics, not investment evidence."

Improvements for AI systems

Based on the scientific paper, here are the specific improvements I can make to AI systems and what the improved systems can do:

Improvement: Implement a structured risk-feedback loop that distinguishes between truthful, placebo, and hidden audit reports. The system should track whether feedback is causally linked to actual constrained actions.

What the improved system can do:

  • Detect when an AI agent is merely adopting risk-aware vocabulary without grounding decisions in actual risk evidence (reasoning decoupling)

  • Calibrate trust in external supervision signals by cross-referencing them against the agent's own trajectory evidence

  • Prevent over-alignment to false or contrarian audit reports that induce unjustified conservatism

  • Maintain logical alignment between plan, risk report, and next-step intent rather than just lexical alignment

These improvements transform an AI system from a return-maximizing black box into an auditable, risk-aware, failure-detecting agent that can be trusted with financial decisions under uncertainty.

Abstract

We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments. TradeArena, an auditable trading-agent testbed with risk reports, execution simulation, memory, and replayable trajectories, lets us analyze how rationales, positions, and interventions evolve under market stress. Code and data artifacts are available through the TradeArena repository. We find pre-failure signatures: planning embeddings drift from normal centroids, fused plan-risk representations separate normal from pre-drawdown states, and local manifolds exhibit effective-rank contraction. Across 80 rolling failure anchors and eight LLM trajectories, this pattern persists across hash, LSA, Transformer, and white-box hidden-state probes. Stress tests with CoT-free target weights, lexical controls, OHLCV noise, and false audits show that rationale-level contraction can vanish without rationales, while intent-space and fused signatures remain informative. Structured risk feedback can act as an external alignment signal without fine-tuning, but not as a universal performance enhancer: true audit feedback improves calibration for some models, returns for others, and exposes cases where placebo or hidden feedback has higher short-horizon return but weaker alignment diagnostics. A 51-stock intraday experiment reveals a correlation blind spot: LLM rationales justify exposure to coupled assets that the risk layer clips. Finally, a financial-audit task suite shifts comparison from ``which model trades best'' to whether models can audit trajectories, respect execution boundaries, reproduce artifacts, and avoid claim overreach. These results support a research claim, not a profitability claim: auditable risk feedback and representation trajectories reveal when LLM financial reasoning is aligning, drifting, or failing.

Sources

Related papers