TradingMoE: Routing the Right Experts in Evolving Markets

arXiv:2608.11785 · cs.LG · Submitted 2026-08-12 · Read on arXiv

Chang Zhou, Xingtong Yu, Minbin Huang, Zhennan Wu, Yuan Fang, Hong Cheng, Xinming Zhang

University of Science and Technology of China · The Chinese University of Hong Kong · The University of Tokyo · Singapore Management University

cs.LG

Submitted: 2026-08-12

Updated: 2026-08-13

License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/

Importance score: 75/100

The gist: TradingMoE: Routing the Right Experts in Evolving Markets proposes a trading-oriented sparse Mixture-of-Experts (MoE) framework that augments a frozen dense LLM with lightweight residual experts for

Terminology

Summary

TradingMoE: Routing the Right Experts in Evolving Markets proposes a trading-oriented sparse Mixture-of-Experts (MoE) framework that augments a frozen dense LLM with lightweight residual experts for direct trading-decision generation.

The paper identifies two key challenges with existing LLM-based trading systems. First, conventional internal MoE routers assign expert scores from token hidden states that are only implicit proxies for expert suitability. The authors find that naive router scores have a Pearson correlation of only −0.015 with the measured gains, and 66.76% of decision tokens leave at least one better expert unselected. Second, these routers receive no direct signal indicating when an inactive expert has become more suitable as market conditions change.

The authors also reveal that token-specific expert usefulness exhibits a compact low-dimensional structure. Specifically, token–expert counterfactual credit matrices exhibit a pronounced low-rank structure: rank-16 reconstructions retain 74.2% and 77.9% of the credit energy for OLMoE and DeepSeek-V2-Lite backbones, respectively, and identify higher-credit experts than native routers on over 98% of decision tokens.

Based on these findings, TradingMoE introduces two key components:

  1. Query–Key router: For each decision token, the router constructs a low-dimensional trading-demand query that represents the expertise required by the token under the current market context. Each expert is associated with a learnable expert key, and routing scores are computed through query–key matching. The complete query–key score matrix has rank at most the query dimension, providing a compact parameterization consistent with the observed low-rank credit structure.

  2. Sparse expert selection update mechanism: At each decision-value token, the method samples a small number of inactive experts as challengers and compares them with the lowest-scored active expert in the current Top-k route. Their first-order relative credits estimate whether a challenger would reduce the trading-decision loss more than the lowest-scored active expert and should therefore replace it. The mechanism uses a detached routing margin and same-step router update that cancels in the forward pass but provides gradients during backpropagation.

The paper provides theoretical justification showing that sampling a small number of inactive experts provides an unbiased estimate of the update averaged over all inactive experts and that the routing-margin update is consistent with the local loss reduction induced by expert replacement.

Experiments are conducted on stock and cryptocurrency markets against 22 baselines from seven families, including tree-based forecasting, neural forecasting, reinforcement learning, financial LLMs, general LLMs, LLM-based trading agents, and external expert-routing methods. Results show that TradingMoE improves cumulative return over the best-performing baselines by 30.89% and 30.7%, respectively, while exceeding the corresponding buy-and-hold benchmarks by 37.81% and 80.17% percentage points. On Stock, it achieves a cumulative return of 49.08% with a Sharpe ratio of 5.091; on Crypto, it reaches 73.79% with a Sharpe ratio of 1.355.

Ablation studies demonstrate that both components are complementary: the query-key router provides more effective token–expert matching, while the sparse expert selection update further refines routing using feedback from sampled inactive experts. Leakage-controlled evaluations using backbones released after the evaluation period confirm that the strong performance does not rely on temporal leakage. Additional analyses cover hyperparameter sensitivity, multi-seed statistical robustness, and transaction-cost sensitivity.

Improvements for AI systems

Improvements to AI systems:

  1. Add a query–key routing layer to frozen dense LLMs that constructs low-dimensional task-specific queries from hidden states and matches them against learnable expert keys, replacing implicit token-hidden-state scoring. This improves expert selection accuracy by directly aligning routing decisions with task demands, as demonstrated by the near-zero correlation of naive routers with actual expert gains.

  2. Implement a sparse challenger-based expert replacement mechanism during inference, where a small set of inactive experts is sampled and compared against the lowest-scoring active expert using first-order credit estimates. This enables dynamic adaptation to changing input contexts (e.g., market regimes) without retraining the base model, correcting the 66.76% of decision tokens where a better expert was missed.

  3. Exploit the low-rank structure of token–expert credit matrices by parameterizing the router’s score matrix to have rank at most the query dimension (e.g., rank-16). This reduces routing overhead and improves generalization, as rank-16 reconstructions retain over 74% of credit energy, allowing the system to approximate optimal expert assignment with far fewer parameters.

  4. Add a detached routing-margin update that provides gradients during backpropagation without altering forward-pass routing. This allows the router to learn from counterfactual expert replacements in a stable, unbiased manner, improving long-term decision quality without destabilizing the frozen backbone.

What the improved AI system can do:

  • Directly generate high-quality decisions (e.g., trading actions) from raw sequential data (prices, volumes) using a frozen LLM augmented with lightweight residual experts, without fine-tuning the base model.

  • Adaptively re-route experts in real time as input distributions shift, by sampling challengers and swapping underperforming active experts—enabling robust performance in non-stationary environments (e.g., financial markets).

  • Achieve higher cumulative returns and Sharpe ratios than specialized forecasting, RL, and LLM-agent baselines, with 30.89% and 30.7% improvements over the best baselines on stock and crypto, respectively, while maintaining computational efficiency via low-rank routing.

  • Operate without temporal leakage even when the backbone LLM was released after the training data period, ensuring the system’s decisions are based on genuine pattern recognition rather than memorized future information.

  • Provide interpretable routing decisions by explicitly linking each decision token’s trading-demand query to expert keys, enabling analysis of which expert capabilities are triggered under specific market conditions.

Sources

Related papers