Large Language Model-Driven Small-Capitalization Trading: Integrating Financial News Sentiment, Macroeconomic Indicators, and Technical Signals

arXiv:2608.12283 · q-fin.PM, cs.CL · Submitted 2026-08-12 · Read on arXiv

Tailstate Intelligence Ltd. · Independent Researcher · Zanista AI Ltd.

q-fin.PM, cs.CL

Submitted: 2026-08-12

Updated: 2026-09-30

License: http://creativecommons.org/licenses/by/4.0/

Importance score: 75/100

The gist: This paper studies small-capitalization trading with LLM-derived news sentiment, macroeconomic indicators, technical signals, and uncertainty-aware portfolio construction.

Terminology

Summary

This paper studies small-capitalization trading with LLM-derived news sentiment, macroeconomic indicators, technical signals, and uncertainty-aware portfolio construction. The central result is that the way the tradable set is defined matters at least as much as the sentiment backend or allocator. Separating the macro-beta trigger into pure alpha and pure beta usually produces stronger Sharpe and return than requiring the stock-side and indicator-side triggers to fire together.

The horizon pattern is economically interpretable. Pure beta works best when the signal is a cross-asset transmission problem: at very short horizons, liquid macro or sector instruments can lead slower small-cap constituents; at 20–40 trading days, aggregate shocks can continue to be mapped into heterogeneous firm fundamentals. Pure alpha works best when the signal is firm-specific and slower to diffuse, especially at 5, 10, and 60 trading days. These results suggest that LLM news sentiment is most useful when embedded in a broader trading architecture that controls the opportunity set, distinguishes macro and idiosyncratic channels, and passes predicted risk into portfolio construction rather than using sentiment only as an expected-return overlay.

The paper addresses two gaps in prior work. First, most studies score isolated headlines rather than related news stories, even though repeated coverage can overweight a single event. Second, sentiment is usually treated as an expected-return input while portfolio risk remains historical and signal-independent. The authors address both gaps with a pipeline that clusters related articles into representative news-summary sentiment features, combines them with macroeconomic and technical indicators, and predicts both expected returns and covariance for portfolio construction. To limit look-ahead bias, sentiment is scored with models whose knowledge cutoffs precede the 2025 test period.

The methodology proceeds as follows. Stock selection compares three ways to define the traded cross-section, each decomposing a macro-beta trigger into a separate event set. A stock's own movement is flagged via its return z-score, and its relationship to a given indicator is measured via beta computed over a rolling window against a panel of macroeconomic releases, commodities, currencies, rates, volatility, world-index, and sector-ETF indicators. The three selection legs are pure alpha (SS SI), pure beta (SI SS), and beta (SI ∩ SS). Pure alpha isolates firm-specific abnormal moves not contemporaneously explained by macro indicators. Pure beta is the anticipatory macro leg: an indicator has moved for a levered stock, but the stock itself has not yet registered its own abnormal move. Beta is the confirmatory leg, where stock and indicator triggers fire together and agree on direction.

For news sentiment, the paper treats sentiment as a property of a story, not an article. Within a trailing 30-day window, article summaries are embedded and clustered by single-linkage agglomerative clustering using cosine distance, merging two articles into the same story once their similarity reaches 0.90. Each resulting cluster is treated as one underlying story, and only the article nearest to the cluster centroid is retained as the representative news item. The daily sentiment feature is computed from these representative articles and paired with a log-scaled story count as a coverage-volume feature. The sentiment model returns a calibrated probability distribution over negative, neutral, and positive outcomes, transformed into a final normalized sentiment surprise through entity-prior correction, strictly trailing rolling demeaning, and group-day cross-sectional standardization. GPT-4o mini is the primary scorer, with FinBERT, Mistral 7B Instruct, and Llama 2 13B Chat as alternatives.

For joint return and risk prediction, the sentiment signal feeds a multimodal network alongside daily price- and volume-derived technical features. The model predicts the full conditional return distribution – its mean and covariance – rather than return alone. Dropout applied to the representation gives a stochastic representation for multiple forward passes; a mean head and covariance head operate on each draw, with the covariance head parameterized as rank-2 low-rank plus diagonal. The aleatoric term captures the market's own conditional randomness, while epistemic uncertainty is captured by keeping dropout active at inference and treating the spread across stochastic passes as confidence in the forecast. Both are combined into a single predictive covariance. The model is trained by maximizing the likelihood of realized returns under its own predicted aleatoric distribution, Gaussian or Student-t, so risk is learned jointly with return rather than fit afterward from residuals.

For portfolio construction, the combined covariance is fed directly into a standard mean-variance optimizer with transaction-cost regularization, full investment, and a per-asset position cap of 40%, at a risk-aversion setting of δ = 2.5. All reported filtered portfolios are long-only. This is benchmarked against five standard alternatives: equal-weighted, risk parity, hierarchical risk parity, and Black–Litterman and Bayesian Black–Litterman.

The evaluation uses Russell 2000 equities with daily OHLCV data from October 2, 2023 through December 31, 2025 and scored news from October 1, 2023 through December 31, 2025. All universe selection, standardization, model fitting, and early stopping use only the in-sample period ending December 31, 2024. The 2025 calendar year is held out for backtesting. The benchmark grid covers three selection regimes, four sentiment backends, two target distributions, eight holding periods (1, 2, 3, 5, 10, 20, 40, 60 trading days), eight transaction-cost scenarios (0 to 100 basis points), and six allocation methods.

The main results are as follows. At 100 bps transaction costs, pure alpha leads at 5, 10, and 60 trading days; pure beta leads at 20 and 40 days. The one-day result is cost-sensitive: pure beta is strongest at low costs, but the edge disappears under the 100 bps stress case. The strongest conservative row is pure beta with GPT-4o mini sentiment, a Student-t target, a 40-day holding period, and risk parity allocation, reaching Sharpe 2.33 at 100 bps. At 60 days, pure alpha remains the strongest filtered regime from zero costs through 100 bps, with its lead over pure beta and the beta intersection widening as costs rise. The beta intersection is generally weaker because requiring both channels to fire removes two useful cases: early macro spillovers before the stock has reacted, and idiosyncratic stock events without a broad factor shock.

Several limitations qualify the results. GPT-4o mini's October 2023 knowledge cutoff helps limit look-ahead bias but does not eliminate distraction from general pre-existing company knowledge, and prevents use of newer models. Long articles are summarized before scoring, so any nuance lost during summarization cannot be recovered downstream. The evaluation is not a live deployment and does not measure real-time retrieval, summarization, scoring latency, or timestamp precision within the trading day. The horizon explanations are financial interpretations of a one-year out-of-sample benchmark, not causal identification of the underlying news events. A longer live sample, formal multiple-testing adjustment, and event-level attribution would be needed before treating any single horizon–selection pair as a persistent anomaly.

Improvements for AI systems

Improvements to AI Systems:

  1. Event-Clustered Sentiment Scoring: Implement a news-processing module that clusters related articles into single underlying stories (using embedding similarity with a 0.90 merge threshold) and scores only the centroid-representative article. This reduces overweighting of repeated coverage and produces more accurate, event-level sentiment signals.

  2. Dual-Channel Alpha/Beta Decomposition: Build a stock-selection layer that explicitly separates firm-specific (pure alpha) from macro-anticipatory (pure beta) signals using rolling beta against a panel of macro indicators. The AI can then dynamically switch between channels based on holding horizon—using pure beta for 20–40 day macro transmission and pure alpha for 5/10/60 day idiosyncratic diffusion—improving risk-adjusted returns.

  3. Joint Return-Covariance Prediction with Uncertainty: Train a multimodal network that outputs both mean and covariance of returns, using dropout at inference for epistemic uncertainty and a rank-2 low-rank plus diagonal covariance for aleatoric uncertainty. This replaces historical risk estimates with signal-dependent, forward-looking risk, enabling more robust portfolio optimization.

  4. Horizon-Aware Portfolio Construction: Integrate a horizon selector that maps predicted holding periods to optimal selection regimes (e.g., pure beta for 20–40 days, pure alpha for 5/10/60 days) and feeds the combined predictive covariance into a mean-variance optimizer with transaction-cost regularization, position caps, and risk aversion. This yields higher Sharpe ratios (e.g., 2.33 at 100 bps costs) than static allocation.

  5. Cost-Sensitive Regime Switching: Add a transaction-cost-aware decision layer that adjusts the selection regime based on cost levels—e.g., favoring pure beta at low costs for short horizons but switching to pure alpha under high costs—preventing performance degradation from overtrading.

  6. Look-Ahead-Bias Mitigation: Use sentiment models with knowledge cutoffs preceding the test period and apply strictly trailing rolling demeaning and group-day cross-sectional standardization to ensure signals are causal and not contaminated by future information.

What the Improved AI System Can Do:

  • Process financial news as coherent stories rather than isolated headlines, reducing noise and event duplication.

  • Distinguish between macro-driven and firm-specific price movements, and automatically select the appropriate trading horizon (1–60 days) and selection regime for maximum Sharpe ratio.

  • Predict both expected returns and forward-looking covariance simultaneously, incorporating model uncertainty (epistemic) and market randomness (aleatoric) into portfolio weights.

  • Dynamically adapt to transaction costs, shifting from beta-led to alpha-led strategies as costs rise, preserving profitability under realistic trading conditions.

  • Generate long-only, fully invested portfolios with per-asset caps, optimized via mean-variance with risk aversion, outperforming equal-weight, risk parity, and Black–Litterman benchmarks in out-of-sample backtests.

  • Operate with minimal look-ahead bias, making it suitable for deployment in live trading environments where timeliness and causal signals are critical.

Abstract

Large language models can extract richer signals from financial news than fixed sentiment lexicons, and recent work has explored feeding such signals into portfolio construction. We study an uncertainty-aware construction that feeds model-predicted risk -- decomposed into aleatoric and epistemic components -- directly into the covariance matrix of portfolio allocators, rather than treating portfolio risk as fixed or adjusting only expected returns. We evaluate the pipeline on Russell 2000 equities under three stock-selection regimes: a pure-alpha trigger that isolates abnormal stock moves not explained by macro indicators, a pure-beta trigger that captures macro-indicator moves before the stock itself fires, and a beta trigger in which both channels agree. Across the full holding-period grid, the separated pure-alpha and pure-beta legs usually dominate the beta intersection on Sharpe and return. Two horizons are especially informative. At one day, pure beta can work under low and moderate transaction costs because it captures immediate lead-lag spillovers from liquid macro and sector indicators into exposed small-cap stocks, but this advantage disappears at 100 bps when turnover and microstructure noise dominate. At 40 days, pure beta works for a different reason: slower macro repricing overtakes the firm-specific pure-alpha channel. The strongest conservative row is pure beta with GPT-4o mini sentiment, a Student-t target, a 40-day holding period, and risk parity allocation, reaching Sharpe 2.33 at 100 bps. The results suggest that stock-selection regime and allocator choice matter at least as much as the sentiment model, and that separating firm-specific and macro-exposure triggers is more informative than requiring both to fire simultaneously.

Sources

Related papers