Financial Language Models as Applied Artificial Intelligence Systems for News-Based Trading under Market Frictions
cs.CL, cs.LG, q-fin.TR
Submitted: 2026-09-20
Updated: 2026-09-20
Comments: 47 pages. Revise and resubmit at Engineering Applications of Artificial Intelligence
Code: https://github.com/meta-llama/llama3
License: http://creativecommons.org/licenses/by/4.0/
The gist: Financial language models can transform unstructured firm-specific news into structured decision signals, but financial AI research lacks an integrated deployment framework for evaluating whether
Terminology
Abstract
Financial language models can transform unstructured firm-specific news into structured decision signals, but financial AI research lacks an integrated deployment framework for evaluating whether those signals remain useful in financial decision systems. Computer science research has developed strong methods for time-series forecasting, text classification, multimodal stock prediction, graph-based market modeling, and machine-learning operations, yet these streams do not provide a domain-specific protocol that jointly tests financial language-model outputs under event-time observability, probability calibration, execution timing, transaction costs, liquidity constraints, capacity limits, operational diagnostics, and statistical inference. We introduce MFAST, a Market-Friction-Aware Sentiment-to-Trading framework that converts timestamped financial text into auditable, reproducible, and market-feasible trading decisions. The application is news-based trading, where firm-specific text must be linked to securities before portfolio decisions can be evaluated. The framework links Refinitiv News Analytics to Center for Research in Security Prices (CRSP) equity data, restricts the primary out-of-sample evaluation to post-release news outside disclosed foundation-model data-freshness periods, and adds a public replication arm using open financial text and public price data. Results show that decoder-only language models outperform encoder baselines and dictionary sentiment in classification, calibration, return prediction, and net portfolio performance, while operational diagnostics reveal trade-offs among accuracy, latency, memory, throughput, and inference cost. The paper shows that credible evaluation of financial language models requires an end-to-end engineering approach combining language understanding, temporal discipline, market-friction-aware deployment, and reproducible validation.
Sources
- Understanding intermediate layers using linear classifier probes
- FinBERT: Financial Sentiment Analysis with Pre-trained Language Models
- Longformer: The Long-Document Transformer
- The Llama 3 Herd of Models
- Financial Statement Analysis with Large Language Models
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
- Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models
- BloombergGPT: A Large Language Model for Finance
- PIXIU: A Large Language Model, Instruction Data and Evaluation Benchmark for Finance
- FinGPT: Open-Source Financial Large Language Models
- OPT: Open Pre-trained Transformer Language Models
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering