LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification
Michael Schlee, Fabian Lukassen, Christoph Weisser
Georg-August-Universität Göttingen · Hochschule Bielefeld - University of Applied Sciences and Arts
cs.CL
Submitted: 2026-08-12
Updated: 2026-08-13
License: http://arxiv.org/licenses/nonexclusive-distrib/1.0/
Importance score: 75/100
The gist: LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification Abstract Summary: The paper investigates whether financial time
Terminology
Summary
LabelFusion-TS: Fusing Large Language Models, Transformer Encoders, and Financial Time Series for Monetary-Policy Stance Classification
Abstract Summary:
The paper investigates whether financial time series are useful as an additional input for classifying sentences from Federal Reserve communication as hawkish, dovish, or neutral. The proposed system, LabelFusion-TS, extends the LabelFusion architecture by adding a market time-series modality. It combines three independently trained components: a fine-tuned RoBERTa encoder, a prompted large language model (LLM), and a fused ensemble of time-series transformers over market series from the months preceding publication. With only about a thousand annotated sentences for training, the RoBERTa encoder is pre-trained on LLM-annotated sentences and then fine-tuned on human labels. Trained on FOMC communication up to 2015 and evaluated on 2015–2022, the fused system achieves 70.2% weighted F1, compared to 64.1% for the zero-shot LLM, and overtakes it with as few as 240 human-labelled sentences. This provides initial evidence for market time series as an input modality in financial text classification.
Introduction Summary:
Financial language is interpreted within a quantitative environment, where readers assess central-bank statements with recent interest rates, inflation, and asset prices in view. Identical wording can support different readings depending on that background. Market data are public, machine-readable, available at daily or finer resolution over several decades, and aligned in time with any dated document. However, classification models in financial NLP operate almost exclusively on text. Where text and numerical data are combined, the direction is typically the reverse—textual features added to numerical models to forecast prices or volatility—while market series as input to text classifiers remain little explored. The LabelFusion architecture of Schlee et al. (2026) fuses a prompted LLM with a fine-tuned transformer encoder for financial news classification and was designed to accommodate additional input sources, proposing price time series as the next modality. This work carries out that proposal, evaluating it on monetary-policy stance classification, where sentences from Federal Reserve communication are labelled hawkish, dovish, or neutral. Central-bank language is deliberately hedged, so a sentence like “the Committee will act as appropriate” is read differently in a tightening than in an easing cycle, and every sentence originates from a dated public document, allowing market data from preceding months to be attached without further assumptions. The fused system exceeds a zero-shot LLM when trained on past communication and tested on future statements. Contributions are: (1) LabelFusion-TS, a recipe for extending intermediate fusion in financial text classification with further input modalities, instantiated for financial time series; (2) initial empirical evidence that financial time series can provide context for financial text classification across varying data availability regimes.
Related Work Summary:
LabelFusion obtains label scores from a prompted LLM and a contextual embedding from a fine-tuned RoBERTa encoder, fusing them with a small MLP. This is hybrid fusion in the taxonomy of Baltrušaitis et al. (2017). On Reuters-21578, the fusion beats standalone LLM and RoBERTa classifiers in the full-data regime, while the prompted LLM dominates in the low-data regime. The design is kept unchanged with one additional expert. For monetary-policy stance classification, the benchmark of Shah et al. (2023) is used: sentences from FOMC meeting minutes, press-conference transcripts, and speeches (1996–2022), labelled hawkish, dovish, or neutral by trained annotators, with baselines up to fine-tuned RoBERTa-large and zero-shot ChatGPT. No published work uses market data as input to stance classification; the connection between central-bank communication and markets is otherwise well documented. For time series, text, and distillation, purpose-built series encoders like PatchTST treat a time series as a sequence of patches, suiting short fixed windows; the multimodal literature predominantly uses text to help forecast series, whereas this work uses series to help classify text. The silver-label stage follows knowledge distillation in its self-training form: the LLM acts as a free annotator of in-domain unlabelled text, producing silver pseudo-labels for the student; unlike Xie et al., where injected noise lets the student surpass the teacher, here a subsequent gold fine-tuning stage plays this role.
Data and Task Summary:
The benchmark covers three forms of Federal Reserve communication published between 1996 and 2022: FOMC meeting minutes, press-conference transcripts, and speeches by Board members. Trained annotators assigned every sentence the policy direction it signals—hawkish for tightening, dovish for easing, neutral otherwise. The combined release contains 2,379 annotated sentences, of which 2,312 remain after removing duplicates and conflicting annotations. Neutral is the largest class (48%), followed by dovish (27%) and hawkish (25%). The task is single-label three-class classification of one sentence at a time. The metric is weighted F1, where class-wise F1 scores are averaged with weights proportional to class frequency in the test set; the frequent neutral class dominates the average. The published sentences state only the year of the document, but the evaluation splits data chronologically and pairs sentences with market data from the period preceding publication. The document dates are recovered by locating the document from which a sentence originates: first by exact match against sentence lists of individual documents, then by searching for the normalised sentence in raw document text. Sentences found in documents of more than one date (mostly formulaic passages) and sentences found in no document are excluded. This dates 1,692 sentences (73% of the benchmark) with a class distribution close to the full release. The evaluation splits dated sentences by time: 1,274 sentences up to September 2015 are used for training (with the most recent 15% held out for model selection), and 418 sentences from 2015–2022 form the test set. September 2015 serves as the cut-off throughout. The chronological split mirrors deployment and prevents leakage. A further 96 benchmark sentences occur only in documents published in 2014 or earlier; their exact day is unknown but they precede the cut-off with certainty, and they are used as additional training sentences for the sentence encoder, which operates on text alone and needs no date. For market context, for each sentence a window of six daily series over the 126 business days ending the day before its document date is built: the federal funds rate, the 2-year Treasury yield, the yield-curve slope, the log equity index, CPI year-over-year inflation, and the unemployment rate, each expressed as its change since the window start. All series are public (FRED). The window ends before the document date, and macro releases enter with a 45-day publication lag, so it contains only what markets actually knew at the time. Expert annotation covers only a fraction of the available text; the training-period documents contain roughly ten times as many sentences as the benchmark annotates. 13,017 of them are labelled with a prompted LLM, using the prompt of Shah et al. (2023)—the same model that is also one of the three classifiers. These silver labels are used for a single purpose: pre-training the sentence encoder. The selection excludes benchmark sentences, duplicates, and any sentence with word overlap above 0.85 with a test sentence, so no test material enters pre-training.
LabelFusion-TS Summary:
LabelFusion-TS keeps the recipe of LabelFusion: three specialised models (experts) are trained on the task individually, then frozen, and a small voting network is trained separately on their output. The text expert is RoBERTa-large trained in two stages: a silver stage (two epochs on the 13k LLM-annotated training-era sentences) and a gold stage (standard fine-tuning on human-labelled training sentences plus 96 recovered undated ones). Its CLS embedding represents the sentence. The LLM expert is an instruction-tuned open LLM (gemma4:31b at temperature 0) that classifies each sentence zero-shot with the classification prompt used by Shah et al. (2023); its vote is encoded as a one-hot vector. The same model with the same prompt produces the silver labels. The fusion is a voting MLP with one hidden layer combining the three views: ŷ = fMLP([h; z; m]) ∈ [0,1]3, which is LabelFusion’s fusion with one added input; removing m recovers LabelFusion exactly. Every component is trained with three different random seeds, and all 3×3×3=27 combinations are evaluated. The mean is reported, plus a probability ensemble that averages, per sentence, the class distributions predicted by the 27 systems and can be deployed as a single classifier.
Results Summary:
The time-split evaluation shows the market expert alone reaches 37.3% weighted F1. The text expert reaches 66.1%, already exceeding the zero-shot LLM (64.1%) whose silver labels it was pre-trained on. The full fusion reaches 67.4% on average, with all 27 seed combinations above the LLM, and the probability ensemble of the 27 systems reaches 70.2%, ahead of every individual component. The label-budget sweep shows LabelFusion-TS overtakes the zero-shot LLM at roughly 20% of the training pool (≈240 sentences) and remains above it for 100%, while the LLM row remains constant by construction. Schlee et al. (2026) report a crossover only after the majority of a substantially larger training set.
Discussion Summary:
The market–label connection is visible before any model is trained: grouping training-year sentences by the federal-funds-rate move over the preceding 126 business days, 66% are hawkish when the rate rose by at least 10 basis points, versus only 38% when it fell as much. Hedging often leaves the words alone underdetermined; “act as appropriate” appears in a dovish 2019 sentence and a neutral 2000 one. A text-only model under a chronological split cannot observe the test period’s environment, while the market window supplies it at test time by construction. This also explains why the market expert is weak alone (37.3%): all sentences of a document share one window, so it contributes to the environment rather than an independent judgement, useful only in combination with the text views. The same logic should extend to other tasks whose label semantics shift with the economic environment, such as sentiment or risk classification across volatility regimes.
Conclusion & Limitations Summary:
Financial text classifiers routinely ignore an input every human analyst consults: the market context. Adding it as an additional expert to an existing fusion architecture, trained with silver labels from the fusion’s own LLM and evaluated on a train-on-the-past, test-on-the-future protocol, beats a modern zero-shot LLM with as few as ∼240 human labels, and a simple seed-grid ensemble adds another three points. Components weak alone can be strong together, and the market time series is a cheap, public, time-aligned signal that there is little reason to keep leaving out. The study is limited to one task and benchmark, a seven-year test span dominated by two unusual regimes, a recovered date subset (73% of sentences), and silver labels drawn from the same LLM used as the zero-shot baseline—a different teacher could shift the picture.
Improvements for AI systems
Improvements to AI Systems:
-
Multimodal Context Integration for Text Classification: Extend any text classifier (e.g., sentiment, stance, or risk assessment) to accept time-series data as an additional input modality. The system can automatically fetch and align relevant market or environmental time series (e.g., interest rates, volatility indices, commodity prices) with each text instance, then fuse the text embedding with a time-series embedding via a small MLP. This enables the classifier to dynamically adjust its predictions based on the prevailing quantitative context, which is critical for tasks where label semantics shift over time (e.g.,
act as appropriate
means different things in easing vs. tightening cycles). -
Low-Resource Fine-Tuning via Silver-Label Pre-Training: For domains with scarce human annotations, the system can use a prompted LLM to generate pseudo-labels on a large corpus of unlabeled in-domain text, then pre-train a smaller encoder (e.g., RoBERTa) on these silver labels before fine-tuning on the limited gold labels. This two-stage process allows the model to learn domain-specific language patterns with as few as 240 human-labeled examples, outperforming a zero-shot LLM. The system can be deployed in any low-resource text classification setting (e.g., legal, medical, or financial documents) where expert annotation is expensive.
-
Seed-Grid Probability Ensembling for Robustness: Instead of relying on a single model, the system can train multiple instances of each expert (e.g., 3 seeds per modality) and combine all permutations (e.g., 27 systems) by averaging their predicted class probabilities per instance. This ensemble approach improves weighted F1 by 3 points over the best single model and reduces variance, making the system more reliable in production. The ensemble can be applied to any fusion architecture, not just this one, to boost accuracy without additional data or architecture changes.
-
Chronological Train-Test Splitting with Time-Aligned Features: The system can be designed to respect temporal causality by training only on data before a cut-off date and testing on future data, while ensuring that any time-series features used at test time contain only information available up to the day before the document's publication (e.g., with publication lags for macro data). This prevents data leakage and makes the system suitable for real-world deployment where future context is unknown. It also enables the model to generalize across different economic regimes, as demonstrated by the 2015–2022 test span.
-
Weak Expert Synergy in Hybrid Fusion: The system can incorporate a weak standalone expert (e.g., a time-series transformer that alone achieves only 37% F1) alongside stronger text-based experts, and still improve overall performance when fused. This suggests that even low-performing modalities can contribute valuable complementary information (e.g., market environment) that is not captured by text alone. The architecture can be generalized to add other weak signals (e.g., audio tone, image context, or metadata) to boost classification accuracy in multimodal tasks.
-
Automatic Date Recovery and Data Filtering for Temporal Alignment: The system can include a preprocessing pipeline that recovers exact publication dates for text instances by matching against known document lists and raw text, then filters out instances with ambiguous dates or high overlap with test data. This ensures clean temporal alignment with time-series features and prevents test-set contamination, improving the reliability of evaluation and deployment. This can be reused in any task where documents have known publication dates but are not explicitly labeled.
What the Improved AI System Can Do:
-
Classify financial or policy texts (e.g., central bank statements, earnings calls, news articles) with higher accuracy than text-only models, especially in low-data regimes, by incorporating market context (e.g., rates, yields, inflation) that human analysts would use.
-
Adapt to changing economic environments automatically—e.g., correctly interpret a sentence as hawkish in a rising-rate period and neutral in a stable period, without retraining.
-
Operate effectively with very few human labels (e.g., 240 sentences), making it feasible for niche domains where annotation is costly or slow.
-
Provide robust predictions via ensembling, reducing the risk of a single bad seed or model variant, and improving performance stability across different random initializations.
-
Deploy in real-time settings where only past data is available, since the system is trained and evaluated on a chronological split and uses only lagged, publicly available time series.
-
Extend to other tasks such as sentiment analysis across market volatility regimes, credit-risk classification with macroeconomic indicators, or political stance detection with polling time series, by swapping the time-series inputs and text corpus.
Sources
- Multimodal Machine Learning: A Survey and Taxonomy
- Distilling the Knowledge in a Neural Network
- RoBERTa: A Robustly Optimized BERT Pretraining Approach
Related papers
- Exploring Solution Divergence and Its Effect on Large Language Model Problem Solving
- Ishigaki-IDS-Bench: A Benchmark for Generating Information Delivery Specification from BIM Information Requirements
- Subliminal Steering: Stronger Encoding of Hidden Signals
- MedStruct-S: A Benchmark for Key Discovery, Key-Conditioned QA and Semi-Structured Extraction from OCR Clinical Reports
- The End of Transformers? On Challenging Attention and the Rise of Sub-Quadratic Architectures
- Untangling the Mechanisms of Misleading Context in Medical Question Answering