Human Agreement and Return Association Are Not Interchangeable Criteria
cs.AI, cs.CL, cs.SI
Submitted: 2026-09-10
Updated: 2026-09-24
License: http://creativecommons.org/licenses/by/4.0/
The gist: Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal.
Terminology
Abstract
Financial NLP has a standard workflow: validate a sentiment tool against human labels, then trust it to extract market signal. This assumes the two evaluations measure the same thing. We test that assumption in a setting where both can be measured at once: a corpus of securities class actions (2002-2025) linking 70,500 X messages to abnormal stock returns, with a single-annotator human labelled gold sample. Running five instruments (VADER, Loughran-McDonald, FinBERT, Twitter-RoBERTa, and an LLM annotator) through one identical pipeline, we find that the relationship between construct and predictive validity depends on the sampling convention and score representation. Under conventional method-specific sampling, human agreement aligns more closely with graded same-day associations than with one-day leads. On a fixed-n panel, however, agreement has similar graded rank correlations at both horizons, while the coarse ordering remains weak. Benchmark agreement therefore establishes semantic validity but does not by itself determine predictive rankings. In a conversation that is 17.6% spam, message volume predicts neither market damage nor settlement size.
Sources
- FinBERT: Financial Sentiment Analysis with Pre-trained Language Models
- Assessing Look-Ahead Bias in Stock Return Predictions Generated By GPT Sentiment Analysis
- Chronologically Consistent Large Language Models
- Sentiment trading with large language models
- Can ChatGPT Forecast Stock Price Movements? Return Predictability and Large Language Models
Related papers
- MAVEN-T: Reinforced Heterogeneous Distillation for Real-Time Multi-Agent Trajectory Prediction
- Model Discovery Agent: LLM-assisted Bayesian experiment design for data-efficient discovery of mechanistic world models
- The Clinician's Veto: Navigating Trust, Liability, and Uncertainty in Autonomous AI Prescribing
- MindHelper: Closed-Loop Embodied Mental-State Reasoning for Precision Intervention
- Incumbent Advantage: Brand Bias and Cognitive Manipulation Dynamics in LLM Recommendation Systems
- VSAL: A Vision Solver with Adaptive Layouts for Graph Property Detection