Are Whitepaper Claims Reflected in Market Structure? A Contamination-Aware Pipeline and a Power-Limited Null
q-fin.CP, cs.LG
Submitted: 2026-01-28
Updated: 2026-09-13
Comments: 15 pages, 3 figures, 5 tables. Corrects volume units, document/composition attribution, corpus-vintage statements and inference. Adds USD turnover reruns and full-space noise sensitivity
Code: https://github.com/studiofarzulla/whitepaper-claims
License: http://creativecommons.org/licenses/by/4.0/
The gist: Do the functional narratives in cryptocurrency whitepapers correspond to their tokens' market behaviour? We compare ten-category topical-emphasis profiles for 43 screened documents with seven market
Terminology
Abstract
Do the functional narratives in cryptocurrency whitepapers correspond to their tokens' market behaviour? We compare ten-category topical-emphasis profiles for 43 screened documents with seven market statistics calculated from 2023--2024 exchange data. The primary specification reconstructs USD notional turnover from hourly bars. Dimension-matched Procrustes congruence is ϕ=0.280 (permutation p=0.507), slightly below its permutation-null mean of 0.284; the zero-padded statistic gives the same non-detection. A four-leg comparison separates document replacement from changes in the assets included. Replacing documents on the 34 common assets changes padded congruence by-0.014 under USD turnover and-0.009 under base-token volume. Entity rankings and threshold crossings depend on both composition and specification, so an earlier contamination-only attribution is withdrawn. The documents are not a verified historical corpus: at least two postdate the market window. Excluding these documents, or excluding all seven assets with shorter histories, does not produce a significant alignment. Fresh numerical simulations distinguish injected signal from fitted congruence and compare noise restricted to the market subspace with noise throughout the text space. At the lowest classifier-agreement scenario, detection remains below 43% even at the largest injected signal. These are conditional checks of the alignment stage, not validation of the text instrument or exclusion bounds on economic effects. The contribution is an auditable non-detection and a specification-sensitive corpus diagnosis, with the inferential limits made explicit.
Sources
- FinBERT: Financial Sentiment Analysis with Pre-trained Language Models
- Information Filtering Networks: Theoretical Foundations, Generative Methodologies, and Real-World Applications
- Graph Regularized PCA
- Do Cryptocurrency Markets Differentiate Infrastructure from Regulatory Shocks? A Multi-Moment Event Study with Dependence-Robust Inference
- Large language models in finance : what is financial sentiment?
Related papers
- AI-Driven Multiscenario Interest Rate Forecasting: A Proof of Concept for Banking Asset Management
- LOB-ID: Evaluating Synthetic Market Data by Inception Distances
- Shapley-based Structural Analysis of Neural Calibration for Stochastic Volatility Models
- Optimal automation under overdispersed discrete risk: thresholds and hysteresis in a Negative Binomial model
- GARCH-Informed Neural Networks for Volatility Prediction in Financial Markets
- Loss Choice or Model Choice? The Role of Forecast Level in Cryptocurrency Volatility Forecasting