Are Whitepaper Claims Reflected in Market Structure? A Contamination-Aware Pipeline and a Power-Limited Null

arXiv:2601.20336 · q-fin.CP, cs.LG · Submitted 2026-01-28 · Read on arXiv

q-fin.CP, cs.LG

Submitted: 2026-01-28

Updated: 2026-09-13

Comments: 15 pages, 3 figures, 5 tables. Corrects volume units, document/composition attribution, corpus-vintage statements and inference. Adds USD turnover reruns and full-space noise sensitivity

Code: https://github.com/studiofarzulla/whitepaper-claims

License: http://creativecommons.org/licenses/by/4.0/

The gist: Do the functional narratives in cryptocurrency whitepapers correspond to their tokens' market behaviour? We compare ten-category topical-emphasis profiles for 43 screened documents with seven market

Terminology

Abstract

Do the functional narratives in cryptocurrency whitepapers correspond to their tokens' market behaviour? We compare ten-category topical-emphasis profiles for 43 screened documents with seven market statistics calculated from 2023--2024 exchange data. The primary specification reconstructs USD notional turnover from hourly bars. Dimension-matched Procrustes congruence is ϕ=0.280 (permutation p=0.507), slightly below its permutation-null mean of 0.284; the zero-padded statistic gives the same non-detection. A four-leg comparison separates document replacement from changes in the assets included. Replacing documents on the 34 common assets changes padded congruence by-0.014 under USD turnover and-0.009 under base-token volume. Entity rankings and threshold crossings depend on both composition and specification, so an earlier contamination-only attribution is withdrawn. The documents are not a verified historical corpus: at least two postdate the market window. Excluding these documents, or excluding all seven assets with shorter histories, does not produce a significant alignment. Fresh numerical simulations distinguish injected signal from fitted congruence and compare noise restricted to the market subspace with noise throughout the text space. At the lowest classifier-agreement scenario, detection remains below 43% even at the largest injected signal. These are conditional checks of the alignment stage, not validation of the text instrument or exclusion bounds on economic effects. The contribution is an auditable non-detection and a specification-sensitive corpus diagnosis, with the inferential limits made explicit.

Sources

Related papers