{"id":"b5ca0d95-b0c5-4d5c-9e33-ff0c13af8eef","arxiv_id":"2601.20336","paper_version":6,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A contamination-aware comparison of 43 content-verified cryptocurrency whitepapers against market statistics finds no significant alignment between whitepaper claims and market structure.","lead":"This paper builds a pipeline that compares what cryptocurrency whitepapers claim to do with how the tokens actually behave in markets, and finds no detectable link in a sample of 37 assets. The result is presented as an honest null result with careful checks on measurement noise and statistical power, rather than as evidence that no link exists.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Claims-matrix reliability is the load-bearing assumption: with inter-model agreement near chance (kappa=0.14, mean r=0.30), the non-significant phi could be attenuation; a human-gold-standard rerun would settle it.","rationale":"Good-faith reading: the paper is an unusually transparent null-result preprint. It flags the ATOM contamination, reports low inter-model agreement, acknowledges that the positive control is mechanically coupled, and explicitly states that it cannot distinguish weak alignment from none. Those admissions are real evidence, not decoration. The central claim is therefore modest: no detectable alignment, with rejection only of alignments above phi = 0.70. That claim is conditionally supported. The condition is the validity of the claims matrix. If BART-MNLI's category weights are noisy or systematically biased, every downstream phi, entity impact, and feature importance is measured with error. The paper's disattenuation (Section 5.10.7) tries to bound this risk but assumes the mean inter-method correlation is a valid reliability estimate. That assumption is not secure when all three methods are pretrained language models with overlapping training data. I therefore do not regard the null as fully established, but I also do not see grounds to reject the paper: its conclusions are already framed as absence of evidence, and the human-gold-standard rerun is a feasible condition rather than a fatal flaw. The reader's conditional verdict is appropriate, and my concern does not move it. The abstract/full-text version mismatch (43 vs 38 assets, phi 0.303 vs 0.246) should be resolved editorially, but it is secondary to the instrument-validity question.","tokens_in":22067,"tokens_out":5920,"duration_ms":58109,"concrete_test":"Obtain two independent expert annotations of the 38 whitepapers (or a stratified 500-chunk subsample) into the ten semantic categories, with Cohen's kappa reported. Average the expert category weights into a gold-standard claims matrix, re-run the Section 4.5-4.7 Procrustes and permutation pipeline, and compare the resulting phi and its permutation or percentile upper bound against the BART-MNLI values. If the expert-based claims-versus-statistics phi is non-significant and its upper bound stays below 0.70, the null survives; if it reaches phi >= 0.65 or p < 0.05, the reported null is an artifact of classifier attenuation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is a null: whitepaper claims do not align with market structure, and the study can reject only strong alignment (phi >= 0.70). The argument for this null runs through the BART-MNLI claims matrix C (Section 4.3). That matrix is the sole text-side input to every alignment statistic, entity-impact analysis, and feature-importance result. Its reliability is not established: exact top-1 agreement with DeBERTa-v3 is 32% (Cohen's kappa = 0.14), and the three-method pairwise correlation on the final claims matrix is only r = 0.30 (Table 5). The paper's own response is a Spearman disattenuation (Section 5.10.7), which uses the mean inter-method correlation 0.30 as reliability (with market-data reliability assumed 0.95) and concludes that the true claims-factors phi is about 0.11. That correction is valid only if the three classifiers are parallel measures with independent errors. They are not: all are language models trained on overlapping public text, so shared inductive bias can make the inter-method correlation either overstate or understate the reliability. Moreover, the positive control (statistics vs factors, p < 0.001) validates the market-side pipeline but not the claims-side instrument, because both sides derive from the same OHLCV data. Thus the 'absence of evidence' conclusion is not yet distinguishable from 'instrument too noisy to see alignment': the observed phi = 0.246 (or 0.303 dimension-matched) is exactly what a true moderate alignment attenuated by an inter-method reliability near 0.15 could look like. The load-bearing assumption is therefore not merely statistical power; it is the content validity of the claims matrix.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript asks whether the functional narratives in cryptocurrency whitepapers correspond to how the tokens actually behave in markets. It proposes a pipeline that combines zero-shot NLP classification of whitepaper text into ten semantic categories, seven cross-sectional market statistics computed from hourly OHLCV data, CP and Tucker tensor decompositions of the market tensor, and Procrustes rotation with Tucker's congruence coefficient to measure alignment. The submission contains two versions of the paper. The arXiv v4 abstract reports a contamination-aware analysis of 43 content-verified whitepapers with dimension-matched phi = 0.303 and zero-padded phi = 0.223, both non-significant, and states that an earlier entity-level signal was an artifact of corpus contamination. The embedded working-paper full text reports an n = 37 analysis with phi = 0.246 (claims-statistics) and phi = 0.058 (claims-factors), together with entity-level findings that specialized tokens such as XMR, CRV, and YFI help alignment. The paper also reports a positive control (statistics versus factors, p < 0.001), extensive robustness checks, a power analysis showing the text instrument is the binding constraint, and a Spearman disattenuation with a corrected claims-factors phi of about 0.11.","tokens_in":22426,"tokens_out":6685,"duration_ms":62996,"significance":"If the v4 null result holds, this is a useful contribution to empirical narrative economics and cryptocurrency research: it provides a transparent, falsifiable negative result with a clear statement of the power limits. The paper is unusually careful in several respects that deserve credit: it reports the full inter-model reliability numbers rather than hiding them, it discloses and diagnoses corpus contamination, it reports power calculations and minimum detectable effects, it provides code and data for replication, and it includes a positive control that shows the pipeline can detect real structure when it exists. The contamination diagnosis is a valuable cautionary result in itself. However, the significance of the central claim depends entirely on the validity of the claims matrix produced by zero-shot classification. With exact top-1 agreement of only 32% between the primary and an alternative classifier, and mean pairwise correlation of 0.30 across three methods, the null result is not yet distinguishable from an instrument-too-noisy-to-see-alignment result.","major_comments":[{"comment":"The manuscript contains two mutually incompatible versions of the central result. The arXiv v4 abstract reports an n=43 content-verified sample with dimension-matched phi=0.303 and zero-padded phi=0.223, and states that no entity-level signal survives contamination removal. The main text reports n=37, phi=0.246 for claims-statistics, phi=0.058 for claims-factors, and entity-level findings that XMR, CRV, YFI, and SOL help alignment while RPL, HBAR, AAVE, and SUSHI hurt it. Appendix D admits only one contaminated document (ATOM containing Binance Smart Chain text), which contradicts the v4 abstract's claim that roughly a quarter of documents were failed-download stubs or wrong-document whitepapers. The body of the paper never explains the 43-asset clean corpus, which documents were removed, or why Tables 6 and 10 still report the older numbers. A reader cannot determine which analysis is the one being claimed, and this must be resolved before the paper can be evaluated.","section":"Title/abstract (arXiv v4) vs. FULL TEXT abstract, Tables 2, 3, 6, 10, and Appendix D"},{"comment":"The claims matrix is the sole text-side input to every alignment result, and its reliability is not established. Exact top-1 agreement between BART-MNLI and DeBERTa-v3 is only 32% (Cohen's kappa = 0.14), the three-method mean pairwise correlation is r = 0.30, and discretized Fleiss kappa is 0.045. The disattenuation in Section 5.10.7 uses the mean inter-method correlation as the reliability estimate rho_XX = 0.30 and assumes the three classifiers are parallel measures with independent errors. That assumption is not defensible: all three are pretrained transformer language models trained on overlapping public text, so shared inductive biases can make inter-method agreement either overstate or understate true reliability. Consequently, the observed phi = 0.303 (or 0.246) is exactly the range that would arise from a moderate true alignment attenuated by measurement error. The paper needs a human gold-standard annotation of a random sample of chunks or assets to estimate criterion validity, and the disattenuation and all conclusions drawn from it must be updated with that estimate.","section":"Section 4.3.3 and Table 5; Section 5.10.7"},{"comment":"The statement that the study 'can reject strong alignment (phi >= 0.70)' is not justified under measurement error. With rho_XX = 0.30 and rho_YY = 0.95, a true phi of 0.70 would be attenuated to approximately 0.37, which is not far above the observed 0.303. The power calculations in Section 4.7.1 treat the observed phi as if it were the true value, and the bootstrap confidence intervals are explicitly acknowledged as upward-biased. The paper can legitimately say it does not detect significant alignment in this sample and that power is limited, but the stronger rejection claim, and the old abstract's statement that 'whitepaper narratives do not meaningfully predict market factor structure,' are not supported by the reported analysis.","section":"Sections 4.7.1 and 5.10.7"},{"comment":"The v4 abstract's central contribution is the contamination-aware result, yet the methods section describes no content-verification protocol. There are no criteria for identifying failed-download stubs or wrong-document whitepapers, no list of excluded documents, no description of the 43-document verified corpus, and no pre/post contamination comparison. The only contamination note in the body is Appendix D, which covers a single ATOM document and concludes it is unlikely to alter conclusions. The discrepancy between 'roughly a quarter' and one document must be resolved with a reproducible verification protocol and the analysis rerun on the verified corpus; without this, the contamination diagnosis is not independently checkable.","section":"Sections 3.2 and 4.3; arXiv v4 abstract"}],"minor_comments":[{"comment":"The text says the corpus was expanded to 38 assets, while the v4 abstract says 43 content-verified whitepapers; the relationship between these numbers needs to be stated explicitly.","section":"Section 4.3.1"},{"comment":"The disattenuation is reported only for claims-factors (phi_disatt about 0.11), but the corresponding disattenuated value for claims-statistics should also be reported for completeness, since that is the comparison emphasized in the abstract.","section":"Section 5.10.7"},{"comment":"The matched-dimension claims-statistics phi of 0.304 and the zero-padded phi of 0.246 from Table 6 appear to correspond to the v4 abstract's 0.303 and 0.223, but the numbers are not identical and the mapping is not explained.","section":"Section 5.10.3 and Table 6"},{"comment":"The title on the first page of the full text, 'Do Whitepaper Claims Predict Market Behavior? Evidence from Cryptocurrency Factor Analysis,' differs from the arXiv title 'Are Whitepaper Claims Reflected in Market Structure? A Contamination-Aware Pipeline and a Power-Limited Null'; the paper should use one title consistently.","section":"Title and running header"}],"recommendation":"major_revision","confidential_remarks":"The submission appears to be a version-controlled artifact that still contains the pre-contamination working paper in the body, with the new contamination-aware abstract attached at the front. This is a serious editorial problem: the central numbers, sample size, and conclusions differ between the two layers of the manuscript. I would ask the authors to submit a single coherent version before the content is reviewed in detail. In addition, the reliability of the claims matrix is the load-bearing assumption for the null result; a human gold-standard study is needed, not just an inter-model correlation between transformer classifiers. The paper's transparency about its limitations is commendable, but the central claim is currently underdetermined by the evidence."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nFirst, the headline: this is a careful null-result paper with a genuinely useful contamination diagnosis, but the arXiv abstract and the full text describe different analyses, and the text-side instrument is too unreliable to support the stronger reading. Treat it as a working paper that needs a major revision, not a finished result.\n\nWhat is actually new: the paper shows that an apparent entity-level alignment signal (specialized tokens aligning better than infrastructure tokens) was entirely a corpus contamination artifact—failed-download stubs and wrong-document whitepapers. That is a useful methodological contribution, and the power-limited null framing is honest: they report the minimum detectable effect (phi ≈ 0.66) and explicitly say they can only reject strong alignment, not distinguish weak alignment from none. The transparency is a real strength: full instrument-reliability numbers, bootstrap caveats, the mechanical coupling of the positive control, and the ATOM contamination note are all in the text.\n\nNow the soft spots, in proportion. The load-bearing problem is the claims matrix. BART-MNLI versus DeBERTa exact agreement is 32% (kappa=0.14); mean pairwise correlation with two other methods is 0.30. The paper's own disattenuation uses that 0.30 as reliability, but these are not parallel measures—all three models are trained on overlapping public text and share inductive biases, so the correlation can overstate or understate reliability. The positive control (statistics vs factors, p<0.001) validates the market side of the pipeline, not the text side. So the observed phi around 0.25 is exactly what a true moderate alignment attenuated by a noisy claims matrix could look like. The split-sample validation (phi=0.449, p=0.051) is suggestive but marginal and still within market data only.\n\nThere is also a version control problem. The arXiv abstract claims 43 content-verified whitepapers, a quarter of documents were failed stubs, and 'no token registers as helping alignment.' The full text says 38 whitepapers, only flags one contaminated ATOM document, and still reports entity-level help from XMR, CRV, YFI, and SOL. Those cannot both be true. A referee would need the authors to reconcile the versions.\n\nWho this is for: people working on crypto factor models, narrative economics, or cross-modal alignment methods. The contamination diagnosis and the power-limited null framing deserve reading. It should go to peer review, but only after the authors align the abstract with the full text, add a human-gold-standard reliability check on a subset of the claims matrix, and soften the 'rigorous evidence of misalignment' language in Section 7.2.\n\nBest","headline":"A careful null-result pipeline with a genuinely useful contamination diagnosis, but the claims matrix is too unreliable to carry the conclusion, and the arXiv abstract and full text describe different studies.","tokens_in":22952,"tokens_out":4492,"would_cite":true,"duration_ms":37849,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"After removing contaminated documents, the paper finds that cryptocurrency whitepaper claims show no statistically significant alignment with cross-sectional market structure, and the study is too underpowered to distinguish weak…","keywords":["cryptocurrency whitepapers","narrative economics","Tucker's congruence coefficient","Procrustes rotation","zero-shot classification","market factor structure","contamination-aware pipeline","statistical power"],"falsifier":"Independently hand-label a random sample of the 43 whitepapers' text chunks into the ten semantic categories and recompute the Procrustes-Tucker alignment; if the human-labeled claims matrix aligns with market statistics at or above the minimum detectable effect of φ ≈ 0.66, or significantly above the permutation null, then the paper's null is an artifact of the NLP instrument rather than a genuine absence of narrative-market correspondence.","tokens_in":21733,"feed_emoji":"📉","tokens_out":8034,"duration_ms":66808,"temperature":0.7,"pith_summary":"The paper sets out to test whether the functional claims in cryptocurrency whitepapers correspond to how the tokens actually behave in markets. It builds a content-verified, contamination-aware pipeline: zero-shot NLP classifies 43 whitepapers into ten functional categories, seven market statistics are computed from two years of hourly prices, and Procrustes rotation plus Tucker's congruence coefficient measures alignment between the two spaces. The central result is a null: dimension-matched alignment is φ = 0.303 and zero-padded alignment is φ = 0.223, both statistically non-significant. The paper also reports that an apparent entity-level signal—specialized tokens appearing to align better—was an artifact of corpus contamination, since roughly a quarter of the earlier corpus was failed-download stubs or wrong documents. The study can reject strong alignment (φ ≥ 0.70) but cannot distinguish weak alignment (φ ≈ 0.3) from no alignment, so the contribution is a well-characterized absence of evidence, not evidence of absence.","feed_headline":"Whitepaper claims don't match crypto market structure","feed_subtitle":"Cleaned 43-coin study finds no significant alignment; strong effects are ruled out, weak ones remain possible.","key_machinery":"The load-bearing machinery is a three-part measurement pipeline: zero-shot NLI classification that converts each whitepaper into a probability-weighted claims matrix over ten functional categories; a cross-sectional market-statistics matrix built from seven metrics (mean return, volatility, Sharpe ratio, max drawdown, average volume, vol-of-vol, and trend); and Procrustes rotation with Tucker's congruence coefficient φ, which finds the best orthogonal alignment between the two spaces and measures per-dimension cosine similarity without mean-centering. A Monte Carlo permutation test and power simulation set the detection limit: at n = 43 the test can reject φ ≥ 0.70 but not φ ≈ 0.3. The contamination-aware content verification—removing failed-download stubs and wrong documents—is what turns the earlier apparent entity-level signal into a diagnosis rather than a finding.","core_discovery":"On a cleaned corpus of 43 content-verified whitepapers, the paper claims that whitepaper narratives carry no detectable structural correspondence with cross-sectional market behavior. Using a claims matrix from zero-shot classification across ten semantic categories and a market-statistics matrix from hourly OHLCV data over 2023–2024, it aligns the spaces by Procrustes rotation and measures similarity with Tucker's congruence coefficient. The observed coefficients—dimension-matched φ = 0.303 and zero-padded φ = 0.223—are both non-significant against a permutation null. A positive control comparing market statistics to latent factors is significant, showing the pipeline can detect real structure when it exists. The paper further reports that the earlier apparent finding that specialized tokens align more strongly than infrastructure tokens disappeared after content verification: on the clean corpus, no single token registers as helping alignment.","pith_inferences":["Inference: if the binding constraint is the low reliability of the text instrument (top-1 agreement of 32%, κ = 0.14, mean pairwise correlation r = 0.30), then improving claims extraction—domain-adapting the classifier or using human adjudication—would raise power more than simply adding assets.","Inference: the same contamination-aware design could be applied to other narrative assets, such as stock prospectuses or green bonds, where document provenance is similarly unreliable; the failed-download-stub failure mode is probably general.","Inference: the paper's static whitepaper corpus leaves open that contemporaneous narratives (social media, governance posts, developer communication) align better with market structure; dynamic narrative tracking is a natural test.","Inference: because φ is a contemporaneous structural measure, these results neither confirm nor refute predictive forecasting from narratives; absence of alignment is compatible with narratives that predict but are too weak to move cross-sectional structure."],"forward_implications":["If alignment is truly absent, portfolios built on whitepaper categories (such as DeFi baskets or Layer-1 sets) are not supported by any detected structural link to the market cross-section.","Regulators relying on whitepaper disclosures as informativeness signals would find no support in this study for those disclosures mapping to market behavior.","The contamination result implies that earlier and future claims of narrative-market alignment in crypto must be rechecked against document provenance before being trusted.","The power analysis sets a concrete bound: this design can rule out strong alignment but cannot adjudicate weak alignment, so investment and policy conclusions should only cite the strong-alignment rejection.","The significant statistics–factors control shows the alignment machinery is not itself broken, strengthening the interpretation that the null is specific to narrative content."],"supporting_citations":[{"why":"Supplies the SVD solution to the orthogonal Procrustes problem used to align claims and market spaces.","marker":"Schönemann, 1966"},{"why":"Introduced the congruence coefficient φ used to measure factor similarity.","marker":"Tucker, 1951"},{"why":"Provides the interpretation thresholds (0.65, 0.85, 0.95) for congruence coefficients used to characterize alignment strength.","marker":"Lorenzo-Seva and ten Berge, 2006"},{"why":"The BART model used for zero-shot classification of whitepaper text.","marker":"Lewis et al., 2020"},{"why":"Establishes the entailment-based zero-shot classification approach (hypotheses like 'This text is about [category]') underlying the claims matrix.","marker":"Yin et al., 2019"},{"why":"Establishes null distributions for congruence coefficients from simulated data, informing significance testing.","marker":"Korth and Tucker, 1975"},{"why":"Examines the distribution of factor congruence under chance conditions after Procrustes rotation, guiding the permutation test.","marker":"Paunonen, 1997"},{"why":"Provides the cryptocurrency factor-model baseline that motivates comparing whitepaper claims to market factor structure.","marker":"Liu and Tsyvinski, 2021"}],"fun_headline_variants":["Whitepaper claims don't survive clean data test","Clean corpus shows no whitepaper-market alignment","Contamination erased apparent whitepaper signal","43 coins: no whitepaper-market link after cleaning"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim depends on treating the content-verified claims matrix produced by the zero-shot classifier as an accurate measure of what each whitepaper actually says; if the classifier's low inter-model agreement means it does not, the null could reflect measurement failure rather than true dissociation.","fun_headline_variants_meta":{"raw":{"variants":["Whitepaper claims don't survive clean data test","Clean corpus shows no whitepaper-market alignment","Contamination erased apparent whitepaper signal","43 coins: no whitepaper-market link after cleaning"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000758,"raw_usage":{"total_tokens":3427,"prompt_tokens":1066,"completion_tokens":2361,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":682,"completion_tokens_details":{"reasoning_tokens":2300}},"tokens_in":682,"tokens_out":2361,"duration_ms":16778,"temperature":1.0,"reasoning_tokens":2300,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:38:25.946798+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Independently hand-label a random sample of the 43 whitepapers' text chunks into the ten semantic categories and recompute the Procrustes-Tucker alignment; if the human-labeled claims matrix aligns with market statistics at or above the minimum detectable effect of φ ≈ 0.66, or significantly above the permutation null, then the paper's null is an artifact of the NLP instrument rather than a genuine absence of narrative-market correspondence.","supporting_citations":[],"review_version":2}