{"id":"c756d398-acbb-4110-bb81-e1654d0fc05f","arxiv_id":"2602.07046","paper_version":4,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"In 15 negative crypto events (2019–2025), cumulative abnormal returns are statistically indistinguishable between infrastructure failures (−7.6%) and regulatory enforcement (−11.1%), difference +3.6 pp, p=0.81 under event-level block bootstrap.","lead":"The body of this preprint reports that after resampling whole events instead of individual asset-days, crypto markets show no statistically significant difference in price impact between infrastructure failures and regulatory enforcement (8 vs 7 events, p=0.81). It matters because the same events appear to trigger much larger volatility responses, suggesting markets process shock types through risk rather than expected returns.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Abstract and body report different primary analyses: copula-GARCH inference and p=0.283 in the abstract are absent from the body, whose block-bootstrap p=0.81 is the only presented test.","rationale":"The reader's CONDITIONAL verdict is sound, but I would weight the concerns differently. The most load-bearing issue for the paper's central claim is the internal inconsistency between the abstract and the body. The abstract describes an elaborate inference apparatus (GJR-GARCH-X, Student-t copula, design-effect calculations, inference ladder, Monte-Carlo size study) and reports a different point estimate and p-value (Δ=7.19 pp, p=0.283; design-effect p≈0.07–0.15) than the body (Δ=3.6 pp, p=0.81; Section 5.1, Table 2). The body's methodology section (4.3) presents only the event-level block bootstrap; the GARCH-X/copula methods are not defined anywhere in the text. This is not a cosmetic mismatch: if the abstract's analysis is the 'inference of record,' then the body's primary table is not the paper's actual test, and the reader cannot assess the claimed methods. If the body is the intended analysis, the abstract misrepresents the results. Either way, the central null claim is not cleanly supported by a single verifiable analysis. The event-exogeneity concern (Section 3.1, limitation 6.4.5) is real but secondary: the paper's own exogenous-only subsample (Table 11, Section 5.12) and the small number of return-threshold-selected events (2/17) limit its bite, and the null persists across all robustness checks (permutation, Ibragimov–Müller, winsorization, window variations). Thus I disagree with the reader's labeling of event exogeneity as the weakest assumption, though I agree it is a limitation. The abstract/body mismatch is more load-bearing because it undermines verifiability of the headline result and the claimed contribution. A check of the GitHub repository for the copula-GARCH scripts would settle which analysis is real. The verdict remains CONDITIONAL: the paper needs revision to harmonize the abstract with the body or to provide the missing methods.","tokens_in":15799,"tokens_out":14268,"duration_ms":150151,"concrete_test":"Inspect the GitHub repository (github.com/studiofarzulla/sentiment-without-structure) and run the analysis scripts that correspond to the abstract's GJR-GARCH-X/Student-t-copula CCC-GARCH-X bootstrap, the inference ladder, and the Monte-Carlo size study. If those scripts do not exist or do not reproduce Δ=+7.19 pp, p≈0.283, design-effect p≈0.07–0.15, the abstract's methodological claims and numbers are unsupported; the body's block-bootstrap result (p=0.81) would then be the only verifiable primary analysis. If the scripts do reproduce the abstract, the body's Section 5.1 must be reconciled with the abstract (sample size, event count, point estimate) or the paper contains two incompatible primary analyses.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim ('no statistically significant difference') is operationalized in the body (Section 5.1, Table 2) as Δ=+3.6 pp, p=0.81, CI [−25.3%, +30.9%], on 8 vs 7 events. The abstract reports a different analysis of record: Δ=+7.19 pp, p=0.283, with a Student-t-copula CCC-GARCH-X bootstrap and design-effect p≈0.07–0.15, referencing 50 events and six assets. The body never describes this GARCH-X/copula model, the inference ladder, or the Monte-Carlo size study; Section 4.3's only method is the event-level block bootstrap, and Section 3.1 says the sample is 31 events/4 assets. The point estimate doubles and the p-value roughly triples between the two. If the abstract's pipeline is the intended record, the body's Table 2 is an unreconciled alternative and the claimed methods are unverifiable; if the body is intended, the abstract misrepresents the paper. Either way, the null claim's evidential basis is ambiguous, and the paper's main methodological contribution is absent from the body. The reviewing rule instructs flagging missing support; this is exactly such a case.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper tests whether negative infrastructure and regulatory shocks produce different cumulative abnormal returns (CARs) in cryptocurrency markets. Using a four-category event classification and an event-level block bootstrap on 31 events (8 versus 7 analyzed negative events) across BTC, ETH, SOL, and ADA, the body reports mean CARs of -7.6% versus -11.1%, a difference of +3.6 pp with 95% CI [-25.3%, +30.9%] and p=0.81. The paper interprets this as an exploratory null: returns do not appear to distinguish the two shock classes. Robustness checks include permutation tests, market model adjustments, winsorization, leave-one-out, non-overlapping events, and Ibragimov-Müller few-cluster inference. The arXiv abstract, however, describes a different analysis: 50 events, six assets, GJR-GARCH-X/copula bootstrap inference, a +7.19 pp difference with p=0.283, and design-effect p-values of 0.07-0.15. That analysis is absent from the body. The body also relies on a companion volatility study for the claim that markets differentiate shock types through the risk channel.","tokens_in":16099,"tokens_out":7895,"duration_ms":81654,"significance":"Both versions of the paper make a useful methodological point: event-level/block-bootstrap inference can overturn naive i.i.d. significance in few-event, cross-sectionally correlated crypto event studies. The body is unusually honest about low power (MDE approximately 40 pp) and selection/dating concerns, and the reproducible code and data statement is a strength. If the block-bootstrap null is the record, the paper is a cautionary empirical contribution. However, the abstract and body report different primary analyses, so the central null's evidential basis is ambiguous. The interpretive claim that markets differentiate through the risk channel depends on an external paper not included here. The event-selection and event-dating procedures condition partly on the outcome being measured, which undermines causal readings unless addressed. These issues are fixable within the manuscript's scope, but they are load-bearing for the paper's central claims.","major_comments":[{"comment":"The submitted abstract and the body report different primary analyses. The abstract claims a GJR-GARCH-X/copula bootstrap with 50 events, six assets, a +7.19 pp CAR difference (p=0.283), and design-effect p-values of 0.07-0.15. The body's Table 2 (Section 5.1) reports Δ=+3.6 pp, p=0.81, CI [-25.3%, +30.9%], from an event-level block bootstrap on 31 events/four assets; Section 4.3 describes only that bootstrap. The GARCH-X/copula model, the inference ladder, and the Monte-Carlo size study are never defined in the body. Because the central null's evidential basis changes between the two versions, the authors must reconcile them: either add the missing analysis and explain the discrepancy, or make the abstract match the body.","section":"Abstract vs. §5.1/Table 2; §4.3"},{"comment":"Event inclusion and dating condition on the outcome that defines the CAR. Criterion 1 in §3.1 admits events with same-day or three-day |BTC| return >5%; §6.4.5 admits that gradual events are dated by 'first major price discontinuity.' The former means the sample is selected partly on the abnormal returns being measured; the latter means the event day t=0 is chosen from the return path. The 'exogenous-only' analysis in §5.12 addresses inclusion for events labeled 'Exogenous' or 'Both,' but Table 13 shows most negative events are 'Both,' and no alternative dating is implemented. Because the point estimate and the bootstrap distribution both depend on t=0, this is a load-bearing identification concern. Please re-date using news timestamps and report sensitivity, or explicitly bound the resulting bias.","section":"§3.1, §5.12, §6.4.5"},{"comment":"Section 1 states that the enforcement-capacity hypothesis was 'specified ex ante,' but Section 4.4 discloses that event selection criteria 'evolved iteratively' and that the four-category classification 'was developed after initial data exploration.' These statements are in direct tension. If the classification and selection rules were not fixed before examining the CARs, the exploratory nature of the null should be stated in the abstract and conclusion, not only in Section 6.4. Please either provide a dated pre-analysis plan or transparently describe which rules were fixed before the data were examined.","section":"§1 vs. §4.4"},{"comment":"The interpretive conclusion that 'markets differentiate shock types through the risk channel' rests entirely on the companion paper (Farzulla, 2025a), which is not included in this manuscript. The body reports no second-moment estimation. The 5.7x variance ratio, p=0.0008, the regime F=45.23, and the flat regulatory coefficient are external claims. Moreover, the methodological note in §6.3 says the companion uses a different data source (CoinGecko, six assets) and return definition (log returns), so the direct comparability is not established. Either include a self-contained version of that analysis or explicitly demote this conclusion to an external implication. Without this, the title/abstract promise of a 'multi-moment' study is unsupported.","section":"§6.1, §6.3"}],"minor_comments":[{"comment":"The arXiv title, 'Do Cryptocurrency Markets Differentiate Infrastructure from Regulatory Shocks? A Multi-Moment Event Study with Dependence-Robust Inference,' does not match the body's title, 'Same Returns, Different Risks.' The body abstract and the arXiv-side abstract also differ in sample size and method. Please harmonize titles and abstracts.","section":"Title and Abstract"},{"comment":"Sample observation counts are inconsistent. Table 14 lists Infra_Pos N assets as 23 and Reg_Pos as 32, whereas Table 3 reports 21 and 31; the Table 14 summary says total event-asset observations = 118, but summing the listed assets gives 113. Please reconcile these counts.","section":"Tables 3, 14, and summary counts"},{"comment":"Section 4.3 step 2 instructs within-event averaging so each event receives equal weight, but Table 2's note says the primary results use the observation-weighted bootstrap. Since the two schemes give similar but not identical p-values (0.81 vs. 0.93), the paper should state clearly which is the primary estimator and why.","section":"§4.3 vs. Table 2 note"},{"comment":"The placebo test reports p=0.08. Describing the constant mean model as 'adequately controls' is too strong for a borderline result; the text already says 'borderline,' but the conclusion could be phrased as 'we cannot reject the null of no drift at the 5% level.'","section":"§5.9"}],"recommendation":"major_revision","confidential_remarks":"The abstract/body mismatch is severe enough that the manuscript should not proceed until it is resolved. The body's block-bootstrap analysis is honest and carefully caveated, but the abstract appears to describe a different, more elaborate study that is absent from the body. If the authors do not add the GARCH-X/copula analysis, the abstract must be rewritten to match the body. The selection/dating issue is also central and should be addressed with alternative dating or explicit bias bounds."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know: the body is a candid, low-powered event study with a robust null result; the arXiv abstract reads like a different paper. The body's central finding—no significant CAR difference between 8 negative-infrastructure and 7 negative-regulatory events under event-level block bootstrap (Δ=+3.6 pp, p=0.81, CI [−25.3%, +30.9%])—is honestly labeled exploratory and survives a broad set of robustness checks: permutation p=0.93, Ibragimov–Müller p=0.93, winsorization, market-model, leave-one-out, exogenous-only subsample p=0.61. The 4-category classification is a modest but real improvement over previous pooled designs, the code and data are public, and the authors are unusually transparent about selection, dating, and their own HARKing.\n\nThe soft spot that matters: the arXiv abstract describes a different analysis—50 events, six assets, Student-t-copula CCC-GARCH-X, p=0.283, design-effect adjustments. None of that appears in the body, which uses 31 events, four assets, and block bootstrap. The full-text abstract matches the body, so the mismatch may be stale metadata or an unreconciled companion, but as submitted the record is ambiguous. If the GARCH-copula toolkit is the contribution, it isn't in this text; if the block bootstrap is the contribution, the abstract overstates what's delivered.\n\nOther soft spots are proportionate. Event selection uses a |BTC|>5% screen and gradual events are dated by 'first major price discontinuity'—conditioning on the returns being measured. The authors acknowledge this and run an exogenous-only subsample, which mitigates but doesn't eliminate the dating issue. The bigger problem is power: with 8 vs 7 events the MDE is roughly 40%, so the null is weak evidence of no difference. They say so themselves. And the substantive interpretation leans on an unreviewed companion variance study—fine as context, not as evidence.\n\nWho should read it: researchers doing crypto event studies, especially anyone tempted to report i.i.d. standard errors on event-asset panels. It's a useful cautionary tale about few-cluster inference and a good example of honest reporting of an exploratory null. But I'd want the abstract and body reconciled before citing it.\n\nMy recommendation: send it to peer review anyway. The empirical question is relevant, the robustness work is real, and the mismatch can be fixed in revision. The first referee report should ask for a single analysis of record and a clear statement of which data and method answer the title question.","headline":"The body is an honest, low-powered null result with solid robustness; the arXiv abstract describes an entirely different GARCH-copula analysis, so the paper needs reconciliation before it can be trusted.","tokens_in":16591,"tokens_out":5306,"would_cite":false,"duration_ms":51570,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Infrastructure failures and regulatory enforcement produce statistically indistinguishable cryptocurrency market returns, according to an event study using block bootstrap inference.","keywords":["cryptocurrency","event study","block bootstrap","cumulative abnormal returns","infrastructure risk","regulatory enforcement","volatility asymmetry","cross-sectional dependence"],"falsifier":"A direct test would re-run the comparison on a pre-registered, news-timestamped event set of at least ~60 negative events per category (the paper's own power math says ~930 per group for 80% power at the observed effect), using intraday prices and named-token abnormal returns. If the CAR gap then cleared zero with a narrow confidence interval, the null would be overturned; if a companion intraday variance analysis found the 5.7× gap, the risk-channel interpretation would be further supported.","tokens_in":15653,"feed_emoji":"📉","tokens_out":5470,"duration_ms":48942,"temperature":0.7,"pith_summary":"The paper asks whether cryptocurrency markets price infrastructure failures (exchange collapses, hacks, protocol depegs) differently from regulatory enforcement (lawsuits, bans). Using 31 events across four large assets from 2019–2025, it finds mean cumulative abnormal returns of –7.6% for infrastructure failures and –11.1% for regulatory enforcement; the 3.6-point difference has a block-bootstrap confidence interval of [–25.3%, +30.9%] and p=0.81. The paper argues this null is methodologically meaningful: naive i.i.d. tests that pool asset-event observations would overstate precision. Read together with a companion conditional-variance analysis reporting 5.7 times larger volatility responses to infrastructure events, the result suggests markets differentiate shock types through risk, not expected returns.","feed_headline":"Block bootstrap finds no crypto return gap for hacks vs regulation","feed_subtitle":"Return impact is indistinguishable; the distinction shows up in volatility, not prices.","key_machinery":"The load-bearing tool is the event-level block bootstrap: resample whole events (keeping the four assets' CARs together) to build the null distribution, rather than pooling asset-event observations as independent. This corrects for the cross-sectional correlation that inflates degrees of freedom in standard event studies. The paper also introduces a 4-category classification (infrastructure/regulatory × positive/negative) so that like-valence shocks are compared, and it triangulates with the Ibragimov–Müller few-cluster t-test on event-level means.","core_discovery":"On the paper's own terms, the discovery is a null result made informative by dependence-robust inference. When the unit of resampling is the event rather than the asset-day, the first-moment response to negative infrastructure shocks (mean CAR –7.6%) is not statistically distinguishable from the response to negative regulatory shocks (–11.1%), with a difference of +3.6 percentage points, 95% CI [–25.3%, +30.9%], p=0.81. The paper shows this null survives market-model adjustment, winsorization, permutation tests, leave-one-out analysis, and few-cluster tests. Because a companion study finds a highly significant 5.7× difference in conditional variance, the paper's conclusion is that the market","pith_inferences":["If the return-level null is genuine, then market efficiency with respect to shock type holds for prices but not for volatility; that asymmetry could imply that option markets, not spot returns, are the place to look for differential pricing of regulatory vs infrastructure risk.","A testable extension: infrastructure CARs should mean-revert faster than regulatory CARs over 60–90 days as bounded uncertainty resolves; the current 35-day window cannot distinguish this.","The null may partly be a design artifact of spillover measurement: using only BTC/ETH/SOL/ADA captures market-wide response but omits directly named tokens (e.g., XRP in SEC v. Ripple), which could show stronger regulatory effects.","The paper's block-bootstrap recipe transfers to any cross-asset event study in heavy-tailed markets; applying it to other shock taxonomies (e.g., monetary vs fiscal news) would test whether the null generalizes."],"forward_implications":["Prior parametric significance claims in crypto event studies may be artifacts of pseudoreplication across correlated assets.","The enforcement-capacity hypothesis (infrastructure hits returns harder) is neither confirmed nor rejected; the observed effect size is far below the minimum detectable effect of ~40 pp with 8 vs 7 events.","Portfolios holding crypto assets need roughly 4–5 times larger capital buffers for infrastructure-event risk than return-based models imply.","Event-type taxonomy should separate positive from negative valence within infrastructure and regulatory categories.","Future confirmatory work requires ~930 events per category to detect the observed effect at 80% power."],"fun_headline_variants":["Crypto hacks vs regulation: no return gap once dependence is modeled","Volatility gap real, return gap not: crypto event study reanalysis","Dependence-robust test finds no crypto shock return gap","Crypto shock returns: no gap between hacks and regulation","Crypto market shocks: return effect vanishes under bootstrap"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The events are treated as dated and selected without conditioning on the return path being measured; in particular, gradual events are dated by 'first major price discontinuity' and some events enter via a 5% return screen, so the abnormal returns that feed the CARs may have influenced which events are in the sample and when.","fun_headline_variants_meta":{"raw":{"variants":["Crypto hacks vs regulation: no return gap once dependence is modeled","Volatility gap real, return gap not: crypto event study reanalysis","Dependence-robust test finds no crypto shock return gap","Crypto shock returns: no gap between hacks and regulation","Crypto market shocks: return effect vanishes under bootstrap"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000547,"raw_usage":{"total_tokens":2529,"prompt_tokens":903,"completion_tokens":1626,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":647,"completion_tokens_details":{"reasoning_tokens":1552}},"tokens_in":647,"tokens_out":1626,"duration_ms":11785,"temperature":1.0,"reasoning_tokens":1552,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-03T04:36:54.089691+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A direct test would re-run the comparison on a pre-registered, news-timestamped event set of at least ~60 negative events per category (the paper's own power math says ~930 per group for 80% power at the observed effect), using intraday prices and named-token abnormal returns. If the CAR gap then cleared zero with a narrow confidence interval, the null would be overturned; if a companion intraday variance analysis found the 5.7× gap, the risk-channel interpretation would be further supported.","supporting_citations":[],"review_version":1}