{"id":"7fb0062b-f06b-45c2-85c3-a019c6b2d9ad","arxiv_id":"2507.07884","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"high","formal_verification":"none","parameter_count":4,"one_line_summary":"Searches for a small subset of conspiracy theories, notably the Great Replacement and the Rothschilds, modestly improve the prediction of weekly hate crimes in Michigan, while most of the 36 theories tested add no predictive signal.","lead":"This paper tests whether Google searches for 36 conspiracy theories can forecast hate crimes, using weekly hate crime counts in Michigan from 2015 to 2019. A smart generalist might read it because it is a quantitative test of the worry that online misinformation about conspiracies is connected to offline violence.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The positive subset may be a selection artifact: 180 unadjusted comparisons on a single split make ~9 false positives expected, and the post-hoc permutation test only shows the model used the feature's temporal shape, not that the improvement is reliable.","rationale":"The paper is transparent, and much of what it reports is plausible and useful: 28 of 36 theories show no below-baseline improvement, the authors flag the keyword-validity problem (Sec. 5.1), hate-crime under-reporting, IP-geolocation error, and they mostly refrain from causal language. The null result for most theories is what a principled exploration should produce, and the feature-permutation idea is a reasonable sanity check on whether the model uses temporal shape. My concern is with the survivorship of the positive subset, and it has three nested parts. (1) Multiplicity: 36 theories × 5 lags = 180 comparisons; at a 5% false-positive rate, ~9 below-baseline cells are expected, and Table 2 shows 16 flagged cells across 8 theories, squarely in the range a null process would produce. (2) The permutation test does not address this. Its null is that the fitted model does not depend on the feature's temporal ordering; a chance alignment of a search series with test-period crime is itself a temporal ordering, so permuting it will tend to increase MAE and the test will 'pass.' The test is applied only to already-flagged cells, uses K = 3 permutations with the minimum as the bar, and has no null distribution, so it cannot deliver false-discovery control. (3) The entire evaluation is a single 80/20 split with a single seed; with T = 262 and a 4-week horizon, the effective number of independent test errors is small, and no confidence interval for the MAE differences is reported. Because early stopping (Sec. 3.2) uses the same final 20% on which the reported errors are computed, there is an optimistic bias that may favor the more flexible conspiracy-augmented models. Two observations strengthen the concern: the authors' own Table 3 eliminates the only flagged lag for Obama Kenya and the Great Reset, so the 'eight theories' of Sec. 4 reduce to six; and the abstract's 2–3 week delayed-effect claim rests mainly on two cells (Great Replacement, lags 2 and 3) with no replication evidence. The settled test is a surrogate null simulation: circularly shift each search series independently (which exactly preserves autocorrelation and marginal distribution while destroying alignment with crime) and run the complete pipeline. This directly calibrates how many 'significant' theories the discovery rule would produce under the null. If the null 95th percentile reaches the observed 6 theories / 12 cells, the central claim is a selection artifact and the positive conclusion should be rejected; if the null count is much lower, the conditional-acceptance position is supported and only the measurement-validity caveats remain.","tokens_in":26855,"tokens_out":24095,"duration_ms":266472,"concrete_test":"Null-calibration of the full discovery pipeline: generate B = 100 surrogate datasets in which each of the 36 Google Trends series is independently circular-shifted by a random offset (preserving marginal distribution and autocorrelation while destroying alignment with the Michigan hate crime series). Run the authors' exact pipeline (Sec. 3.3) on each surrogate, with the same seed and 80/20 split: 5 lags per theory, 36 1D-CNN fits, baseline comparison (Eq. 13), and permutation test with K = 3 (Eq. 14). Record, per surrogate, the number of theories and theory–lag cells passing both criteria, and take the 95th percentile over B. Observed counts are 8 theories / 16 cells before the permutation rule and 6 theories / 12 cells after it.","verdict_should_be":"UNVERDICTED","load_bearing_attack":"The central claim is selected from 180 theory–lag comparisons (36 theories × 5 lags) against a single baseline MAE, computed on one 80/20 chronological split with one seed; under a 5% null rate, ~9 below-baseline cells are expected by chance, and Table 2 flags 16 cells (8 theories). The post-hoc permutation test (Sec. 3.3, Eqs. 9–14) cannot certify these cells because its null is 'the model does not rely on the series' temporal order,' not 'the improvement over baseline is due to chance.' A chance alignment is itself temporal structure; permuting the series destroys it, so PI > 0 is expected even for spurious cells, especially since the test is applied only to already-flagged cells with K = 3 and no null distribution. It compares a model trained and evaluated on original data with models trained and evaluated on permuted data, so it largely restates that the feature contributed to the fitted model. Compounding this, early stopping (Sec. 3.2) selects weights on the same final 20% used to report the 'out-of-sample' errors, and a conspiracy-augmented model has extra capacity to fit validation noise. After the authors' own permutation rule, only six theories retain any validated lag: the sole flagged lags of Obama Kenya (lag 0) and the Great Reset (lag −1) fail, and Q-Anon survives only at lag 3. The abstract's temporal claim ('effects emerging two to three weeks after fluctuations') thus rests mainly on the Great Replacement lags 2 and 3, whose stability across splits is never shown. The keyword-validity concern acknowledged in Sec. 5.1 is real but secondary: it affects the interpretation of an association, whereas the selection problem threatens whether any association exists.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"Using weekly Google search volumes for 36 conspiracy theories in Michigan (2015–2019) as inputs, the paper trains 1D-CNN models to forecast weekly hate crime counts and compares them to a baseline model that uses only crime history and seasonality. For each theory, models are trained at lags −1 to 3. Eight theories produce at least one lag with lower scaled MAE than the baseline, and a feature permutation test is applied to those flagged cells. The authors interpret the positive subset (including the Great Replacement, Rothschilds, and Q-Anon) as evidence that demand for certain conspiracy theories carries incremental predictive signal for hate crimes, with effects appearing two to three weeks after search fluctuations.","tokens_in":27185,"tokens_out":10074,"duration_ms":97352,"significance":"This is an important and timely research question, and the paper demonstrates a thoughtful attempt to bring deep learning to criminological time series. The authors are transparent about several limitations (e.g., search proxies, IP geolocation) and state that the data are publicly available (although the URL is missing). However, the current evidence is not robust enough to support the central claim. The combination of a single 80/20 chronological split, no multiple-comparison correction across 180 cells, early stopping on the same final 20% used for evaluation, and a permutation test whose null hypothesis does not address chance alignment leaves the reported association vulnerable to selection artifacts. The paper would need a substantially more careful validation strategy before its conclusions can be accepted.","major_comments":[{"comment":"The study evaluates 36 theories × 5 lags = 180 predictions against a single baseline MAE. Under a 5% null type I error, about 9 cells are expected to fall below baseline by chance alone, yet Table 2 flags 16 cells (8 theories) with no multiple-comparison correction. The permutation test in Table 3 is applied only to the already-flagged cells, so it cannot mitigate this selection problem. Please report false-discovery-rate-adjusted thresholds or validate the selected theories on an independent temporal holdout.","section":"Section 3.3, Tables 2 and 3"},{"comment":"The final 20% of the time series serves both as the validation set for early stopping and as the test set for the MAE values reported in Tables 2 and 3. The early-stopping rule (patience 15, restoring best weights by validation loss) and the learning-rate reduction select the model state on this same period, and the baseline MAE (12.18) is computed on it as well. As a result, Eq. (13) compares models whose hyperparameters and stopping point were chosen using the evaluation data. A genuinely out-of-sample evaluation would require a last-segment test set that is not used for any training-related decision, or nested/rolling-origin cross-validation.","section":"Sections 3.2 and 3.3"},{"comment":"The permutation test does not test whether the improvement over baseline is reliable. Its null is that the model does not depend on the feature's temporal order; permuting the series destroys any chance alignment with the target, so PI > 0 is expected even for a spurious feature that happens to align in the original data. Moreover, Eq. (10) evaluates a model retrained and retested on the permuted data, so the test largely restates that the feature contributed to the fitted model. With K = 3 and the min-selection in Eq. (11), there is no null distribution and no p-value. I recommend testing the null of no association (e.g., by permuting the target series or using a block bootstrap) on all 180 cells before any selection, and reporting the resulting distribution.","section":"Section 3.3, Eqs. (9)–(14)"},{"comment":"The abstract and conclusions cite Q-Anon as a theory that improves prediction with effects at two to three weeks, but Table 3 shows that Q-Anon's lag-2 cell fails the permutation test (true MAE 12.13 vs. lowest permuted MAE 11.96); the only validated Q-Anon lag is 3. In addition, the Great Reset at lag −1 and Obama Kenya at lag 0 fail the permutation test, so three of the eight flagged theories lose their only flagged lags. The text in Section 4 acknowledges this, but the abstract and conclusion do not. Please revise the summary claims to state explicitly which theories and lags survive the permutation test.","section":"Table 3 and Abstract"},{"comment":"The Google Trends features are central to the interpretation, but several search terms are ambiguous. 'Rothschilds' refers to a real family, 'Tuskegee Syphilis Study' is a real historical event, and 'Obama Kenya' may reflect news coverage of the birther debate rather than conspiracy demand. The authors acknowledge this in Section 5.1 and argue that non-conspiracy searches would bias the estimate downward, but no quantitative assessment is provided. Because the paper's central interpretation is that online demand for conspiracy theories predicts hate crime, this measurement ambiguity should be addressed, for example by comparing results to models with news-cycle controls or by using more narrowly defined search terms.","section":"Section 3.1 and Section 5.1"}],"minor_comments":[{"comment":"'Ethic intimidation' should be 'ethnic intimidation'.","section":"Section 3.1"},{"comment":"The interleaved permuted and true MAE values are difficult to parse; add explicit column labels and a note indicating which value is which.","section":"Section 3.3, Table 3"},{"comment":"The text states the URL is specified in the Methods, but no URL appears in Section 3.1 or elsewhere; please provide the link.","section":"Data Availability"},{"comment":"Hyperparameters for the LSTM and other candidate models are not reported, limiting the reproducibility of the model selection step.","section":"Section 3.2"},{"comment":"The paper uses a single fixed seed for weight initialization; given the small sample, please show that results are stable across multiple seeds (e.g., a small seed-sensitivity table).","section":"Section 3.2"},{"comment":"A few reference names contain encoding artifacts (e.g., 'Müller' appears as 'M¨ uller'); please correct the LaTeX/Unicode encoding.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The manuscript states it is a preprint of a paper with a DOI (10.1007/s10610-025-09629-w). If this is intended as a new submission, the overlap with the published version should be clarified. Otherwise, the statistical issues above need to be resolved in a major revision. The topic is suitable for the journal and the paper could become a useful contribution after re-analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a legitimate extension of the authors' earlier four-theory study to 36 theories, and it is honest about many limitations. But the central claim—that a specific subset of conspiracy searches improves hate-crime prediction—is not established. It rests on a single 80/20 split, a single seed, 180 unadjusted theory-lag comparisons, and a permutation test that cannot distinguish a chance alignment from a real one.\n\nWhat the paper does well: the expansion to 36 theories is a real step up in scope, and the null result for most theories is plausible. I also credit the authors for acknowledging, in Section 5.1, that Google Trends keywords can proxy non-conspiracy searches and that IP geolocation is imperfect. The modeling is transparent—1D-CNN is a reasonable choice—and the literature is engaged seriously. The citation pattern is fine; they build on their own 2024 study and the hate-speech literature without overclaiming novelty.\n\nThe soft spot is load-bearing. With 36 theories × 5 lags, you expect about nine cells to beat the baseline by chance. Table 2 flags sixteen. The post-hoc permutation test (Eqs. 9–14) doesn't rescue those cells: its null is that the feature's temporal order doesn't matter, not that the improvement over baseline is due to chance. A spurious time alignment is exactly the kind of temporal structure the test would preserve. And three flagged cells—Great Reset at lag −1, Obama Kenya at lag 0, Q-Anon at lag 2—fail the authors' own test. The abstract's \"two to three weeks\" claim therefore leans heavily on Great Replacement lags 2 and 3, and we never see whether those survive alternative splits. Early stopping on the same final 20% used to report errors is a further concern, though minor relative to the selection issue.\n\nThe keyword-validity concern is real but secondary: it affects interpretation, not existence. The selection problem affects whether there is an association at all.\n\nFor whom: this is a useful paper for people working on online extremism and crime forecasting, especially as a cautionary example of how easy it is to over-read deep-learning feature-importance results. I would take it seriously in a reading group. Recommendation: send it to peer review, but a good referee should demand multiple chronological splits, corrected comparisons or a pre-registered shortlist, and code plus processed data. With those, the finding could be solidified; without them, the abstract overclaims.","headline":"Legitimate extension and honest about limitations, but the subset finding is a selection artifact risk and the abstract overclaims.","tokens_in":27759,"tokens_out":3003,"would_cite":true,"duration_ms":32490,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Online searches for a handful of racially charged conspiracy theories—notably the Great Replacement, Q-Anon, and Rothschilds—improve machine-learning forecasts of weekly hate crimes in Michigan, with effects appearing two to three weeks…","keywords":["information pollution","conspiracy theories","hate crime","Google Trends","deep learning","1D-CNN","permutation test","Great Replacement"],"falsifier":"Run the same 1D-CNN pipeline on the same 262 Michigan weeks with a matched placebo set of non-conspiracy, non-racial search terms (e.g., 'Rockefellers', 'Olympics', weather terms) and require that the eight validated theories beat the placebo set on permutation-validated error reduction; if placebo terms produce equal or larger improvements, the reported association is likely an artifact of the method rather than of conspiracy content.","tokens_in":26650,"feed_emoji":"📈","tokens_out":11732,"duration_ms":113739,"temperature":0.7,"pith_summary":"This paper asks whether online demand for conspiracy theories is tied to offline hate crimes. Using Google Trends searches for 36 racially and politically charged conspiracy theories in Michigan from 2015 to 2019, the authors train a one-dimensional convolutional neural network to predict weekly reported hate crimes and compare its error to a baseline trained only on crime history and seasonality. They find that eight theories—most clearly the Great Replacement, Q-Anon, and the Rothschilds—repeatedly improve prediction accuracy, with improvements emerging two to three weeks after search fluctuations, while the other 28 show no clear link. The paper interprets this as partial empirical support for the idea that specific conspiracy narratives, rather than conspiracy belief in general, are associated with real-world hate violence, in line with neutralization and differential association theories. If correct, the finding would make a small set of search trends a practical early-warning input for a narrowly defined class of crimes.","feed_headline":"Eight conspiracy search trends improve hate-crime forecasts","feed_subtitle":"Great Replacement, Q-Anon, and Rothschilds show the clearest links, with effects at two-to-three-week lags.","key_machinery":"The load-bearing machinery is a lightweight one-dimensional convolutional neural network (1D-CNN) with three convolutional layers (32, 64, and 128 filters), ReLU activations, dropout, early stopping, and two fully connected layers, trained to forecast hate crime counts four weeks ahead from five-week windows of inputs. For each of 36 conspiracy search series, separate models are trained at lags -1, 0, 1, 2, and 3, and their scaled mean absolute error is compared against a baseline that sees only historical hate crimes and week/month seasonal dummies. To ensure any improvement comes from the time structure rather than incidental numeric properties, the authors rerun each model with the conspiracy series randomly permuted and require the original series to beat the best of three permuted runs (positive permutation importance).","core_discovery":"The paper's central claim is that online search interest in a small set of racially charged conspiracy theories carries incremental predictive signal for weekly hate crime counts in Michigan, beyond what past crime and seasonality alone provide. In a 1D-CNN trained separately for each of 36 Google Trends series, eight theories—Ten Days of Darkness, Obama Kenya, Q-Anon, the Great Replacement, the Tuskegee Syphilis Study, Rothschilds, RAHOWA, and the Great Reset—produced lower held-out mean absolute error than a conspiracy-free baseline at one or more lags. The largest gains came from the Great Replacement, with error reductions of 3.29% and 6.42% at lags of two and three weeks, and the Rothschilds improved forecasts at four of the five tested lags. A feature permutation test that shuffled each search series and retrained the model showed that for most of these eight theories the original time order, not static numerical properties, carried the predictive power. The authors state the relationship is not proven causal and could reflect unmeasured confounders or reverse dynamics, but they argue it aligns with theories of neutralization and differential association.","pith_inferences":["A placebo version of the same pipeline using non-conspiracy, non-racial search terms (e.g., 'Rockefellers', 'Olympics', weather terms) would test whether the predictive gains are specific to conspiracy content or a general artifact of adding any time series to a small-sample CNN.","Because the paper aggregates all bias types, the predictive concentration in antisemitic and anti-replacement narratives suggests a testable disaggregation: antisemitic conspiracy searches should predict antisemitic offenses more strongly than other bias categories.","If believers and news-driven curious searchers are mixed in the Google Trends series, the true effect may be larger than measured; using conspiracy-specific phrases or filtering for co-occurring radical content could separate the two and sharpen the estimated lags."],"forward_implications":["If the claim holds, a handful of conspiracy search terms—not all conspiracy content—can sharpen short-term hate crime forecasting in a state like Michigan.","The two-to-three-week delay between search fluctuations and improved prediction points to a concrete surveillance window for platforms or law enforcement, though the paper stops short of prescribing interventions.","The null results for most theories imply that general interest in conspiracies is not a reliable risk indicator; the content of the narrative matters.","The fact that some theories help only at forward lags (trend shifted earlier) signals that part of the observed association may run from hate crimes to searches, not only from searches to crimes."],"supporting_citations":[{"why":"The direct predecessor, which analyzed four conspiracy-search trends and introduced the feature-contribution approach that this study scales to 36.","marker":"Lo Giudice et al. (2024)"},{"why":"Provides the main empirical precedent that online anti-refugee sentiment predicts crimes against refugees, the template for connecting online expression to offline hate.","marker":"Müller and Schwarz (2021)"},{"why":"Shows anti-Black and anti-Muslim social media posts predict offline racially and religiously aggravated crime; the paper draws on it to justify extending the hate-speech-to-crime link to conspiracy theories.","marker":"Williams et al. (2019)"},{"why":"Supplies neutralization theory, the principal explanation for how conspiracy narratives let individuals justify deviant acts.","marker":"Sykes and Matza (1957)"},{"why":"Supplies differential association theory, the second theoretical frame describing how pro-crime values and rationalizations are transmitted in groups.","marker":"Sutherland (1947)"},{"why":"Documents QAnon-linked violence and the short exposure windows before action, used to motivate the 1D-CNN's sensitivity to local peaks and the interpretation of the delayed effect.","marker":"Amarasingam and Argentino (2020)"},{"why":"Origin of the permutation-importance idea that the paper adapts into a retrained feature-permutation test for time series.","marker":"Breiman (2001)"},{"why":"Provides methodological grounding for permutation methods in time-series analysis, supporting the shuffled-series test.","marker":"Cánovas and Guillamón (2009)"},{"why":"Supports the premise that feature importance from predictive models can signal causal relationships in time-series settings.","marker":"Castro et al. (2023)"},{"why":"Provides the standard definition of conspiracy theories used to frame the study.","marker":"Douglas et al. (2019)"}],"fun_headline_variants":["Eight conspiracy search trends signal future hate crimes","Q-Anon and Great Replacement searches forecast hate crimes","Online conspiracy searches hint at hate crime patterns weeks later","Search trends for 8 conspiracy theories improve hate crime forecasts","Google searches for conspiracy theories predict hate crime incidents"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The inference that below-baseline prediction error reflects a conspiracy-to-violence link assumes that Google Trends search volumes for terms like 'Rothschilds' or the 'Tuskegee Syphilis Study,' geolocated to Michigan by IP address, measure genuine conspiracy-theory demand rather than news-driven curiosity, historical reference, or unrelated uses of the same words.","fun_headline_variants_meta":{"raw":{"variants":["Eight conspiracy search trends signal future hate crimes","Q-Anon and Great Replacement searches forecast hate crimes","Online conspiracy searches hint at hate crime patterns weeks later","Search trends for 8 conspiracy theories improve hate crime forecasts","Google searches for conspiracy theories predict hate crime incidents"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000442,"raw_usage":{"total_tokens":2227,"prompt_tokens":918,"completion_tokens":1309,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":534,"completion_tokens_details":{"reasoning_tokens":1235}},"tokens_in":534,"tokens_out":1309,"duration_ms":12709,"temperature":1.0,"reasoning_tokens":1235,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:30:33.541162+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same 1D-CNN pipeline on the same 262 Michigan weeks with a matched placebo set of non-conspiracy, non-racial search terms (e.g., 'Rockefellers', 'Olympics', weather terms) and require that the eight validated theories beat the placebo set on permutation-validated error reduction; if placebo terms produce equal or larger improvements, the reported association is likely an artifact of the method rather than of conspiracy content.","supporting_citations":[{"cited_title":", Shadman Yazdi, A","cited_arxiv_id":null,"evidence_quote":"The direct predecessor, which analyzed four conspiracy-search trends and introduced the feature-contribution approach that this study scales to 36."},{"cited_title":", Burnap, P","cited_arxiv_id":null,"evidence_quote":"Shows anti-Black and anti-Muslim social media posts predict offline racially and religiously aggravated crime; the paper draws on it to justify extending the hate-speech-to-crime link to conspiracy theories."},{"cited_title":"\\ Matza, D","cited_arxiv_id":null,"evidence_quote":"Supplies neutralization theory, the principal explanation for how conspiracy narratives let individuals justify deviant acts."},{"cited_title":", Mendes Júnior, P R","cited_arxiv_id":null,"evidence_quote":"Supports the premise that feature importance from predictive models can signal causal relationships in time-series settings."},{"cited_title":", Uscinski, J E","cited_arxiv_id":null,"evidence_quote":"Provides the standard definition of conspiracy theories used to frame the study."}],"review_version":1}