{"id":"d1ce6499-c87a-4924-867c-141f6a8326e2","arxiv_id":"2502.12161","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic review of 141 AI-based earthquake forecasting studies concludes that most oversimplify the problem and only a few demonstrate value against strong seismological benchmarks.","lead":"This review of 141 AI-based earthquake forecasting studies finds that most are not tested against strong statistical seismology baselines and oversimplify the forecasting task. It argues that pairing AI with geophysical knowledge, proper benchmarks, and realistic data handling is the credible path forward.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The review's central promise of gains from 'data from many varied sources' rests on an unvalidated assumption that non-seismic precursors carry predictive information beyond the seismic catalog; the paper itself cites criticism of this premise.","rationale":"The reader's weakest_assumption identifies exactly the same load-bearing concern: the review presupposes that non-seismic and non-mechanical data contain predictive information not already in seismic catalogs. This is the right spot because the review's central proposal has two pillars: (1) benchmark AI against strong statistical seismology models and (2) expand inputs to diverse non-seismic data. Pillar (1) is well supported by the six cited studies, and the review's diagnosis of weak baselines and oversimplification is consistent with Mignan and Broccardo. Pillar (2) is asserted in the abstract and Section 1 but never tested; the paper's own Section 2.2 acknowledges the lack of a clear mechanism and of rigorous statistical validation for such precursors. The survey data in Section 4.4.3 show sparse use of non-seismic inputs (27 of 141 studies) and no evidence of incremental skill. Thus the 'enhance predictive accuracy' half of the central claim rests on an unvalidated empirical assumption. The proposed concrete test is directly tied to the review's own benchmark framework (Section 4.2) and would settle whether the premise holds in the most favorable setting. The review is a perspective/review piece, so a CONDITIONAL verdict with this stated caveat is appropriate; the concern does not require rejection, only explicit acknowledgement of the unvalidated premise. Hence the verdict should remain UNCHANGED relative to the reader's conditional acceptance, and the agreement is complete.","tokens_in":45319,"tokens_out":2814,"duration_ms":30507,"concrete_test":"Rebuild the Section 4.2 pseudo-prospective benchmark (RELM polygon, 2010–2020 test period; FCN and ETAS baselines with reported Area Skill Scores) and add the most commonly cited non-seismic inputs—radon, TEC, outgoing longwave radiation, and geomagnetic variations—as auxiliary channels, with strict temporal separation of training and test data. If Area Skill Scores for M≥4 and 30/60-day windows do not improve over the seismic-only FCN and ETAS baselines, the incremental-information premise underlying the review's 'diverse data sources' recommendation is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim (Section 1) is that integrating AI with 'data from many varied sources augmented by geophysical knowledge' will enhance predictive accuracy and reveal precursor mechanisms. The 'many varied sources' half of this claim depends on the assumption that ionospheric TEC, electromagnetic emissions, thermal infrared anomalies, radon, groundwater, and similar non-seismic inputs (I54–I73, Section 4.4.2) contain incremental predictive information. The review never validates this assumption: Section 2.2 explicitly notes 'the lack of a clear physical mechanism linking non-seismic precursors... has led to widespread criticism' and 'lack of rigorous statistical testing methodology' for proposed precursors. Furthermore, the survey's own Section 4.4.3 shows only 27 of 141 studies used non-seismic inputs, and none of the six highlighted benchmark-comparison studies demonstrates that non-seismic variables improve forecast skill after conditioning on the past catalog. If non-seismic data carry no information beyond seismicity, the recommended expansion of inputs and the anticipated discovery of precursor mechanisms lose their principal expected benefit. The other half of the proposal—injecting seismological structure into models, as in Zlydenko et al. and Zhan et al.—is independently supported, so the load-bearing weakness is specifically the multi-source data premise, not physics-informed AI in general.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This review surveys 141 (or, inconsistently, 142/140/136) AI-based earthquake prediction and forecasting studies published from 1994 to late 2024 and assesses them from a seismological perspective. It catalogs model outputs, input types, loss functions, and evaluation metrics; argues that most surveyed studies oversimplify the forecasting problem and evaluate against weak baselines; and emphasizes that meaningful progress requires benchmarking against strong statistical-seismology models such as ETAS and injecting seismological structure into models and loss functions. The authors single out six AI studies that compare against geophysical or statistical-seismology benchmarks, and they present their own FCN-vs-ETAS comparison on California data as an additional reference framework. The paper concludes with recommendations for pseudo-prospective testing, awareness of catalog incompleteness, and seismologically informed loss-function and feature design.","tokens_in":45555,"tokens_out":2948,"duration_ms":29796,"significance":"The review fills a useful niche: it gives AI researchers a concrete, well-organized map of earthquake-forecasting output types, inputs, losses, and metrics (Sections 4.3-4.6), and it forcefully restates the benchmarking message of Mignan and Broccardo with additional 2023 examples. It also ships reproducible assets: the GitHub corpus list, the FCN benchmark code/data link, and the explicit acknowledgment of short-term catalog incompleteness in Section 5.3 are genuine strengths. However, the paper's headline promise—that adding 'data from many varied sources' (especially non-seismic precursors) will materially improve forecasting and reveal precursor mechanisms—is not actually supported by the survey evidence it presents. The six highlighted ETAS-comparison studies all use seismic inputs; none demonstrates that non-seismic variables add skill beyond a catalog-based baseline. The paper itself cites the lack of physical mechanism and of rigorous statistical testing for non-seismic precursors (Section 2.2). Thus the contribution is strongest as a critical methodology review and a physics-informed-AI agenda, but weaker as an evidence-based case for multi-source precursor integration.","major_comments":[{"comment":"The surveyed corpus size is stated inconsistently: Section 1 states 142 papers, Section 2.3 states 141, Section 3.3 refers to 'a review of 136 papers' and later '140 studies', and Section 4.4.3 says 'Out of 141 investigated works.' The percentages in Section 2.3 (66.9% compared to a baseline) and the claim that only six studies provide meaningful benchmark comparisons depend on a precisely defined corpus. Please reconcile these numbers and provide the survey selection protocol (databases, search terms, inclusion/exclusion criteria) so that the representativeness of the corpus can be assessed.","section":"Section 1 and Section 2.3"},{"comment":"The sentence 'four studies from 2023 compared their models with reasonably strong geophysical models, six of which made comparisons with some versions of the general class of ETAS models' is internally contradictory: four and six cannot both be the denominator and the subset. The following paragraph in Section 3.3 then lists six highlighted studies, several of which are not from 2023. Please clarify which studies are being counted, with references and years, and align the two counts.","section":"Section 2.3, paragraph 2"},{"comment":"Table 1 is referred to as reporting Area Skill Scores for the ETAS and FCN models across 12 time-magnitude windows, plus a runtime comparison, but the table body appears empty in the manuscript: no numerical scores or speed values are given. Since the FCN-versus-ETAS pseudo-prospective experiment is presented as a central benchmark contribution and as evidence for 'performance parity with speed,' the absence of the actual numbers makes the claim unverifiable. Please include the full Table 1 with all Area Skill Scores and runtime measurements, and check that the GitHub link resolves to these results.","section":"Section 4.2, Table 1"},{"comment":"The load-bearing claim that integrating 'data from many varied sources' (ionospheric TEC, electromagnetic emissions, thermal infrared, radon, groundwater, etc.) will enhance predictive accuracy and uncover earthquake-precursor mechanisms is not supported by the surveyed evidence. Section 2.2 itself notes the lack of a clear physical mechanism and of rigorous statistical testing for non-seismic precursors; Section 4.4.3 shows that only 27 of 141 studies used non-seismic inputs; and none of the six highlighted benchmark-comparison studies demonstrates that non-seismic variables improve forecast skill after conditioning on past seismicity. Stronger versions of this claim should be reframed as an open hypothesis, with explicit experimental tests proposed (e.g., ablation or Shapley-value analysis against an ETAS-conditioned baseline). The physics-informed-AI half of the proposal is independently supported by the six highlighted studies and should be separated from the multi-source-precursor half.","section":"Sections 1, 4.4.2-4.4.3, and 5.5"},{"comment":"The review states that 'none [of the 136 papers] had conducted prospective testing' and that the six highlighted studies were evaluated pseudo-prospectively. This is a strong negative claim about the entire corpus. Given the corpus-size inconsistencies noted above, please either verify this claim with a per-paper testing-mode table in the Supplement or soften the statement to 'none of the papers we could verify' with a clear audit trail.","section":"Section 2.2 and Section 3.3"}],"minor_comments":[{"comment":"The abstract says 'precursors' and 'many varied sources' before the body has established the evidentiary status of those sources; consider aligning the abstract with the more cautious discussion in Section 2.2.","section":"Abstract"},{"comment":"The text reads 'This ANN model performs on par with, or even surpasses, a standard ETAS model with isotropic spatial kernels ... in terms of the average information gain per earthquake'; please specify whether the comparison is on Japanese data for all magnitude thresholds and report the exact information-gain values that support 'on par' versus 'surpasses.'","section":"Section 3.3, paragraph on Zlydenko et al."},{"comment":"Figure 2 is described as a cumulative frequency of output types, but the axes and the exact cumulative definition are not explained in the caption. Please add a self-contained caption with axis labels and the counting convention.","section":"Section 4.3.10, Figure 2"},{"comment":"Several equations have notational typos that could confuse readers: Eq. (3) for Magnitude Density omits the volume factor in the text definition, Eq. (11) places the denominator inside the summation in the prose, and Eq. (18) appears to define eta as a sum rather than an average. Please proofread all input equations against the cited sources.","section":"Section 4.4, Input definitions"},{"comment":"The L-test, S-test, and N-test definitions use oi,j both for the predicted seismicity and later for the forecasted number; please rename the predicted rate to avoid confusion with the observed count omega.","section":"Section 4.6.3, Eq. (90)-(96)"},{"comment":"The phrase 'fast-food research' is informal for a Physics Reports review; consider replacing it with a neutral descriptive term such as 'superficial research practice.'","section":"Section 5.1"},{"comment":"Some references are incomplete or carry placeholder text: the citation for 'Liu et al. (reference)' in the Input 14 definition and the truncated reference in Section 5.5 need to be completed.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is a broad review and will likely attract interest, but the editorial bar for Physics Reports presumably requires the survey's central claims to be internally consistent and evidence-backed. The corpus-size inconsistencies and the empty Table 1 are easy-to-fix but currently undermine the paper's own quantitative claims. More substantively, the 'many varied sources' promise needs to be either evidenced or explicitly demoted to a research hypothesis; otherwise the review overstates what the surveyed literature shows. I would support major revision rather than rejection, because the physics-informed-AI benchmarking message is valuable and the taxonomy sections are useful even if the multi-source precursor premise is not yet established."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Dear colleague,\n\nRead this one if you care about the AI-seismology interface. The paper's core diagnosis is right: most AI-based earthquake forecasting studies benchmark against other ML models or trivial Poisson nulls, ignore clustering and class imbalance, and rarely compare to serious statistical seismology baselines like ETAS. The six studies the authors single out that did run ETAS comparisons form a genuinely informative list, and the taxonomy of 20 output types and 73 input types is a practical contribution beyond Mignan and Broccardo's earlier review. The guidance on pseudo-prospective testing, catalog incompleteness, and loss functions is sound and concrete.\n\nThe soft spots are real but not fatal. The paper gives inconsistent corpus counts (142, 141, 140, 136) and never discloses its literature-selection protocol; for a review, that is a reproducibility gap. Table 1, which is supposed to carry the Area Skill Scores for the ETAS-vs-FCN benchmark, arrives without the numbers in the manuscript text—either an image, a missing table, or a placeholder; a referee should insist on seeing those numbers plus code with a commit hash. Absolute statements like 'none had conducted prospective testing' and 'AI becomes indispensable' need moderation; the 'none' is especially hard to verify given the selection protocol is missing.\n\nThe biggest conceptual weakness is the promise attached to non-seismic inputs (ionospheric, electromagnetic, radon, and so on). The opening sells 'data from many varied sources' as the route to new insight, and Section 5.5 recommends expanding into those variables. But the paper itself cites the lack of a clear physical mechanism and the absence of rigorous statistical testing for these precursors, and none of the six ETAS-comparison studies uses non-seismic inputs. So the claim that these data carry incremental predictive information is asserted, not argued. This doesn't sink the review, because the physics-informed-AI half of the program is independently supported; but the multi-source half is speculative and should be labeled as such.\n\nI'd send this to a serious referee. The taxonomy, the benchmark critique, and the seismological guidance are worth having in the literature; the reporting problems and the overreach on precursors are fixable in revision.","headline":"A useful, well-organized review of AI for earthquake forecasting whose core benchmarking message holds up; the multi-source non-seismic data promise is its weakest leg.","tokens_in":46087,"tokens_out":2627,"would_cite":true,"duration_ms":23341,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"AI earthquake forecasting will not advance by swapping one neural architecture for another; it needs rigorous benchmarking against physics-based models such as ETAS and the injection of seismological knowledge into features, network…","keywords":["Statistical physics","Earthquake forecasting","Statistical seismology","Machine learning","Geophysics","Deep learning","ETAS benchmark","Pseudo-prospective testing"],"falsifier":"Run a prospective test under standard statistical-seismology protocols in California in which the same model is trained once on seismicity-only features and once with ionospheric, radon, electromagnetic, and thermal-infrared channels added, then scored with the Molchan area skill and the L-test against the inhomogeneous ETAS benchmark; if the augmented model does not beat the seismicity-only model beyond sampling noise over a multi-year window, the paper's premise about non-seismic precursors loses its empirical footing.","tokens_in":45086,"feed_emoji":"🌍","tokens_out":8401,"duration_ms":72368,"temperature":0.7,"pith_summary":"This review argues that most AI-based earthquake forecasting research evaluates models against the wrong baselines. Across the 141 surveyed studies, roughly a third include no baseline comparison at all, and most that do compare only against other machine-learning methods or a Poisson null hypothesis, not against the best statistical-seismology benchmarks such as the ETAS model. The paper's central claim is that useful progress will come less from novel architectures than from benchmarking AI against physics-based forecasting models and from building seismological structure into inputs, network design, and loss functions. It presents a taxonomy of model outputs, inputs, loss functions, and evaluation metrics, and uses an 11-year pseudo-prospective comparison in Southern California to show that a fully convolutional network can match ETAS skill while running 2,000 to 4,000 times faster.","feed_headline":"AI earthquake forecasts need real seismology benchmarks","feed_subtitle":"Only six of 141 AI studies compare against strong statistical models like ETAS; the review shows why that matters.","key_machinery":"The machinery carrying the argument is the comparison against the ETAS model, the epidemic-type aftershock sequence model from statistical seismology, together with a proposed pipeline of physics-informed design choices: feature engineering based on magnitude-frequency statistics and aftershock decay parameters, network architectures that embed the triggering structure of ETAS, and loss functions that reweight samples according to seismological priors such as magnitude-frequency balance. The paper uses its own Southern California benchmark, run on a 0.1-degree grid over 11 years with the ETAS model as reference, to demonstrate that a fully convolutional network reaches ETAS-level skill with far lower computational cost. That benchmark, along with the six comparable studies, carries the claim that AI contributes to earthquake forecasting mainly through speed parity, supremacy, or complementary discovery rather than through architectural novelty alone.","core_discovery":"On the paper's own terms, the central discovery is that the field's bottleneck is evaluation and integration, not AI capability. Only six of the reviewed studies compare an AI model with a reasonably strong statistical-seismology model, and those six show a consistent pattern: models that copy the mathematical structure of ETAS, or that learn a flexible neural point process, can match or exceed ETAS in forecasting skill while cutting computation time by orders of magnitude. The review uses these cases to argue that AI will earn its place in seismology in one of three ways: performing on par with leading geophysical models but far faster, surpassing them, or revealing new physical patterns. It also demonstrates that many studies treat earthquake forecasting as a plain binary classification or regression problem, ignoring the severe imbalance between earthquake and non-earthquake samples and the strong spatio-temporal clustering of seismic events.","pith_inferences":["A testable consequence the review leaves implicit is that if non-seismic precursors carry independent signal, the improvement should appear first in forecasting background and mainshock events, since ETAS and its AI imitators already capture triggered aftershocks well.","A meta-analysis of the surveyed studies might show that models using seismological feature engineering report smaller gains over strong baselines, because those baselines already exploit the same features; that would sharpen the claim that integration, not input diversity, is the bottleneck.","By the review's own standard, AI's most defensible immediate contribution is operational speed: even at prediction parity, replacing a slow ETAS calibration with a fast neural model makes real-time ensemble hazard mapping feasible.","The review does not assess whether transformer-based or foundation models trained on large multi-modal datasets would change its benchmark conclusions; that remains an open extension of its argument."],"forward_implications":["Future AI earthquake forecasts should be judged by whether they match or beat ETAS-class models, not by gains over Poisson nulls or other machine-learning baselines.","Standardized outputs, such as the probability of at least one earthquake above a magnitude threshold in a space-time bin, would make models directly comparable.","Loss functions that encode seismological weighting, balancing rare large events, background versus triggered events, and spatio-temporal clustering, will matter more than network depth.","Models that embed the triggering structure of ETAS can retain geophysical interpretability while running roughly a thousand times faster, making near-real-time operational forecasting more practical.","Pseudo-prospective testing must account for short-term catalog incompleteness; otherwise, comparisons between AI models and ETAS are unreliable."],"supporting_citations":[{"why":"Supplies the prior survey finding that most ANN earthquake-forecasting studies lack meaningful baselines and have not demonstrated new predictive insights.","marker":"[12]"},{"why":"Shows a single-neuron logistic regression reproduces the results of a deep aftershock-forecasting network, establishing that simple statistical baselines can match elaborate AI.","marker":"[14]"},{"why":"Is the deep-learning aftershock study reanalyzed in [14], exemplifying the architecture-driven work the review criticizes.","marker":"[13]"},{"why":"Provides the testing principles (clear parameterization, objective success criteria, prior-data use) and operational forecasting benchmarks the review adopts as standards.","marker":"[11]"},{"why":"Defines the ETAS model class that serves as the main geophysical benchmark throughout the review.","marker":"[17]"},{"why":"Supplies the fully convolutional network model and the 11-year Southern California pseudo-prospective comparison showing ETAS-level skill with far lower computation.","marker":"[53]"},{"why":"Demonstrates an ANN encoder that imitates the mathematical structure of ETAS, achieving parity or better with a 1000-fold speedup and learning anisotropic spatial kernels.","marker":"[49]"},{"why":"Documents short-term catalog incompleteness, which undermines naive AI-versus-ETAS comparisons in aftershock forecasting.","marker":"[123]"}],"fun_headline_variants":["AI quake forecasts need seismology-grade benchmarks","Only 6 of 141 AI quake studies benchmark against ETAS","AI earthquake forecasting lacks strong baseline comparisons","Why AI quake prediction needs geophysics benchmarks","Earthquake AI predictions need seismology-informed baselines"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The proposal assumes that non-seismic and non-mechanical precursor signals, such as ionospheric total electron content, electromagnetic emissions, thermal infrared anomalies, radon, and groundwater levels, carry information about future earthquakes that is not already contained in the seismic catalog; if they do not, widening the input space will not improve forecasts.","fun_headline_variants_meta":{"raw":{"variants":["AI quake forecasts need seismology-grade benchmarks","Only 6 of 141 AI quake studies benchmark against ETAS","AI earthquake forecasting lacks strong baseline comparisons","Why AI quake prediction needs geophysics benchmarks","Earthquake AI predictions need seismology-informed baselines"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.0005,"raw_usage":{"total_tokens":2453,"prompt_tokens":960,"completion_tokens":1493,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":576,"completion_tokens_details":{"reasoning_tokens":1418}},"tokens_in":576,"tokens_out":1493,"duration_ms":9960,"temperature":1.0,"reasoning_tokens":1418,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-08T14:25:34.282952+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run a prospective test under standard statistical-seismology protocols in California in which the same model is trained once on seismicity-only features and once with ionospheric, radon, electromagnetic, and thermal-infrared channels added, then scored with the Molchan area skill and the L-test against the inhomogeneous ETAS benchmark; if the augmented model does not beat the seismicity-only model beyond sampling noise over a multi-year window, the paper's premise about non-seismic precursors loses its empirical footing.","supporting_citations":[{"cited_title":"Mignan, M","cited_arxiv_id":null,"evidence_quote":"Supplies the prior survey finding that most ANN earthquake-forecasting studies lack meaningful baselines and have not demonstrated new predictive insights."},{"cited_title":"Mignan, M","cited_arxiv_id":null,"evidence_quote":"Shows a single-neuron logistic regression reproduces the results of a deep aftershock-forecasting network, establishing that simple statistical baselines can match elaborate AI."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Is the deep-learning aftershock study reanalyzed in [14], exemplifying the architecture-driven work the review criticizes."},{"cited_title":"Ogata, Statistical models for earthquake occurrences and resid- ual analysis for point processes, J","cited_arxiv_id":null,"evidence_quote":"Defines the ETAS model class that serves as the main geophysical benchmark throughout the review."},{"cited_title":"Zlydenko, G","cited_arxiv_id":null,"evidence_quote":"Demonstrates an ANN encoder that imitates the mathematical structure of ETAS, achieving parity or better with a 1000-fold speedup and learning anisotropic spatial kernels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Documents short-term catalog incompleteness, which undermines naive AI-versus-ETAS comparisons in aftershock forecasting."}],"review_version":1}