{"id":"aca82875-e533-4c27-bb30-4e404569fff6","arxiv_id":"2607.24109","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A conditional diffusion model that forecasts aftershock rate and maximum magnitude as spatial fields outperforms the USGS Reasenberg-Jones model globally and matches tuned ETAS on southern California daily forecasts.","lead":"This paper trains a diffusion model, QuakeGen, to generate maps of how many aftershocks and how large they will be in each location after a big earthquake, conditioned on seismicity already recorded. It reports better spatial forecasts than the USGS's operational model on global sequences and comparable daily forecasts to a tuned ETAS model in southern California.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Global comparison leaks the first 1.6 h of target aftershocks into QuakeGen's conditioning, inflating spatial skill at every horizon; re-scoring with target truncated at tc is needed.","rationale":"The central claim has two parts: global outperformance and regional parity with ETAS. The regional part is credible: the regional model has no post-anchor conditioning, uses a held-out split, and matches a tuned ETAS on multiple metrics. The global part is the load-bearing claim for 'outperforms,' and it is compromised by the target-conditioning overlap described above. I agree with the reader's weakest assumption and would sharpen it: the contamination is not limited to short horizons. Because Omori-type decay places a large share of events in the first hours, the 1.6 h conditioning window is a non-negligible fraction of scored events even at 168 h and 720 h, so the long-horizon global skill metrics are also suspect until re-scored. The proposed check—truncating the target to (t0+1.6h, t0+T]—directly isolates this effect. If the QuakeGen advantage persists, the central claim survives with a short-horizon caveat; if it disappears, the global claim is not established. The paper's other gaps (no confidence intervals, no quantitative max-magnitude evaluation) are real but secondary. The manuscript is honestly written, code/weights are available, and the framework is promising; the issue is addressable. Therefore the existing CONDITIONAL verdict is appropriate and no verdict change is needed.","tokens_in":42440,"tokens_out":8998,"duration_ms":79242,"concrete_test":"Recompute all global rows of Table 2 with the target restricted to events with te > t0 + 1.6 h for each horizon, leaving QuakeGen's forecasts unchanged, and compare IG, W1, MSESS, and Npred/Nobs against USGS-RJ. Also tabulate the fraction of scored events falling inside (t0, t0+1.6 h] per horizon. If QuakeGen's advantage is eliminated or reversed at several horizons, the headline 'outperforms' claim is unsupported; if it persists at 168 h and 720 h, the long-horizon spatial claim survives and only the short-horizon numbers need caveats.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim rests on the global comparison with USGS-RJ (Table 2), but the evaluation is contaminated by an information leak. The target is defined in Eq. 1 as all events in (t0, t0+T], yet the global conditioning (Section 3.1, Table 1) includes post-anchor windows ending at tc = 1.6 h. Thus every event in the first 1.6 h after the mainshock is both an input channel and part of the scored target for every horizon. Because QuakeGen receives per-cell log-count and max-magnitude fields over these windows, it can reproduce those events nearly exactly, whereas USGS-RJ, which also uses early aftershocks as sources, still smooths them through its kernel. The effect is not confined to 3 h/12 h: the early aftershock rate is highest, so the 1.6 h interval contains a large fraction of the events scored at 48 h, 168 h, and even 720 h. The reported QuakeGen advantage in IG, W1, and MSESS could be driven by these directly observed events. A clean evaluation must score only events after the conditioning cutoff, or equivalently move the forecast issue time to t0 + tc and score (t0+tc, t0+T].","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces QuakeGen, a conditional denoising diffusion model that recasts aftershock forecasting as generation of gridded log-count and maximum-magnitude fields. The model is trained once on global mainshock–aftershock sequences (1990–2023, tested on 2024–2025) and on the QTM southern California catalog (2009–2015, tested on 2016–2017), and is compared with the operational USGS Reasenberg–Jones model and with regionally tuned ETAS. The central claims are that QuakeGen outperforms USGS-RJ on global sequences, recovers fault-aligned anisotropic structure that isotropic kernels cannot express, and matches ETAS on regional daily forecasting without per-sequence or per-region refitting. The paper also presents a cross-catalog experiment in which the global model is applied to the enhanced QTM catalog without retraining.","tokens_in":42746,"tokens_out":7369,"duration_ms":67404,"significance":"If the global result is confirmed, this is a meaningful contribution: a single generative model, trained without per-sequence or per-region refitting, producing probabilistic spatial forecasts and beating an operational baseline on held-out years. The evaluation design has genuine strengths: a strictly held-out global test period (2024–2025), held-out regional years (2016–2017), public code and archived forecast outputs, and a regional comparison against benchmark-tuned ETAS. However, the global evaluation has a scoring/conditioning overlap: events in the first 1.6 h after the mainshock are both conditioning inputs and part of the scored target for every horizon, so QuakeGen can partially copy the target while USGS-RJ cannot. This contaminates the headline global comparison in Table 2 and Fig. 6. The regional result is unaffected by this overlap and provides independent support for the framework. The paper is also honest about its main limitation—reliance on large training catalogs—and shows a clear failure case (Kamchatka).","major_comments":[{"comment":"The global target in Eq. (1) is all events in (t0, t0+T], while the conditioning includes post-anchor windows ending at tc = 1.6 h (Table 1). For the 3 h and 12 h horizons, 53% and 13% of the scored window are also input channels; for longer horizons, the early high-rate interval still contains a large fraction of the scored events. Because QuakeGen receives per-cell log-count and max-magnitude fields over (0, 1.6 h], it can reproduce those events nearly exactly, whereas USGS-RJ uses the same early aftershocks only as smoothed kernel sources. This inflates the QuakeGen advantage in Table 2 and Fig. 6 at every horizon. The stress-test concern lands. Re-score with the target truncated to (t0+tc, t0+T] or, equivalently, move the issue time to t0+tc and report all metrics on that clean target. Also report sensitivity to tc.","section":"§3.1, Table 1, Eq. (1)"},{"comment":"The text says that when USGS and ISC both record the same earthquake, the authors 'keep both records as separate sequences rather than choosing between them.' If the same mainshock appears in both catalogs, it can be scored twice as two of the 80 test sequences, and its aftershocks can be duplicated in the observed and predicted fields. This would bias the per-event metrics in Table 2 and Fig. 6 and make the 80 test sequences non-independent. Please clarify the de-duplication procedure, and if duplicates exist, merge or remove them and recompute the global comparison.","section":"§2.3, global dataset construction"},{"comment":"The global comparison is also confounded by an input asymmetry: QuakeGen conditions on 768 h of pre-mainshock seismicity and on the first 1.6 h of aftershocks, whereas USGS-RJ starts at the mainshock and uses only the first 1.6 h of aftershocks as kernel sources. This is acknowledged in S1.2 but not controlled. To attribute the gain to learned spatiotemporal structure rather than to additional input information, the paper should include an ablation—e.g., QuakeGen without the pre-mainshock windows, or an RJ variant with a background term—or otherwise qualify the causal claim.","section":"§S1.2, §3.1"}],"minor_comments":[{"comment":"The scaling constants (α, β, γ), the conditioning-window schedule, and the cutoff tc are chosen ad hoc, as the paper states for tc. The paper should explicitly list these as hyperparameters and, if possible, give a small sensitivity analysis; otherwise the 'no prescribed functional form' claim should be softened.","section":"§2.1/Table 1"},{"comment":"The text says QuakeGen 'matches' ETAS regionally, but Table 2 and Fig. 10 show QuakeGen consistently higher in TLL and IG and closer to a productivity ratio of 1. Consider saying 'slightly outperforms' or 'is competitive with' rather than 'matches.'","section":"§3.2/Fig. 10/Table 2"},{"comment":"The maximum-magnitude channel is a claimed output of the model but is only shown qualitatively. If it is part of the contribution, add a quantitative evaluation (e.g., MAE or rank histogram for the P90 field, or a proper score for the magnitude channel). Otherwise state that the field is illustrative.","section":"§3.1/Figs. 7, S7–S29"},{"comment":"The M 4.4 San Jacinto sequence used in the cross-catalog experiment is below the global model's M ≥ 4.5 training range. This should be mentioned in the main text when drawing conclusions about generalization to enhanced catalogs.","section":"§3.3/Fig. S18"},{"comment":"The acknowledgements contain a personal memorial and a passage in another language. This is unusual for a journal article and should be moved to a footnote or removed. The AI-use declaration is appropriate and should stay.","section":"Acknowledgements"}],"recommendation":"major_revision","confidential_remarks":"The global comparison is the hinge of the paper. The overlap between the conditioning window and the scored target is a real, fixable flaw, not a matter of taste: the Table 2 IG/W1/MSESS advantages could shrink substantially after truncating the target at tc. I would not reject because the regional benchmark and the qualitative fault-alignment examples are independent of this leak, but the headline claim must be re-estimated on a clean target. Please also check the USGS/ISC duplicate-sequence issue before resubmission, as it affects the effective test-sample size and the independence of the per-event scores."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Zhu's QuakeGen paper is worth a look, but the headline global comparison has a real evaluation leak that most readers will miss. The forecast target in Eq. 1 is all events in (t0, t0+T], while the global conditioning includes the first 1.6 hours of aftershocks after t0. So for every horizon, QuakeGen is scored on a target that partly contains its own input. At 3 h, 53% of the window is conditioning; at 12 h, 13%; and because aftershock rates are highest right after the mainshock, the first 1.6 h contributes a large share of the events scored at 48, 168, and 720 h as well. USGS-RJ also sees those early events as sources, but it smooths them through an isotropic kernel; QuakeGen can effectively copy the per-cell counts and magnitudes it was fed. That likely explains the large short-horizon gains in IG, W1, and MSESS. The fix is simple: issue the forecast at t0+tc and score only events after tc.\n\nDon't get me wrong: the paper is not a throwaway. It is, as far as I can tell, the first to frame aftershock forecasting as conditional generation of a two-channel field (count and max magnitude) with a diffusion model, trained once and applied across sequences and regions. The regional comparison is clean—no post-anchor conditioning—and a single QuakeGen matches region-tuned ETAS on the 2016–2017 QTM data. That is a real result. The code and weights are public, the global test set (2024–2025) is genuinely held out, and the writing is honest about scope.\n\nThe other soft spots are proportionate. The global comparison has a second asymmetry: QuakeGen gets 768 h of pre-mainshock background, which USGS-RJ does not use. The reported medians lack confidence intervals or significance tests; with 80 global events, paired bootstrap is easy. And the max-magnitude channel is never scored quantitatively—only shown as maps. All of those are fixable.\n\nBottom line: the regional result and the framework deserve a serious referee. The paper should not be accepted as is. It needs a leakage-free global scoring, uncertainty estimates, and at least one quantitative evaluation of the max-magnitude forecasts. If that gets done, the long-horizon spatial advantage may survive in weaker form, but the current headline is overstated.","headline":"A promising diffusion-based aftershock forecaster, but the global comparison with USGS-RJ is contaminated because the first 1.6 h of scored aftershocks are also conditioning inputs; the clean regional result is what should carry the paper.","tokens_in":43246,"tokens_out":6260,"would_cite":true,"duration_ms":52490,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A single diffusion model, trained once on seismic catalogs, can forecast aftershock fields that beat the operational USGS Reasenberg-Jones model globally and match region-tuned ETAS on daily regional seismicity.","keywords":["aftershock forecasting","conditional diffusion model","spatiotemporal fields","Reasenberg-Jones model","ETAS","earthquake catalogs","fault-aligned anisotropy","forecast ensemble"],"falsifier":"Re-score the global 3-hour and 12-hour forecasts using only events that occur after the conditioning cutoff of 1.6 hours, or retrain QuakeGen without the post-anchor conditioning windows. If its spatial advantage over USGS-RJ disappears, the claimed fault-aligned anisotropy gain is largely an artifact of the overlap rather than a learned forecasting skill.","tokens_in":42277,"feed_emoji":"🌋","tokens_out":3170,"duration_ms":30499,"temperature":0.7,"pith_summary":"This paper tries to establish that aftershock forecasting is better cast as conditional generation of a spatiotemporal field than as event-by-event point-process modeling. It introduces QuakeGen, a diffusion model trained once on seismic catalogs, which generates the future field of earthquake counts and maximum magnitudes conditioned on recent seismicity and the forecast horizon. On 80 held-out global mainshock sequences from 2024-2025, QuakeGen outperforms the operational USGS Reasenberg-Jones forecast, reproducing fault-aligned, anisotropic spatial patterns that fixed isotropic kernels cannot express. On southern California daily seismicity, the same single model matches separately tuned ETAS baselines on the regional benchmark. If this holds, a learned generative model could replace per-sequence, per-region calibration in operational forecasting and improve as denser catalogs accumulate.","feed_headline":"Diffusion model tops USGS aftershock forecasts globally","feed_subtitle":"A single trained model draws fault-aligned aftershock fields and matches tuned ETAS on daily regional seismicity.","key_machinery":"A conditional denoising diffusion model with a U-Net denoiser. The forward process gradually corrupts the forecast field to noise; the reverse process learns to denoise it conditioned on stacked log-count and maximum-magnitude maps over nested time windows, the forecast horizon as a scalar embedding, and optional physical fields. Ensemble sampling via DDIM produces multiple plausible future fields whose mean and spread form the probabilistic forecast. This machinery lets the model learn the Omori-like time decay, sequence-dependent productivity, and anisotropic fault-aligned geometry from data rather than from prescribed functional forms.","core_discovery":"QuakeGen treats a future aftershock sequence as a two-channel spatial field - event count and maximum magnitude per grid cell - and learns its conditional distribution with a denoising diffusion model. Conditioned only on observed seismicity over nested past windows, the first 1.6 hours of early aftershocks, and the forecast horizon, one trained model generates an ensemble of future fields. On the held-out global test set it leads USGS-RJ on spatial information gain, Wasserstein-1 distance, and mean-squared-error skill score at every horizon scored, and on the regional daily benchmark it matches ETAS tuned separately for each region. The load-bearing claim is that the fixed isotropic spatial","pith_inferences":["Editorial extension: the author leaves implicit that the short-horizon comparison can be confounded - 53% of the 3-hour window and 13% of the 12-hour window are already inside the conditioning windows, so QuakeGen can partly copy the fields it was conditioned on while Reasenberg-Jones cannot. A clean test would score only events after the 1.6-hour conditioning cutoff.","Editorial extension: because the model is field-based rather than point-process-based, it should transfer to fluid-driven swarms, induced sequences, and slow-slip-driven activity, where pattern migration rather than a single mainshock controls the sequence; this is directly testable with injection-rate or geodetic conditioning fields.","Editorial extension: adding finite-fault geometry, mapped fault traces, or InSAR/GPS deformation as extra conditioning channels would likely sharpen the earliest hours after a mainshock, where aftershock observations are sparse and the current model is least constrained.","Editorial extension: making the forecast autoregressive - re-conditioning on newly arrived events after each horizon - would let the same model track a sequence continuously and align better with operational updating practice."],"forward_implications":["On 80 held-out global mainshocks of 2024-2025, QuakeGen beats the operational USGS Reasenberg-Jones forecast on spatial skill at every horizon, with larger gains at short horizons.","A single QuakeGen model trained once over all of southern California matches ETAS models tuned separately for the Salton Sea and San Jacinto regions at all daily forecast horizons.","The model reproduces Omori-like decay and Gutenberg-Richter magnitude statistics without being given these laws, and one trained network serves all forecast horizons from hours to a month.","The generative ensemble provides spread as a measure of forecast uncertainty, and the joint maximum-magnitude channel forecasts where the largest expected aftershocks will fall.","Applied without retraining to an enhanced template-matching catalog of southern California, the global model matches region-specific ETAS, suggesting it has learned sequence evolution rather than catalog-specific statistics."],"fun_headline_variants":["Diffusion model beats USGS aftershock forecasts on global test","QuakeGen tops USGS-RJ globally, matches ETAS regionally","Diffusion aftershock fields outperform USGS fixed-kernel forecasts","Fault-aligned aftershock forecasts from a single diffusion model","QuakeGen: diffusion forecasts beat USGS, tie ETAS on regional daily"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The short-horizon comparison is fair only if scoring QuakeGen on aftershocks that were already visible to it before the forecast starts does not inflate its apparent skill; 53% of the 3-hour window and 13% of the 12-hour window overlap its conditioning.","fun_headline_variants_meta":{"raw":{"variants":["Diffusion model beats USGS aftershock forecasts on global test","QuakeGen tops USGS-RJ globally, matches ETAS regionally","Diffusion aftershock fields outperform USGS fixed-kernel forecasts","Fault-aligned aftershock forecasts from a single diffusion model","QuakeGen: diffusion forecasts beat USGS, tie ETAS on regional daily"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000414,"raw_usage":{"total_tokens":2020,"prompt_tokens":834,"completion_tokens":1186,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":578,"completion_tokens_details":{"reasoning_tokens":1090}},"tokens_in":578,"tokens_out":1186,"duration_ms":10035,"temperature":1.0,"reasoning_tokens":1090,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T23:01:32.922255+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-score the global 3-hour and 12-hour forecasts using only events that occur after the conditioning cutoff of 1.6 hours, or retrain QuakeGen without the post-anchor conditioning windows. If its spatial advantage over USGS-RJ disappears, the claimed fault-aligned anisotropy gain is largely an artifact of the overlap rather than a learned forecasting skill.","supporting_citations":[],"review_version":1}