Pith. sign in

REVIEW 4 major objections 5 minor 20 references

Earthquake Detection Using Benford's Law

T0 review · 4 major / 5 minor · reviewed 2026-08-01 · deepseek-v4-flash

Pith's one-line read A simple digit-counting rule can detect local earthquakes and pinpoint P-wave onsets without training.

desk verdict The Benford-detection idea is appealing, but the central Goodness-of-Fit definition in Eq. 3 is mathematically inconsistent, so the reported 97.5% recall is not reproducible. read the letter →

arxiv 2607.27821 v1 pith:YHAVP74Y submitted 2026-07-30 physics.geo-ph physics.data-an

classification physics.geo-phphysics.data-an
keywords Benford'sLawearthquakedetectionP-waveonsetgoodnessoffitseismicwaveformsleadingdigitschi-squareSTA/LTAcomparison
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that the leading digits of seismic waveform amplitudes follow Benford's Law during the onset of earthquake energy but not during pre-event noise, and that this contrast can serve as a detection and onset-timing tool. Across a year-long intraplate array in the Deccan Volcanic Province and a Himalayan aftershock network, a sliding-window goodness-of-fit measure rises sharply at P-wave arrivals, detecting 268 of 275 cataloged events (≈97.5%) with no false positives. Crucially, the method requires no training, templates, or amplitude thresholds—only a window length—making it a candidate pre-filter for noisy or data-sparse settings such as planetary seismology. The paper positions the result as a statistical-transition detector rather than a sample-level phase picker.

What carries the argument

The carrying object is the sliding-window first-digit goodness-of-fit: each 1-second-stepped window tallies the leading digits of absolute amplitudes, compares them to Benford's distribution via a chi-square statistic, and converts the misfit into a Goodness-of-Fit measure φ(t). Two derived metrics—the noise-normalized deviation Δφ_N and the gradient-emphasizing Δφ′_N—convert raw φ into a detection trigger. Together they transform a digit-frequency histogram into an onset detector whose only tunable parameter is window length; the work this does is to separate 'noise-like' digit statistics from 'Benford-like' signal statistics.

What would settle it

Compute φ on 200 seconds of pre-event noise from the same stations using the paper's exact Eq. 3: if the 95th percentile φ_th approaches the 90% detection threshold, or if φ frequently exceeds 90% on noise, the detector's baseline is not separating noise from signal. Alternatively, compare BL-derived onsets against analyst-picked P arrivals for a set of high-SNR local events; if the median |Δt| exceeds one window length, the onset-alignment claim fails.

Watch

Extended reading notes

Core claim

The central claim is that local earthquake waveforms consistently conform to Benford's Law—the logarithmic distribution of first significant digits—during the onset of seismic energy, while pre-event noise does not. Using a sliding window to track the temporal evolution of a chi-square-based goodness-of-fit, the authors define normalized deviation metrics that compare signal windows against station-specific noise baselines. Applied to two contrasting tectonic datasets, the metrics show maximum BL conformity within about one window length of theoretical P-wave arrivals, and a combined detection criterion recovers 97.5% of cataloged events with no false positives. The authors conclude that BL

Load-bearing premise

The results assume that the Goodness-of-Fit formula φ=(1−χ)×100% really converts the chi-square misfit into a percentage near 100% for signal windows, and that the theoretical TauP/AK135 P-wave arrival times used as reference onsets are accurate to within a sliding-window length; if either fails, the 97.5% recall and onset-alignment claims are not established.

Editorial extensions

If this is right

  • A threshold-free, training-free detector that can be run on continuous records at trivial computational cost.
  • BL-based onset timing aligns with theoretical P arrivals within one window length, so it can seed phase association or ML pickers.
  • Because only window length matters, the method transfers across tectonic environments and noise conditions without retuning.
  • In planetary or remote deployments with no templates or labeled data, the method offers a first-pass event detection layer.
  • Compared with STA/LTA, it avoids false triggers that arise from amplitude-threshold tuning.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If BL conformity is genuinely a property of transient broadband seismic energy, the same sliding-window test might be adapted to detect tremor, volcanic signals, or debris flows—any source that changes the amplitude distribution's statistical character—without modifying the core metric.
  • The claim that noise does not conform to BL could be made into a sharper test: measure φ on noise-only records from different stations; a stable φ_th near 100% would undermine the separation.
  • The onset precision limited by window length suggests a two-pass design: BL to find candidate windows, then a short-window re-analysis or an ML picker inside those windows—implicit in the paper but not developed.
  • The 97.5% recall is measured against cataloged events; the method's practical value for unknown events depends on false positives over long continuous time, which the paper only evaluates in event-triggered windows.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes a Benford's Law (BL) based framework for local earthquake detection in continuous seismic waveforms. Using sliding windows, the authors compute a chi-square misfit between observed first-digit distributions and the BL prediction, convert it to a Goodness-of-Fit measure φ, and derive two normalized deviation metrics Δφ_N and Δφ'_N. They apply the method to two Indian datasets (Deccan Volcanic Province and NAMASTE/Himalaya), report that earthquake onsets show enhanced BL conformity relative to pre-event noise, claim detection of 268/275 cataloged events (97.5%) with no false positives, and argue that BL-based detection is parameter-light and training-free compared to STA/LTA and machine-learning pickers.

Significance. If the claims were supported, the paper would offer an appealing low-cost, training-free detector with potential use in noisy or data-sparse environments. The use of two contrasting tectonic datasets is a strength, and the idea that BL conformity changes at signal onset is worth investigating. However, the central methodological definition of φ is internally inconsistent, the performance evaluation is circular, and the 'no false positives' claim is not meaningful given the experimental design. These issues are load-bearing, so the contribution as presented cannot be accepted or reproduced from the stated methods.

major comments (4)
  1. [§2.2, Eq. (3)] Eq. (2) defines χ² as a sum of nonnegative terms with expected value ≈8 for 9 digits. Eq. (3) then defines φ = (1−χ)×100%, which for any χ>1 gives a negative percentage. The paper reports φ values near 90–100% throughout (e.g., Figures 2, 4, 5, 6), so the equation as written cannot produce those curves. Either χ in Eq. (3) denotes some unstated normalized quantity (e.g., reduced χ² or a p-value) or the figures were generated with a different formula. Because every downstream metric, threshold, and detection rate depends on φ, this internal inconsistency makes the central claim non-reproducible.
  2. [§3.3 and §3.4] The detection thresholds (φ>90%, Δφ_N>0.5, Δφ'_N>0.5) are described in §3.4 as 'empirically determined,' and the same DVP waveforms are then used to report the 97.5% recall and 'no false positive' results in §3.3 and §3.4. There is no train/validation split, independent test set, or out-of-sample check. The false-positive claim is especially problematic: §2.1 states that only waveform segments around cataloged events are analyzed, so the study never scans continuous noise-only intervals; 'no false positives within the analyzed time windows' does not establish a false-positive rate for continuous detection.
  3. [§3.3] The detection rate is reported as 268 out of 275 cataloged events, but Figure 3a shows 2403 DVP waveforms. It is not stated how the 275 events map to these waveforms (multi-station recordings per event), nor what criterion counts an event as detected (any one station? a minimum number of stations? a single window?). Without this aggregation rule, the 97.5% recall figure is ambiguous and cannot be independently verified.
  4. [§3.4, Eq. (7)] Eq. (7) defines Δφ'_N(t) = 1 − (φ_max − φ(t))/(100 − φ(t)). Algebraically this is (100 − φ_max)/(100 − φ(t)), which measures how close φ(t) is to its maximum in units of the distance to 100%; it is not a temporal gradient. The text repeatedly states that Δφ'_N 'emphasizes the gradient of φ' and 'highlight[s] the initial transition from noise-dominated to signal-dominated behavior,' but the formula contains no time derivative or finite-difference operator. This mismatch between the stated purpose and the actual definition is not merely cosmetic, since Δφ'_N is one of the three required detection criteria.
minor comments (5)
  1. [§4] Typo: 'parameter-lite' should be 'parameter-light' (also in the abstract the same phrase is correctly hyphenated).
  2. [§3.3/§3.4] A full paragraph beginning 'While Δφ_N effectively quantifies...' is repeated nearly verbatim in §3.3 and §3.4. Please remove the duplicate.
  3. [§2.2, Eq. (4)] Δφ_N(t) is defined using φ_max, the maximum of the entire trace, so at times before the event the metric uses future information. If the intended use is real-time detection, the authors should state whether φ_max is replaced by a causal running maximum in practice; as written, the definition is not causal.
  4. [Figure 3c] The color scale is labeled φmax(%) but the axis label and caption do not explain how the color value relates to the plotted points; please clarify.
  5. [§3.1] The statement that 'φ values remain consistently low during the pre-event noise window' is only supported by example figures; no aggregate statistics for pre-event φ are provided. A histogram or median curve over all waveforms would strengthen the claim.

Circularity Check

3 steps flagged · score 6.0 of 10

In-sample threshold fitting and a test set containing no event-free windows make the 97.5% recall and zero-false-positive claims partially circular; the raw Benford-vs-noise comparison is not.

  1. fitted input called prediction [Section 3.1 and Section 3.3-3.4]
    "Across the DVP dataset, the majority of waveforms (90.7%) yield(∆ϕ N )max >0.5, (Figure 3a). Only a small fraction of cases fall below this value, primarily corresponding to low SNR observations. So, we can use(∆ϕN )max = 0.5as a benchmark for our further analysis."

    The 0.5 benchmark is selected after observing that 90.7% of the same DVP waveforms exceed it, and the later 'conformity' and detection results use this same threshold on the same dataset. Choosing a cutoff so that most waveforms pass, then reporting the pass rate as evidence that earthquake signals conform to Benford's Law, is in-sample self-confirmation rather than an independent test.

  2. fitted input called prediction [Section 3.3-3.4]
    "These thresholds are empirically determined on the basis of stability across a wide range of events and noise conditions and are chosen to ensure that detections correspond to statistically significant and temporally localized deviations from background noise. ... Out of a total of 275 cataloged events within the study period and distance range, it successfully detected 268 events, corresponding to a detection rate of approximately 97.5%."

    The combined thresholds are fit to the same DVP waveforms on which the 97.5% recall and onset residuals are then reported. No independent validation set or cross-validation is described, so the headline numbers are in-sample performance after parameter selection rather than predictions. This does not undermine the raw comparison of digit distributions to the fixed Benford law, but it makes the detector-performance claims fitted.

1 more flagged steps
  1. self definitional [Section 2.1 and Section 3.3]
    "In each case, we analyze the continuous waveform data around known events ... Notably, no false positive detections were observed within the analyzed time windows."

    Every analyzed time segment is selected around an ISC-cataloged event, so the test set contains no event-free windows. A trigger anywhere in such a segment can be attributed to the known event, making a false positive impossible by construction. The claim of zero false positives is therefore an artifact of the test-set definition, not a measured false-trigger rate.

full rationale

The central Benford-vs-noise comparison is not circular: first-digit counts in each sliding window are compared with the fixed Benford distribution of Eq. 1 through a chi-square misfit, and the contrast between signal and pre-event windows is an empirical observation. The circularity is concentrated in the thresholding and evaluation layer. Section 3.1 selects (Δφ_N)_max=0.5 after seeing that 90.7% of the same DVP waveforms exceed it; Section 3.4 declares φ>90%, Δφ_N>0.5, Δφ'_N>0.5 to be 'empirically determined'; Section 3.3 then reports 97.5% recall and zero false positives on the same event set without an out-of-sample or cross-validated test. The zero-false-positive claim is definitional because only waveforms around cataloged events are analyzed, so no event-free windows exist in which a false trigger could be counted. The Discussion's stated limitations (window-length trade-off, weak emergent signals) are honest and do not add circularity. Separately, Eq. 3 as written (φ=(1−χ)×100% with χ from Eq. 2) is arithmetically inconsistent with the reported φ≈90–100% values, but this is a reproducibility/correctness defect rather than a circularity and is not counted in the score.

Assumptions & free parameters 4 free parameters · 6 assumptions · 0 invented entities

No new physical entities are introduced. Δφ_N and Δφ'_N are constructed statistics, not independently testable entities. The load-bearing choices are the window length, the detection thresholds, the noise-baseline definition, and an unstated normalization in the Goodness-of-Fit formula.

free parameters (4)
  • Sliding window length L = ~1/10 of trace duration; 40–130 s in sensitivity tests
    Window length controls the trade-off between detection stability and onset-timing precision (§2.2 step 1, §3.4, Fig. 7). The 'optimal' length is selected per epicentral distance without a defined optimization rule.
  • Detection thresholds: φ>90%, Δφ_N>0.5, Δφ'_N>0.5 = 90%, 0.5, 0.5
    Described as 'empirically determined' (§3.4) on the same DVP waveforms used to report 97.5% recall. Changing these thresholds changes the detection rate.
  • φ_th baseline percentile and duration = 95th percentile of φ over first 200 s of pre-event noise
    Defines the station- and event-specific noise baseline used in Δφ_N (§2.2 step 4). The choice of 95th percentile and 200 s window is ad hoc.
  • Unstated normalization of χ in Eq. 3
    Eq. 3 as written cannot map a standard chi-square statistic to the reported 80–95% φ values; an unstated normalization or alternative definition of χ is required to reproduce the results.
assumptions (6)
  • domain assumption Benford's Law applies to finite, filtered, unit-corrected seismic amplitude samples
    The method assumes first-digit counts in short sliding windows are governed by BL for signals and not for noise; no scale-invariance or multiplicative-process justification is given for these specific processed waveforms (§2.2).
  • domain assumption Pre-event noise does not conform to BL while earthquake signal does
    This empirical hypothesis is the central discriminator; it is asserted from the data rather than independently motivated (§3.1).
  • domain assumption AK135/TauP theoretical travel times are accurate enough as ground truth for P-wave onset
    Manual picking was explicitly rejected as infeasible (§3.2); onset-alignment claims rely on model predictions that may be inaccurate for small local events.
  • domain assumption Catalog event locations and origin times (ISC, Mendoza et al.) are correct
    Event windows and P-wave predictions inherit catalog accuracy; mislocation would shift the comparison baseline (§2.1).
  • standard math Chi-square expected counts E_d = N·P(d) with Benford probabilities
    Standard chi-square statistic with expected BL counts, but it is not integrated correctly with Eq. 3.
  • ad hoc to paper One-tenth trace duration windowing yields stable results
    The relative windowing strategy is an empirical choice made for this study, not derived from theory or prior literature (§2.2 step 1).

how reviews work

0 comments
Cite this review

Pith. "Pith review of Earthquake Detection Using Benford's Law." pith.science (2026). https://pith.science/paper/YHAVP74Y

@misc{pith2026260727821,
  author       = {Pith},
  title        = {Pith review of: Earthquake Detection Using Benford's Law},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YHAVP74Y}},
  note         = {Machine review of arXiv:2607.27821}
}
read the original abstract

Reliable detection of local earthquakes and accurate identification of P-wave onsets are fundamental tasks in seismology, yet many existing methods to accomplish them require extensive parameter tuning or large training datasets. In this study, we investigate the applicability of Benford's Law - a logarithmic distribution governing the occurrence of leading digits in naturally occurring data - as a statistical framework for local earthquake detection. We analyze continuous seismic waveform data from two contrasting tectonic environments in the Indian subcontinent: the intraplate Deccan Volcanic Province and the actively deforming Himalayan region. Using a sliding-window approach, we quantify the temporal conformity of first-digit distributions of seismic amplitudes to Benford's Law and apply an adaptive normalization scheme to account for station-specific noise. Our results show that local earthquake waveforms consistently conform to Benford's Law during the onset of seismic energy, while pre-event noise does not. Benford-derived statistical anomalies align closely with theoretical P-wave arrival times, with detection performance primarily controlled by window length. These results establish Benford's Law as a computationally inexpensive, parameter-light, and training-free tool for local earthquake detection.

Figures

Figures reproduced from arXiv: 2607.27821 by the authors.

Figure 1
Figure 1. Station and event distribution for the datasets used in this study. (a) Map of the DVP dataset showing [PITH_FULL_IMAGE:figures/full_fig_p004_1.png] view at source ↗
Figure 2
Figure 2. (a-b) Temporal variation of the Goodness of Fit ( [PITH_FULL_IMAGE:figures/full_fig_p006_2.png] view at source ↗
Figure 3
Figure 3. (a-b) Pie chart illustrating the distribution of [PITH_FULL_IMAGE:figures/full_fig_p007_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Comparison of BL–based detection metrics and STA/LTA performance for a representative noisy waveform. [PITH_FULL_IMAGE:figures/full_fig_p011_4.png]
Figure 5
Figure 5. Figure 5: Performance of the normalized Benford’s Law deviation parameter [PITH_FULL_IMAGE:figures/full_fig_p012_5.png]
Figure 6
Figure 6. Figure 6: Sensitivity of Bl–based P onset detection to moving window length. Panels [a–i] show BL parameters ( [PITH_FULL_IMAGE:figures/full_fig_p013_6.png]
Figure 7
Figure 7. Figure 7: Variation of optimal sliding window length for BL–based P onset detection as a function of epicentral [PITH_FULL_IMAGE:figures/full_fig_p014_7.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

20 extracted references · 3 canonical work pages

  1. [1]

    The American Mathematical Monthly , volume=

    The first digit problem , author=. The American Mathematical Monthly , volume=. 1976 , publisher=

  2. [2]

    Proceedings of the American Mathematical Society , volume=

    Base-invariance implies Benford’s law , author=. Proceedings of the American Mathematical Society , volume=

  3. [3]

    Statistical science , pages=

    A statistical derivation of the significant-digit law , author=. Statistical science , pages=. 1995 , publisher=

  4. [4]

    American Journal of mathematics , volume=

    Note on the frequency of use of the different digits in natural numbers , author=. American Journal of mathematics , volume=. 1881 , publisher=

  5. [5]

    Computers & Geosciences , volume=

    SAIPy: A Python package for single-station earthquake monitoring using deep learning , author=. Computers & Geosciences , volume=. 2024 , publisher=

  6. [6]

    Seismica , volume=

    Picking regional seismic phase arrival times with deep learning , author=. Seismica , volume=

  7. [7]

    Seismological Research Letters , volume=

    Karplus, Marianne S and Pant, Mohan and Sapkota, Soma Nath and N. Seismological Research Letters , volume=. 2020 , publisher=

  8. [8]

    Journal of Geophysical Research: Solid Earth , volume =

    Zhou, Yijian and Ding, Hongyang and Ghosh, Abhijit and Ge, Zengxi , title =. Journal of Geophysical Research: Solid Earth , volume =. doi:https://doi.org/10.1029/2025JB031294 , url =

Show all 20 references
  1. [9]

    Seismological Research Letters , volume =

    Zhou, Yijian and Yue, Han and Fang, Lihua and Zhou, Shiyong and Zhao, Li and Ghosh, Abhijit , title =. Seismological Research Letters , volume =. 2021 , month =. doi:10.1785/0220210111 , url =

  2. [10]

    , title =

    Allen, Rex V. , title =. Bulletin of the Seismological Society of America , volume =. 1978 , month =. doi:10.1785/BSSA0680051521 , url =

  3. [11]

    2024 , publisher=

    Zhou, Qi and Tang, Hui and Turowski, Jens M and Braun, Jean and Dietze, Michael and Walter, Fabian and Yang, Ci-Jian and Lagarde, Sophie , journal=. 2024 , publisher=

  4. [12]

    Geology , volume=

    Geyer, Adelina and Mart. Geology , volume=. 2012 , publisher=

  5. [13]

    1995 , publisher=

    Kennett, Brian LN and Engdahl, ER and Buland, Raymond , journal=. 1995 , publisher=

  6. [14]

    Seismological Research Letters , volume=

    The TauP Toolkit: Flexible seismic travel-time and ray-path utilities , author=. Seismological Research Letters , volume=. 1999 , publisher=

  7. [15]

    Mathematical Geosciences , volume=

    Benford’s Law in time series analysis of seismic clusters , author=. Mathematical Geosciences , volume=. 2012 , publisher=

  8. [16]

    2019 , publisher=

    Mendoza, MM and Ghosh, A and Karplus, MS and Klemperer, SL and Sapkota, SN and Adhikari, LB and Velasco, A , journal=. 2019 , publisher=

  9. [17]

    Geophysical research letters , volume=

    Sambridge, Malcolm and Tkal. Geophysical research letters , volume=. 2010 , publisher=

  10. [18]

    On the ability of the

    D. On the ability of the. Seismological Research Letters , volume=. 2015 , publisher=

  11. [19]

    2023 , publisher=

    Saha, Gokul and Kumar, Vivek and Chaubey, Dipak K and Rai, Shyam S , journal=. 2023 , publisher=

  12. [20]

    Proceedings of the American Philosophical Society , pages=

    The law of anomalous numbers , author=. Proceedings of the American Philosophical Society , pages=. 1938 , publisher=

Pith tools

Reviewed August 1, 2026 · model on record in the stance chip above.