{"id":"4cef9bfa-147c-491f-b244-f347802bac03","arxiv_id":"2509.13600","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A two-dimensional C/N0 and calibrated received-power metric from a low-cost u-blox receiver distinguishes nominal, jammed, spoofed, and blocked GNSS conditions, validated in three regions.","lead":"This paper shows that a low-cost off-the-shelf GPS receiver can detect jamming and spoofing by combining signal-to-noise values with a calibrated power estimate. The method is tested at a controlled jammer exercise in Norway and in real interference zones in Poland and the Mediterranean Sea.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Fixed detection thresholds may not transfer across sites: only the mean of the nominal distribution is recalibrated, while the covariance is assumed site-invariant.","rationale":"The reader's weakest assumption identifies exactly the load-bearing concern: the nominal distribution's covariance is assumed transferable with only a mean recalibration. I agree. This is not an internal inconsistency but an unvalidated distribution-shift assumption. The controlled Norway test is real supporting evidence for detection under jamming, and the real-world deployments are suggestive, but the paper's own Section 3.3/4 wording commits to fixed thresholds after centering, making the covariance assumption load-bearing for the 1e-6 false-positive claim. The concrete test uses data the authors already have and would settle the issue: if local nominal covariances match Stanford's, the concern is resolved and the method is stronger; if not, the paper should either re-fit the threshold per site with local nominal spread or quantify the achieved false-positive rate. Reader verdict remains CONDITIONAL, so no verdict change is needed, but the paper should include this cross-site covariance check or explicitly limit the fixed-threshold claim.","tokens_in":12515,"tokens_out":3771,"duration_ms":45524,"concrete_test":"Using the already-collected datasets, identify nominal epochs at each site (e.g., Norway's 2-hour no-RFI window, quiet periods in Northern Poland, and stationary/docked periods in the Mediterranean). Apply only the Section 4 mean-shift calibration, then count the fraction of local nominal samples that fall outside the Stanford-optimized detection region. If empirical false-alarm rate at any site is orders of magnitude above 1e-6 (say >1e-3), threshold transferability fails. Additionally, estimate the 2x2 covariance of each site's nominal cluster and compare eigenvalues/principal axis to the Stanford covariance used in the threshold optimization (e.g., bootstrap or F-test). This separates 'shape changed' from 'mean shifted' and settles whether the fixed-threshold claim holds.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central viability claim depends on the nominal C/N0-over-received-power distribution being portable after a mean shift. Section 3.1 fits a multivariate Gaussian to Stanford data and optimizes the disturbance threshold to a 1-in-1,000,000 false-positive rate; Section 3.3 states that once the nominal region is correctly centered, 'no changes are needed to the RFI detection thresholds'; Section 4 applies only a mean re-centering based on local nominal data. If the covariance (spread/orientation) of the nominal distribution changes with site, antenna, multipath, or local RF environment, the optimized threshold no longer provides the claimed false-positive rate. The Norway no-RFI period in Table 2 (0 false positives) is one compatible site, but it does not establish covariance transferability across Poland, the Mediterranean deployment, or other installations. The paper contains no numerical comparison of nominal covariance across sites, and no released threshold parameters to re-derive the analysis.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a low-cost GNSS RFI monitoring methodology based on a two-dimensional metric formed from carrier-to-noise density (C/N0) and a calibrated received-power estimate derived from u-blox F9P SPAN data. The authors describe a calibration chain (AGC correction, temperature compensation, PSD-weighted bin aggregation, and absolute power conversion), define detection and classification regions in the C/N0-versus-received-power plane, and set a disturbance threshold using a Gaussian fit and importance-sampling falsification to target a 1-in-1,000,000 false-positive rate. The method is validated with a controlled jamming/spoofing exercise in Norway, which provides ground truth, and applied to uncontrolled data from northern Poland and the southeastern Mediterranean. The central claim is that, with proper calibration, COTS-based monitoring can reliably detect and distinguish nominal, jammed, spoofed, and blocked signal conditions.","tokens_in":12774,"tokens_out":3838,"duration_ms":47052,"significance":"If the result holds, the paper offers a meaningful contribution to scalable RFI monitoring: a dense network of sub-$400 receivers could fill a gap between expensive airport-grade systems and satellite/ADS-B-based approaches. The controlled Norway validation is a genuine strength: the detection thresholds were developed from Stanford data alone, and the Norway evaluation is therefore independent. The overall jamming detection accuracy of 99.4% (Table 2) and the near-perfect RFI detection in the spoofing test (Table 3) are encouraging. The paper is also candid about known weaknesses, including low spoofing-characterization sensitivity (42.1%) and the heuristic jamming/spoofing boundary. However, the central portability claim—that threshold transfer across sites requires only a mean shift—rests on an unexamined assumption of covariance invariance, and the calibration curves are presented without uncertainty quantification. These issues do not invalidate the core concept but need addressing before the method can be claimed as a turnkey monitoring solution.","major_comments":[{"comment":"The threshold-transferability claim is load-bearing and is not supported by the presented evidence. Section 3.3 states that 'once the nominal region is correctly centered for a given setup, no changes are needed to the RFI detection thresholds.' Section 4 then applies only a mean re-centering based on local nominal data. The threshold in Section 3.1 was optimized using the Stanford nominal distribution's full shape (mean and covariance) to achieve a 1-in-1,000,000 false-positive rate. If the spread or orientation of the nominal distribution changes with site, antenna, multipath environment, or local RF conditions, the optimized threshold will not deliver the claimed false-positive rate. The Norway no-RFI result (Table 2: 0 false positives out of 7000 samples) is far too small to validate a 1e-6 false-positive probability, and no nominal-covariance comparison across Stanford, Norway, Pola","section":"3.3/4"},{"comment":"The false-positive rate estimation relies on importance sampling with a proposal distribution and 'fuzzing', but the manuscript does not specify the proposal distribution, the number of rollouts, or the convergence criteria. The nominal distribution is fit to a multivariate Gaussian and discretized to a 1-dB/1-dB-Hz grid, yet no goodness-of-fit test or sensitivity analysis is presented. Given that the 1-in-1,000,000 false-positive rate is a headline claim for the threshold design, the statistical procedure needs more detail: at minimum, report the proposal distribution, sample size, and confidence bounds on the estimated false-positive rate, and justify the Gaussian assumption against the observed multipath-influenced data shown in Figure 1.","section":"3.1"},{"comment":"The calibration chain (AGC offset multiplier 3.7x SPAN PGA in Section 2.2.1, temperature curve in Section 2.2.2, SPAN-to-dBW/Hz conversion in Section 2.2.4) is essential to the classification regions, but the fitted parameters are presented without error bars, validation on multiple receiver units, or analysis of how calibration uncertainty propagates into the detection/classification boundaries. For a paper claiming a framework adaptable to other COTS receivers, the receiver-specific nature of these constants and the resulting uncertainty in the absolute received-power measurement must be quantified. Otherwise, a systematically biased power estimate could shift points across the jamming/spoofing or nominal/jamming boundaries without the user knowing.","section":"2.2"},{"comment":"The spoofing-characterization sensitivity of 42.1% means that the method fails to characterize the majority of spoofing events, even though it detects RFI nearly perfectly. The abstract and conclusion claim the method can 'differentiate' spoofed signal conditions, which overstates the results. The authors acknowledge the limitation in the text, but the central claim of a monitoring methodology for 'jamming and spoofing' should be balanced by a clear statement that spoofing characterization is currently unreliable for weak or non-capturing spoofing signals. The proposed pseudorange-consistency checks are listed as future work; until then, the paper should either downgrade the classification claim or present the work as primarily a jamming/blockage detector with an ancillary spoofing indicator.","section":"Table 3/Conclusion"}],"minor_comments":[{"comment":"The abstract says 'southeast soars' should be 'southeast shores' (also appears in the Introduction).","section":"Abstract/Intro"},{"comment":"Equation labeling is inconsistent: the false-positive estimation equation is labeled '(1)' but it is the third numbered equation in the paper (after Eq. (1) in Section 2.2.3 and Eq. (2) in Section 2.2.3). Renumber accordingly.","section":"Section 3.1"},{"comment":"The indicator function in Eq. (3) uses a nonstandard symbol '⊮'; use the standard indicator notation 1{...} or define the notation.","section":"Section 3.1"},{"comment":"Caption typo: 'GPA L1 C/A' should be 'GPS L1 C/A'.","section":"Figure 5"},{"comment":"Figure 9 is described as a 24-hour observation, but Figure 10 shows data sampled from ten months. It would be helpful to explicitly state that Figure 9 is the same-day elevation-filtered subset from which the move to SBAS was motivated.","section":"Section 2.3"},{"comment":"The description of Table 2's 'Full Day' row could mention that the false positives (54) are dominated by the step-RFI recovery period; currently the text explains this only qualitatively.","section":"Section 4.1"},{"comment":"References to 'Kochenderfer et al., 2025' and 'Kochenderfer & Wheeler, 2019' are appropriate, but the page numbers or chapter sections would help readers locate the importance-sampling formulation.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope for a GNSS/signal-processing venue and the controlled Norway test is a valuable contribution. The main reservation is that the portability claim of the detection thresholds is not backed by data; this is fixable but requires either a comparative covariance analysis or a revised methodology. I would also encourage the authors to release the calibration curves and threshold parameters to enhance reproducibility. The spoofing-characterization sensitivity is a genuine limitation but not grounds for rejection provided the claims are appropriately tempered."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a genuinely useful engineering paper. The new bit is replacing the AGC-based power metric from Lo et al. 2021 with a calibrated received-power metric derived from u-blox SPAN FFT data, then combining that with C/N0 into a 2-D detection space. The full calibration chain—AGC offset, temperature correction, PSD weighting, and the empirical conversion to dBW/Hz—is real work, and it pays off. The controlled Norway jamming test gives an honest independent validation: 99.4% overall accuracy, with the low sensitivity on the ramp test correctly attributed to weak signals and receiver response delays. Applying the same thresholds to Poland and the Mediterranean without re-fitting to those sites is a strong move, and the fact that the paper states the thresholds were developed only from Stanford data makes that evaluation genuinely out-of-sample.\n\nThe soft spots are real, but fixable. The spoofing characterization sensitivity is 42.1%, and the paper still claims the method can \"identify and distinguish nominal, jammed, spoofed, and blocked signal conditions\" as if the spoofing result were comparable to the jamming result. The spoofing boundary is a heuristic line, and the paper admits targeted attacks near the boundary will be misclassified. That claim needs to be toned down.\n\nThe bigger concern is the transferability of the nominal distribution's shape. The threshold is optimized to a 1-in-a-million false-positive rate using the Stanford covariance. Section 3.3 says that once the nominal region is re-centered, \"no changes are needed to the RFI detection thresholds,\" and Section 4 only shifts the mean. If the covariance changes with site, antenna, multipath, or local RF environment, that false-positive rate will not hold. Norway's zero false positives during the no-RFI period is one compatible site, not a demonstration across deployments. The paper gives no numerical comparison of nominal covariance across sites and does not release the calibration parameters or data, so this cannot be checked. That is the load-bearing assumption, and the authors need to address it, not just assert it.\n\nAlso minor: the calibration curves in Figures 3 and 7 lack error bars, and the AGC offset (3.7x) and the temperature curve are fitted to Stanford hardware without clear uncertainty.\n\nBottom line: the jamming detection part is well supported and the calibration methodology is a legitimate contribution. The spoofing claim and the cross-site covariance assumption need revision, and code/data release is needed for full reproducibility. This is exactly the kind of paper that should go to peer review—an editor should not desk reject it. I would read a revised version and would cite the jamming detection approach.","headline":"A solid, practical GNSS RFI monitoring paper whose jamming detection is well validated; the spoofing differentiation and cross-site threshold transferability claims are the weak spots, but it deserves serious peer review.","tokens_in":13268,"tokens_out":1495,"would_cite":true,"duration_ms":19579,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Low-cost GNSS receivers can detect and classify jamming, spoofing, and blockage by plotting carrier-to-noise ratio against a calibrated received-power estimate.","keywords":["GNSS","RFI monitoring","jamming detection","spoofing detection","C/N0","received power","COTS receivers","SBAS"],"falsifier":"Collect a week of nominal (known interference-free) data from a receiver deployed at a site with substantially different multipath or antenna characteristics, recenter the nominal region to the local mean, and count the fraction of samples falling outside the paper's detection threshold. If that fraction significantly exceeds 10^-6, the transferability assumption is falsified.","tokens_in":12421,"feed_emoji":"📡","tokens_out":3689,"duration_ms":38520,"temperature":0.7,"pith_summary":"The paper claims that a low-cost commercial GNSS receiver, once its internal power measurement is calibrated, can reliably detect and classify radio-frequency interference: jamming, spoofing, and simple signal blockage. The proposed method plots each satellite's carrier-to-noise ratio (C/N0) against a calibrated received-power estimate, forming a two-dimensional space where nominal, jammed, blocked, and spoofed conditions fall in distinct regions. Using geostationary SBAS satellites keeps the nominal distribution tight, and thresholds are tuned to a one-in-a-million false-positive rate. The method is validated with controlled jamming tests and two real-world deployments, reporting detection accuracy above 99% in the controlled setting.","feed_headline":"Spot GNSS jamming and spoofing with a low-cost receiver","feed_subtitle":"A calibrated C/N0-vs-power plot lets cheap monitors tell nominal, jammed, and spoofed signals apart.","key_machinery":"The load-bearing object is the two-dimensional detection region map in C/N0-versus-received-power space, anchored by a nominal distribution built from geostationary SBAS satellites. The received-power metric is produced by a four-step calibration chain: automatic gain control adjustment, temperature compensation to a 300 K reference, weighting of 500 kHz FFT bins by the L1 C/A power spectral density, and a lab-calibrated translation to absolute dBW/Hz. Detection thresholds are set by optimizing a boundary around the nominal distribution to meet a target false-positive rate, using importance-sampling/falsification to estimate rare-event probabilities. The \"jamming path\" — a roughly 1:1 slope","core_discovery":"The central claim is that the two-dimensional metric C/N0 over received power separates four signal states using only observables a low-cost receiver already outputs. The paper shows that under jamming, C/N0 falls about 1 dB for every 1 dB of added noise, tracing a predictable \"jamming path\"; under blockage, C/N0 drops while noise power stays nominal; under spoofing, the receiver reports C/N0 values too high for the measured power, occupying an implausible region. The received-power axis is constructed from the receiver's internal FFT (SPAN) by applying AGC correction, temperature calibration, PSD-weighted integration over the GPS L1 C/A band, and a laboratory-derived conversion to dBW/Hz. T","pith_inferences":["A testable extension is to re-fit the full covariance of the nominal distribution at each new site rather than shifting only its center; comparing false-positive rates would tell whether the threshold shape is truly universal.","The paper's linear \"jamming path\" suggests the method could estimate interference power from the displacement of a measured point, which the authors do not fully exploit.","For moving platforms (e.g., the ship deployment), motion-induced C/N0 variation blurs the nominal region; a motion-compensated version might be needed for mobile monitoring."],"forward_implications":["If the method holds up, a dense network of sub-$400 receivers can monitor airports and other critical infrastructure for GNSS interference without expensive multi-beam antenna systems.","Thresholds can be tuned to any desired false-positive rate, letting operators trade sensitivity against nuisance alarms.","Spoofing can be flagged without pseudorange-level authentication whenever measured power and C/N0 are inconsistent; targeted weak spoofing remains a limitation.","The same two-dimensional space can flag installation problems (the \"unrealistic\" region), helping maintain data quality in a monitoring network."],"fun_headline_variants":["Low-cost receiver separates jamming, spoofing, and blockage","Cheap GNSS receiver detects spoofing via C/N0 and power","C/N0 versus power: cheap way to spot GNSS attacks","Poor-man's GNSS monitor tells jamming from spoofing","Calibrated C/N0 and power plot flags spoofing cheaply"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The spread of the nominal C/N0-versus-power distribution measured at the development site is assumed to transfer to new locations after only shifting its center, so if local multipath, antenna differences, or receiver variability change that spread, the claimed one-in-a-million false-positive rate will not hold.","fun_headline_variants_meta":{"raw":{"variants":["Low-cost receiver separates jamming, spoofing, and blockage","Cheap GNSS receiver detects spoofing via C/N0 and power","C/N0 versus power: cheap way to spot GNSS attacks","Poor-man's GNSS monitor tells jamming from spoofing","Calibrated C/N0 and power plot flags spoofing cheaply"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000188,"raw_usage":{"total_tokens":1138,"prompt_tokens":684,"completion_tokens":454,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":428,"completion_tokens_details":{"reasoning_tokens":359}},"tokens_in":428,"tokens_out":454,"duration_ms":5041,"temperature":1.0,"reasoning_tokens":359,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T16:26:42.534738+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Collect a week of nominal (known interference-free) data from a receiver deployed at a site with substantially different multipath or antenna characteristics, recenter the nominal region to the local mean, and count the fraction of samples falling outside the paper's detection threshold. If that fraction significantly exceeds 10^-6, the transferability assumption is falsified.","supporting_citations":[],"review_version":1}