{"id":"74d0d77d-b259-4380-8ddb-2dc047e71f8c","arxiv_id":"2507.06688","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A supervised encoder-decoder network trained on simulated GRAND signals and realistic Gaussian noise recovers air-shower radio pulses from real detector noise with >95% efficiency at SNR≈4 and low false positives.","lead":"GRAND's radio detector team trained a neural network to strip noise from air-shower radio pulses. In simulations, the network recovers pulses at signal-to-noise ratios around 4 with over 95 percent efficiency, suggesting it could lower GRAND's detection threshold.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"No sensitivity baseline: the 95% denoising efficiency may not exceed raw thresholding, so the claimed sensitivity enhancement is unestablished.","rationale":"The reader's weakest assumption is the fidelity of ZHAireS/GRANDLib simulated signals to real GRANDProto300 pulses, which is an external-validity concern about deployment. I agree that this is a real limitation, and the paper's own statement that it has 'successfully' bridged the reality gap is only partially supported because only the noise, not the signal, is real. However, I see a more immediate, internally checkable gap: the paper's central quantitative claim of sensitivity enhancement lacks a baseline comparison. The denoising efficiency is defined solely on the denoised trace, and no raw-trace or conventional-filter efficiency is reported. Without such a baseline, the 95% number is uninterpretable as an enhancement, regardless of simulation fidelity. This concern does not require new field data to test; it can be settled by re-analyzing the existing test set. Because the paper is otherwise coherent and the architecture/training are sound, the appropriate outcome remains conditional acceptance with the baseline comparison and confidence intervals required before the sensitivity claim is used in production.","tokens_in":8481,"tokens_out":3994,"duration_ms":49725,"concrete_test":"On the same held-out test set used for Figures 3 and 4, compute the fraction of traces whose raw noisy peak amplitude exceeds 15 ADC in each SNR bin, alongside the same fraction for a simple spectral Wiener filter and for the denoiser. Also compute binomial confidence intervals for each curve. If the raw or Wiener-filter efficiency is already at or above the denoised efficiency for SNR > 4, the sensitivity-enhancement claim is not supported; if the denoiser clearly exceeds both baselines, the claim is strengthened.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim in the abstract and conclusion is that the denoiser produces a 'sensitivity enhancement' and achieves 95% denoising for SNR > 4. However, the reported denoising efficiency (Section 5, Conditions 1 and 2) is computed only on the denoised output: a trace is counted as successfully denoised if the clean-signal peak exceeds 15 ADC and the denoised peak exceeds 15 ADC. The paper never reports the analogous efficiency for the raw noisy traces, nor for a conventional filter such as a Wiener or matched filter. This matters because at SNR > 4, a raw threshold trigger may already achieve high efficiency: with 768 time bins, the maximum of pure Gaussian noise can reach several sigma, and the authors state that 15 ADC corresponds roughly to SNR = 1. If the raw trace itself already meets the 15-ADC peak condition for most SNR > 4 traces, then the 95% denoising efficiency does not demonstrate sensitivity enhancement; it merely shows the denoiser does not destroy signals that a simple threshold would have found. The amplitude-ratio and peak-timing comparisons in Figures 3-5 do show genuine improvements over the noisy traces on those secondary metrics, but the headline efficiency claim is not benchmarked against any baseline. Thus, even under the paper's own simulation assumptions, the quantitative basis for 'sensitivity enhancement' is missing.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a supervised encoder-decoder convolutional network for denoising radio pulses from extensive air showers in the GRANDProto300 experiment. Signals are simulated with ZHAireS and processed through GRANDLib; noise is generated as Gaussian noise matching the power spectra of measured ADC noise traces, while real noise traces are used for testing. The model is trained with an L1 loss on 768-bin windows. The authors report a denoising efficiency above 95% for signals with SNR above about 4, improved peak-timing accuracy relative to noisy traces, amplitude recovery close to unity, and false-positive fractions of 1.4% (S-N) and 1.1% (E-W) at a 15-ADC threshold on pure-noise traces.","tokens_in":8653,"tokens_out":4937,"duration_ms":55755,"significance":"If the central claim of sensitivity enhancement were established, this work would be a useful contribution to the GRAND software pipeline and to radio-detection experiments more generally. The paper has several genuine strengths: it uses a realistic noise model derived from measured ADC noise spectra, it randomizes the pulse position to avoid the network learning a fixed time bin, it evaluates on a held-out test set with real noise traces, and it demonstrates large improvements in peak-timing accuracy and amplitude ratio over the noisy traces. However, the headline quantitative claim of a sensitivity enhancement is not yet supported because the denoising efficiency is not compared with the raw traces or a standard filter baseline; additionally, the test set uses simulated signals, so the real-world deployment claim requires qualification. The approach is promising and the gaps are addressable within the scope of a revision.","major_comments":[{"comment":"The denoising efficiency is defined through Conditions 1 and 2, which compare only the clean-trace peak and the denoised-trace peak against the ADC threshold. No analogous efficiency is reported for the raw noisy traces, nor for a conventional filter such as a Wiener or matched filter. This is load-bearing for the claimed 'sensitivity enhancement': with 768 time bins, the maximum of pure Gaussian noise can reach several sigma, and the paper itself states that 15 ADC corresponds roughly to SNR = 1. Many traces with SNR > 4 may already satisfy the 15-ADC condition on the raw trace, in which case the 95% denoising efficiency does not demonstrate additional sensitivity. Please add the raw-trace efficiency curve, a simple amplitude-threshold baseline, and ideally a standard filter baseline, and recast the sensitivity claim based on the comparison.","section":"Section 5, Figures 3-4"},{"comment":"The test set uses simulated ZHAireS signals added to real noise traces, so the evaluation validates denoising of simulated pulses under realistic noise, not of real air-shower pulses. The conclusion states that the model was 'successfully applied to real noise traces from the experiment' and 'paves the way for its integration into ongoing and future experiments,' which overstates the evidence. The reality gap for the signal component—pulse shape, polarization, and RF-chain response—remains untested. Please state this limitation explicitly in the conclusion and temper the deployment claim accordingly.","section":"Section 3 and Section 6"},{"comment":"The text is ambiguous about whether the false-positive test uses real AN noise traces or simulated Gaussian noise generated from AN power spectra. The earlier description says real noise traces are used for the testing set, but the false-positive paragraph says 'pure Gaussian AN noise.' This distinction is important because the model was trained on Gaussian noise, so evaluation on simulated Gaussian noise could underestimate the false-positive rate on real, non-Gaussian noise. Please clarify which dataset was used and, if real noise traces are available, report the false-positive fraction on them.","section":"Section 5, false positive fraction"},{"comment":"The efficiency, timing, amplitude-ratio, and false-positive results are quoted without statistical uncertainties. The test and validation sets contain 21,145 traces, so binomial confidence intervals on the 95% efficiency and the 1.4% false-positive fraction would be simple to compute and would strengthen the quantitative claims. Please add error bars or report the relevant counts.","section":"Section 5, Figures 3-5"}],"minor_comments":[{"comment":"The sentence 'Out of the many events that are simulated, many antennas produce very weak signals. In order to have a balanced dataset, we only retain traces for which the maximum amplitude in one of the two polarizations is above 15 ADC' describes a selection on the clean signal amplitude. This is consistent with Condition 1 of the efficiency metric, but the resulting bias toward high-amplitude events should be stated explicitly when interpreting the efficiency.","section":"Section 3"},{"comment":"The SNR definition uses 'the maximum of the Hilbert envelope of the noiseless trace'; please clarify that this is the maximum over time of the envelope, and specify the units of the standard deviation in the denominator.","section":"Equation (1)"},{"comment":"The figure captions refer to 'South-North axis' and 'East-West axis'; please define these axes in the text or caption, as the two polarizations are not otherwise described.","section":"Section 5, Figures 3 and 4"},{"comment":"The statement 'the denoising efficiency is above 95% for signals with SNR≈4' should be accompanied by the exact SNR bin or the functional form of the efficiency curve, so the reader can see how sharply the transition occurs.","section":"Section 5, paragraph on efficiency"},{"comment":"The phrase 'As we will illustrate in Section 5, our approach successfully meets this challenge' is forward-looking; consider moving the supporting evidence to the results section or rephrasing to avoid a dangling promise.","section":"Section 3, 'Noise' paragraph"},{"comment":"Reference [16] is cited as 'NUTRIG proceedings'; please provide the full arXiv or journal reference so that the ADC noise dataset can be located.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This is a proceedings contribution with a modest scope, but the missing sensitivity baseline is a fundamental issue: the paper's central quantitative claim is not supported without a comparison to raw thresholding or a standard filter. The ambiguity about the false-positive dataset (real noise vs. simulated Gaussian noise) also needs clarification. Both are fixable in revision. The work is otherwise careful in its treatment of noise and its evaluation on held-out data."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The denoiser does what it claims on simulated signals buried in real GRAND noise, and the use of real AN noise power spectra to train on matched noise is a genuinely nice choice. The secondary metrics—peak-time error and amplitude-ratio recovery—show clear improvements over the noisy traces, and the false-positive characterization on pure noise is a step often skipped in this literature. As an ICRC proceedings contribution, this is a competent status report from the GRAND collaboration.\n\nWhat is actually new here is modest but real: the frequency-domain encoder branch, the residual-block architecture, and the specific application to GRANDProto300 with its noise environment. The training/test separation is clean, with simulated Gaussian noise for training and real measured noise for testing, and the authors avoided letting the network learn the pulse position by using random temporal windows. Those are good methodological instincts.\n\nThe soft spots are in the headline claim. The abstract and conclusion say the method produces a 'sensitivity enhancement,' and the 95% efficiency at SNR>4 is the supporting number. But that efficiency is computed only on the denoised output, with no comparison to a raw threshold trigger or a matched/Wiener filter. At SNR>4, a simple amplitude threshold on the raw trace may already pass for most events; the false-positive rate of that raw threshold is never reported. Without a ROC curve, or at least efficiency at fixed false-positive rate for both raw and denoised traces, the sensitivity-enhancement claim is unestablished, even under the paper's own simulation assumptions. The stress-test note gets this right.\n\nTwo smaller issues: there are no error bars on the efficiency curves, and the test set uses simulated signals with real noise, so the 'reality gap' is bridged only on the noise side, not the signal side. The conclusion overstates things slightly when it claims the method 'successfully meets' the reality-gap challenge. The peak-time and amplitude-ratio results do show real reconstruction gains, so I do not think the core method is flawed—just the sensitivity framing.\n\nFor an ICRC proceedings, I would accept this as-is with minor comments. If the authors want to claim sensitivity enhancement in a journal version, they need to add the baseline and a proper ROC comparison. I would not desk-reject it; a referee can push them to quantify the claimed gain and soften the conclusion. If I were in the radio-detection subfield, I would want this comparison in the paper before citing the sensitivity number.","headline":"A solid ML denoising feasibility study for GRANDProto300 with real noise in the test set, but the claimed sensitivity enhancement is not benchmarked against a raw threshold trigger or a conventional filter.","tokens_in":9263,"tokens_out":2674,"would_cite":false,"duration_ms":31185,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper reports that an encoder-decoder neural network can denoise GRANDProto300 radio traces well enough to recover air-shower pulses at signal-to-noise ratios around 4 with over 95% efficiency and a false-positive rate near 1%.","keywords":["denoising","encoder-decoder","convolutional neural network","GRANDProto300","air-shower radio detection","signal-to-noise ratio","cosmic rays","ZHAireS"],"falsifier":"Use the trained model on real GRANDProto300 traces into which calibration pulses of known amplitude and arrival time have been injected at signal-to-noise ratio about 4; if fewer than about 95% of them come out above the 15-ADC threshold, or if the recovered peak is off by more than 10 ns in a non-negligible fraction, the central claim fails. On the noise side, run the model on a long stream of pure recorded noise and count threshold crossings; a false-positive fraction above about 1.4% at 15 ADC would contradict the reported control.","tokens_in":8233,"feed_emoji":"📡","tokens_out":11549,"duration_ms":158273,"temperature":0.7,"pith_summary":"This paper asks whether machine learning can make the GRANDProto300 radio array detect fainter cosmic-ray air showers by cleaning voltage traces before they are used for triggering and reconstruction. The authors build an encoder-decoder convolutional network, train it on simulated ZHAireS air-shower signals mixed with Gaussian noise whose power spectra are taken from real detector noise recordings, and test it on real noise traces. They report that for signals with signal-to-noise ratio above about 4, more than 95% of traces come out with a detectable peak, while pure-noise traces cross the detection threshold only about 1.4% of the time in the south-north channel and 1.1% in the east-west channel. They also report that denoising keeps the recovered peak close to the true peak position and amplitude, where the raw noisy traces fail. If these numbers survive on real events, the method lowers the effective detection threshold of the array without a large false-trigger penalty.","feed_headline":"AI denoiser finds air-shower pulses at SNR 4 in 95% of cases","feed_subtitle":"Trained on simulated showers and real noise power, it keeps false positives near one percent.","key_machinery":"The load-bearing mechanism is an encoder-decoder convolutional network with two parallel encoder branches: one operates in the time domain with a convolutional layer and two residual blocks; the other applies a fast Fourier transform to the input, stacks the real and imaginary parts, and feeds them through the same residual-block structure. The decoder reconstructs the trace from the latent representation, residual connections allow gradient flow in the deep network, and full convolutions make the model independent of trace length. The training set is built from 8000 ZHAireS air-shower simulations processed through GRANDLib to mimic the GRANDProto300 electronics, with only traces whose peak exceeds 15 ADC retained; the pulse position is decorrelated from the trace by training on random 768-bin windows taken from 1024-bin traces. Noise is synthesized as Gaussian traces matching the average power spectra of real ADC-noise recordings, and the network minimizes the L1 distance between the denoised and noiseless trace, a choice that suppresses false reconstructions relative to L2 losses.","core_discovery":"The central claim is that a fully convolutional encoder-decoder, trained with an L1 reconstruction loss on 138,769 simulated signal-plus-noise trace pairs, can denoise GRANDProto300 radio data: for clean peak amplitudes above 15 ADC, the denoised trace exceeds the same threshold in over 95% of cases once the signal-to-noise ratio reaches about 4, on both polarizations. The recovered peak amplitude stays close to the true amplitude as the signal-to-noise ratio drops, whereas the raw noisy traces' peak amplitudes diverge; the fraction of badly timed peaks (offset by more than 10 or 20 ns) is sharply reduced. On pure noise, the false-positive fraction at the 15-ADC threshold is 1.4% for the south-north channel and 1.1% for the east-west channel, and it vanishes when the threshold is raised to 40 ADC. The authors conclude that the denoiser's success on real noise traces, despite being trained on Gaussian noise, supports its use on real detector data.","pith_inferences":["The paper stops at simulated signal pulses, so the open question is whether ZHAireS and GRANDLib trace shapes match real air-shower pulses; if they do, the same recipe should transfer to other radio arrays by swapping the RF-chain simulation.","An ablation removing the frequency-domain branch would isolate how much of the gain comes from the FFT path, a test the paper does not run.","Because false positives vanish at threshold 40 ADC, roughly SNR 3, the denoiser could plausibly serve as an online filter on the trigger path rather than only as an offline analysis step.","Training on power-spectrum-matched Gaussian noise avoids the hidden-signal contamination that raw noise traces would introduce, a subtlety that could be measured by comparing denoisers trained on the two noise types."],"forward_implications":["The GRANDProto300 detection threshold can be pushed down to pulses with signal-to-noise ratio around 4 without paying more than about one percent false triggers.","Peak-time recovery becomes reliable at the level of the detector's GPS timing precision (10–20 ns) for pulses that would otherwise be lost in noise.","Because the model is fully convolutional, the same trained network can denoise traces of arbitrary length, including continuous readout windows.","A denoiser trained only on simulated Gaussian noise demonstrably works when applied to real noise traces, so the pipeline does not require contaminating training data with hidden real signals."],"supporting_citations":[{"why":"Supplies the 8000 ZHAireS air-shower simulations whose electric fields define the signal traces to be recovered.","marker":"[11]"},{"why":"GRANDLib turns the simulated fields into voltage traces that reproduce the GRANDProto300 RF chain and electronics.","marker":"[12]"},{"why":"NUTRIG ADC-noise recordings provide the measured power spectra used to generate Gaussian training noise and the real noise test set.","marker":"[16]"},{"why":"Defines GRANDProto300 timing precision, the 10–20 ns scale used to judge whether denoising preserves peak position.","marker":"[14]"},{"why":"Describes the GRANDProto300 detectors whose trace characteristics the simulated data must match.","marker":"[13]"},{"why":"Residual blocks in the encoder and decoder are the architectural ingredient that lets the network train deeply.","marker":"[9]"},{"why":"Fully convolutional design is what allows the denoiser to handle traces of arbitrary length.","marker":"[10]"}],"fun_headline_variants":["AI denoising recovers air-shower pulses at SNR 4 with 95% hit rate","Neural net denoiser finds air showers at low SNR, false positives ~1%","ML denoiser from GRAND simulations detects 95% of pulses at SNR 4","Deep learning clears noise, air-shower pulses stand out at SNR 4","AI-based radio pulse denoiser: 95% detection at SNR 4, ~1% false alarms"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reported performance numbers assume that simulated ZHAireS air-shower pulses, after GRANDLib's model of the detector electronics, have the same shape and polarization behavior as real pulses recorded by GRANDProto300 in the field.","fun_headline_variants_meta":{"raw":{"variants":["AI denoising recovers air-shower pulses at SNR 4 with 95% hit rate","Neural net denoiser finds air showers at low SNR, false positives ~1%","ML denoiser from GRAND simulations detects 95% of pulses at SNR 4","Deep learning clears noise, air-shower pulses stand out at SNR 4","AI-based radio pulse denoiser: 95% detection at SNR 4, ~1% false alarms"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000607,"raw_usage":{"total_tokens":2799,"prompt_tokens":886,"completion_tokens":1913,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":502,"completion_tokens_details":{"reasoning_tokens":1794}},"tokens_in":502,"tokens_out":1913,"duration_ms":12801,"temperature":1.0,"reasoning_tokens":1794,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T18:57:17.235298+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Use the trained model on real GRANDProto300 traces into which calibration pulses of known amplitude and arrival time have been injected at signal-to-noise ratio about 4; if fewer than about 95% of them come out above the 15-ADC threshold, or if the recovered peak is off by more than 10 ns in a non-negligible fraction, the central claim fails. On the noise side, run the model on a long stream of pure recorded noise and count threshold crossings; a false-positive fraction above about 1.4% at 15 ADC would contradict the reported control.","supporting_citations":[{"cited_title":"Alvarez-Muñizet al","cited_arxiv_id":null,"evidence_quote":"Supplies the 8000 ZHAireS air-shower simulations whose electric fields define the signal traces to be recovered."},{"cited_title":"Heet al.in2016 IEEE CVPR, pp","cited_arxiv_id":null,"evidence_quote":"Residual blocks in the encoder and decoder are the architectural ingredient that lets the network train deeply."},{"cited_title":"Shelhameret al","cited_arxiv_id":null,"evidence_quote":"Fully convolutional design is what allows the denoiser to handle traces of arbitrary length."}],"review_version":1}