{"id":"8a4deb05-f32e-414d-ac61-c54931e3e131","arxiv_id":"1909.00840","paper_version":3,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"After 4.5 years of operation, the ARIANNA test bed found no neutrino candidates and set a 90% confidence upper limit of E^2 Phi = 1.7e-6 GeV cm^-2 s^-1 sr^-1 at 10^18 eV.","lead":"A seven-station radio telescope in Antarctica ran for 4.5 years, saw no ultra-high-energy neutrinos, and set a new upper limit on how many such neutrinos can be arriving from space. The result is an order of magnitude stronger than the same team's earlier measurement and shows the detector technology works reliably in polar ice.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Limit's systematic bias depends on unvalidated simulation inputs; a Veff re-run with the measured basal reflection is the decisive check.","rationale":"The reader's weakest_assumption is essentially the same concern: Veff accuracy and the reflection coefficient discrepancy. I agree that this is the most load-bearing point because Eq. (4.1) is linear in Veff, and Fig. 5 shows reflected signals dominate the effective volume. The measured sqrt(R)=0.82 from Sec. 2.2 (power 0.67) sits well below the simulation input R=0.9, and the paper does not propagate this into a systematic uncertainty on the limit. This is not an internal inconsistency in the sense of a mathematical error; it is an unquantified systematic bias in a headline number. The paper deserves credit for the transparent template matching, conservative Feldman-Cousins choice with zero expected background, and the measured livetime, but the missing Veff systematics is exactly what the reader flagged and what should be resolved before this limit is used as a precise benchmark. A single simulation re-run with the measured reflection coefficient would settle the magnitude of the bias. One caveat: the text says correcting for the reflection coefficient gives ~500 m attenuation, which is used in Table 1; it is possible the authors intended R=0.9 as an upper-bound/conservative choice for the shelf model, but the direction of the resulting bias is not stated or quantified, so the concern stands. I would therefore keep the CONDITIONAL verdict rather than ACCEPT or REJECT.","tokens_in":17847,"tokens_out":1563,"duration_ms":14328,"concrete_test":"Re-run the ShelfMC effective-volume simulation twice, once with the Table 1 value R = 0.9 and once with the measured value R = 0.67 (or sqrt(R) = 0.82, scaling the reflected ray amplitudes accordingly), keeping all other parameters in Table 1 fixed; then recompute the limit in Eq. (4.1). If the revised limit at 1e18 eV moves by more than about 20%, the quoted limit needs a systematic band or a corrected central value.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is the 90% CL upper limit E2Phi = 1.7e-6 GeV cm^-2 s^-1 sr^-1 at 1e18 eV, which is inversely proportional to Veff from Eq. (3.2). The single most load-bearing input is the basal reflection coefficient used in ShelfMC: Table 1 sets R = 0.9, while Sec. 2.2 reports an independently measured electric-field reflection coefficient sqrt(R) = 0.82 +/- 0.07, i.e., power R = 0.67 +0.13/-0.11. The measured value is not merely lower but below the simulation value with a discrepancy exceeding the measurement uncertainty. Because reflected-path neutrinos dominate the sensitivity (Fig. 5), the analysis efficiency and Veff scale strongly with R. If the true R is near the measured 0.67, Veff could be overestimated by tens of percent, making the quoted limit correspondingly optimistic. The paper does not quantify how Veff depends on R, nor does it quote a systematic band on the limit. Since the headline number is meant as a benchmark for ARIANNA sensitivity, this unquantified dependence is the weakest load-bearing point. A second, related weakness is the power-law fit to the high-chi_ave tail of the background with no explicit goodness-of-fit or systematic check, but the dominant concern remains the reflection coefficient entering Veff.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a search for ultra-high-energy neutrinos with the seven-station ARIANNA test bed at Moore's Bay, using data collected between December 2014 and February 2019. After defining a signal region in the (chi_ave, SNR) plane with a template-matching procedure, no neutrino candidates remain, with a combined signal efficiency of 79%. The central result is a 90% confidence upper limit on the diffuse neutrino flux of E^2 Phi = 1.7e-6 GeV cm^-2 s^-1 sr^-1 in a decade-wide bin centered at 1e18 eV, an order-of-magnitude improvement over the previous ARIANNA limit. The limit is obtained from the simulated effective volume, the measured livetime, the analysis efficiency, and a Feldman-Cousins factor.","tokens_in":18035,"tokens_out":6477,"duration_ms":67739,"significance":"If correct, the result demonstrates the long-term reliability of the ARIANNA surface-radio architecture and improves the ARIANNA diffuse-flux limit by an order of magnitude near 1e18 eV. The analysis is transparent and conservative in several respects: it excludes periods of camp occupation and HiCal pulser operation, corrects for DAQ deadtime, uses only transferred data while estimating the recoverable fraction, and conservatively adopts zero expected background in the limit calculation. The main caveat is that the quoted limit scales inversely with the simulated effective volume, whose key inputs are not systematically varied or assigned uncertainties.","major_comments":[{"comment":"The ShelfMC simulation uses a basal power reflection coefficient of R = 0.9 (Table 1), while Sec. 2.2 quotes the measured electric-field reflection coefficient sqrt(R) = 0.82 +/- 0.07 from [30], corresponding to a power coefficient of roughly 0.67 +/- 0.13. Since the reflected path dominates the effective volume (Fig. 5) and the limit in Eq. (4.1) is inversely proportional to Veff (Eq. (3.2)), using R = 0.9 in the simulation could overestimate Veff and make the quoted limit optimistic if the measured value is correct. The paper does not quantify d ln Veff/dR nor report a systematic band from varying R. I request a rerun of ShelfMC with R = 0.67 (or at least a scan over R within the measurement uncertainty) and a corresponding revision of the limit, or a clear justification for why 0.9 is the appropriate input despite the quoted measurement.","section":"Table 1 and Sec. 2.2"},{"comment":"The signal-region boundary is set by extrapolating a power-law fit to the high-chi_ave tail of the background, with no reported goodness-of-fit, fit uncertainty, or comparison with an alternative background model. The expected background of 0.5 events depends on this extrapolation. Although the analysis subsequently sets the expected background to zero, which is conservative for the limit, the claimed signal efficiency of 79% and the statement that the region contains no events are tied to the fitted boundary. I ask the authors to show the fit quality explicitly and to test the sensitivity of the boundary to the number of tail points used and to the choice of functional form.","section":"Sec. 4.2 and Fig. 10"},{"comment":"No systematic uncertainties are propagated into the quoted limit. The inputs entering Eq. (4.1) all carry uncertainties: the attenuation length (Sec. 2.2, 460 +/- 20 m at low frequency), the basal reflection coefficient (0.82 +/- 0.07), the ice thickness (576 +/- 8 m), the livetime, and the analysis efficiency. Because the result is an upper limit with no observed candidates, the numerical value of the limit is the main physics output and should be accompanied by a systematic band, or at least by a statement of how much Veff, and hence the limit, changes under these variations.","section":"Sec. 4.3 / Eq. (4.1)"}],"minor_comments":[{"comment":"The abstract says '4.5 years of data' while the livetime is given as 2906.9 days (7.96 station-years); please clarify that 4.5 years refers to the calendar span and not the total detector livetime.","section":"Sec. 3.2 / Abstract"},{"comment":"The two panels of Fig. 9 use different color scales (events per bin up to 10^3 and 10^4, respectively); a common color scale would make the comparison between the 100-series and 200-series stations clearer.","section":"Fig. 9"},{"comment":"In the sentence 'was operated for less that 40 days,' 'less that' should be 'less than.'","section":"Sec. 4.5"},{"comment":"The sentence beginning 'At a distance of 110 km...' is grammatically awkward ('is relatively close proximity to'); it should be reworded for clarity.","section":"Sec. 2.2"}],"recommendation":"major_revision","confidential_remarks":"The main load-bearing issue is the inconsistency between the simulated basal reflection coefficient R = 0.9 and the measured value sqrt(R) = 0.82 +/- 0.07 quoted in the same paper. This should be resolved before acceptance, either by rerunning the simulation with the measured value or by providing a quantitative systematic treatment. The absence of any systematic band on the headline limit is also worth insisting on for a journal paper. The analysis itself appears sound in its central logic, and the conservative choices in livetime and background treatment are commendable."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Mike—\n\nQuick take: this is the sort of paper that should get published after a round of small fixes. It reports a new 90% CL upper limit of E^2Phi = 1.7e-6 GeV cm^-2 s^-1 sr^-1 at 10^18 eV from 4.5 years of seven-station ARIANNA data, roughly an order of magnitude better than the collaboration’s 2015 bound, and it documents long-term station reliability at Moore’s Bay and the South Pole. The result is new, and the analysis is easy to follow.\n\nWhat it does well: the limit calculation is transparent, livetime corrections are careful (excluding camp operations and HiCal, correcting for DAQ deadtime), and they deliberately set expected background to zero instead of using their 0.5-event estimate, making the limit slightly conservative. The Monte Carlo, template matching, and signal-region construction are standard, and the paper does not oversell the physics reach—it explicitly says the limit is not competitive with IceCube, ANITA, Auger, or ARA. The citation pattern is fine; ref. [30] is the right source for the measured reflection coefficient, and the comparison limits are current for 2019.\n\nSoft spots, in order of importance:\n\n1. Reflection coefficient ambiguity. Table 1 lists “Reflection Coefficient 0.9” but Sec. 2.2 quotes a measured electric-field coefficient sqrt(R) = 0.82 +/- 0.07. The paper never states whether the simulation parameter is amplitude or power. If it is power, the simulation is using a notably more reflective bed than measured and Veff could be overestimated enough to move the limit by tens of percent. If it is amplitude, the tension is about one sigma and much less concerning. A one-line clarification plus a quick Veff run at the measured value would settle this. It is the only load-bearing issue I see.\n\n2. No systematic band on the quoted limit. At this sensitivity, missing systematics does not change the conclusion, but if 1.7e-6 is going to be used as an ARIANNA benchmark, people need to know what input variations do to it.\n\n3. Minor: the signal region is defined from the same data that is then inspected. I do not think it is a real problem here because the cut is driven by a fitted background tail with tiny expected background, but it should be acknowledged explicitly. The power-law tail fit also deserves a goodness-of-fit statement.\n\nThe stress-test note’s concern about the reflection coefficient is on target only if “0.9” is a power coefficient. The paper never says, so the fix is presentational and numerical, not a reason to reject.\n\nWho this is for: experimental astroparticle neutrino people, especially those benchmarking radio-array sensitivity or planning the next ARIANNA/ARIA step. It deserves a serious referee; I would send it to review and ask for the clarification, not desk-reject.","headline":"A transparent, honest ARIANNA search that improves their own limit by an order of magnitude; the one load-bearing caveat is an ambiguous bed-reflection coefficient in the simulation, which needs stating and a quick systematic check, not a rejection.","tokens_in":18757,"tokens_out":3323,"would_cite":true,"duration_ms":35997,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["95.55.Vj","95.85.Ry"],"model":"deepseek-v4-flash","headline":"After 4.5 years of listening from the Ross Ice Shelf, the seven-station ARIANNA test bed found no neutrino candidates and set a 90% confidence upper limit of $E^2\\Phi = 1.7\\times 10^{-6}$ GeV cm$^{-2}$ s$^{-1}$ sr$^{-1}$ in the decade…","keywords":["ultra-high-energy neutrinos","cosmogenic neutrinos","Askaryan effect","radio detection in ice","Ross Ice Shelf","diffuse flux upper limit","ARIANNA","Monte Carlo effective volume"],"falsifier":"A calibration experiment would settle it: fire a pulser at a known depth—or use the stations' own heartbeat transmitter—repeatedly, and compare the measured trigger rate with the Monte Carlo prediction for the same geometry. If measurements fall systematically below prediction, the effective volume is overestimated and the flux limit should move upward; recomputing the limit with the measured basal reflection coefficient $\\sqrt{R}=0.82$ instead of 0.9 is a concrete immediate version of that test.","tokens_in":17617,"feed_emoji":"📡","tokens_out":12291,"duration_ms":116175,"temperature":0.7,"pith_summary":"Ultra-high-energy neutrinos—ghostly particles produced when the most energetic cosmic rays collide with background radiation—should produce a nanosecond radio flash when they interact in Antarctic ice. ARIANNA's seven solar-powered stations on the Ross Ice Shelf listened for that flash between December 2014 and February 2019 and found no neutrino candidates. The paper converts the null result into a 90% confidence upper limit on the diffuse flux: $E^2\\Phi = 1.7\\times 10^{-6}\\,\\mathrm{GeV\\,cm^{-2}\\,s^{-1}\\,sr^{-1}}$ in the decade centered at $10^{18}\\,\\mathrm{eV}$, an order of magnitude stronger than the collaboration's previous limit. The larger point is that a sparse, surface, radio-quiet detector can run reliably for years, reject all backgrounds with a simple template match, and serve as the basis for a much larger telescope.","feed_headline":"No cosmogenic neutrinos found in 4.5 years of Antarctic radio data","feed_subtitle":"New 90% confidence flux bound at 10^18 eV is ten times stronger than the previous ARIANNA limit.","key_machinery":"The argument rides on two linked objects. The analysis tool is template matching in the two-dimensional space of $\\chi_{\\rm ave}$ versus SNR: simulated neutrino signals concentrate at high correlation and high amplitude, while thermal and environmental backgrounds fall off as a power law in that space. The background tail is fit and extrapolated to set a signal-region boundary that yields an expected background of 0.5 events, and the same boundary is applied to the data. The flux conversion then uses the simulated effective volume $V_{\\rm eff}$ from the ShelfMC Monte Carlo, which generates Askaryan emission, propagates direct and ice-water-reflected rays through the ice shelf, and applies the stations' 2-of-4 trigger logic; Eq. (4.1) divides the Feldman-Cousins factor by $V_{\\rm eff}$, livetime, and efficiency. The effective volume is what converts 'no events seen' into a flux limit, so its accuracy is the mechanism that carries the whole result.","core_discovery":"Over the full data set the seven-station test bed collected 2906.9 days of livetime, corrected for readout deadtime and for periods when field camp or calibration activities could contaminate data. For each triggered waveform the analysis computes $\\chi_{\\rm ave}$, the best Pearson correlation against a library of simulated neutrino templates averaged over co-polarized antenna pairs, and plots it against signal-to-noise ratio (SNR). The signal region is defined from the data itself: the background tail in each SNR bin is fit to a power law and extrapolated to the threshold where only 0.5 background events would be expected over the whole livetime, with the threshold weighted by where simulated neutrinos fall. The signal region retains 81% of weighted simulated neutrinos for the older amplifier series and 78% for the newer series—a combined efficiency of 79%—and no triggered event falls inside it. Using the Feldman-Cousins 90% upper limit with zero observed and zero expected background, the authors obtain $E^2\\Phi \\le 1.7\\times 10^{-6}\\,\\mathrm{GeV\\,cm^{-2}\\,s^{-1}\\,sr^{-1}}$ for a decade-wide bin centered at $10^{18}\\,\\mathrm{eV}$.","pith_inferences":["Editorial extension: combined with current cosmogenic neutrino models, a limit this low starts to squeeze the most optimistic proton-dominated scenarios for the highest-energy cosmic rays, although the test bed alone cannot exclude them.","Editorial extension: the simulated effective volume uses a basal power reflection coefficient of 0.9 while the paper's own site measurement implies $\\sqrt{R}=0.82$; rerunning the simulation with the measured value is a direct way to test how much the limit would soften.","Editorial extension: the Moore's Bay field of view sweeps across the declination band containing the flaring blazar TXS 0506+056, so a scaled array at this site could follow up the same class of neutrino-emitting blazars."],"forward_implications":["At 90% confidence, the true diffuse ultra-high-energy neutrino flux in the decade centered at $10^{18}\\,\\mathrm{eV}$ lies below $1.7\\times 10^{-6}\\,\\mathrm{GeV\\,cm^{-2}\\,s^{-1}\\,sr^{-1}}$, assuming the simulated effective volume is correct.","A simple template-matching cut with an expected background of 0.5 events can run for 7.96 station-years and admit zero background events, so the same analysis method scales cleanly to a much larger array.","Sustained operation at Moore's Bay and at the South Pole demonstrates that the solar-powered, self-contained station design is reliable enough and radio-quiet enough to serve as the building block for a large-area telescope.","The test bed's transient-source sensitivity is already comparable to previous instruments' sensitivity in the direction of GW170817, so even a pilot array can contribute to multi-messenger follow-up campaigns.","A future 130-station array based on this technology, run for five years, would be sensitive enough to constrain the proton fraction of ultra-high-energy cosmic rays to 10% or less."],"supporting_citations":[{"why":"The previous ARIANNA search whose limit this work improves by an order of magnitude, and the origin of the template-matching approach.","marker":"[16]"},{"why":"Defines the ShelfMC Monte Carlo code that produces the effective volume and simulated neutrino signals used throughout the analysis.","marker":"[28]"},{"why":"Supplies the measured ice thickness, depth-averaged attenuation length, and basal reflection coefficient that set the site parameters in Table 1.","marker":"[30]"},{"why":"Provides the parametrization of Askaryan radio emission from high-energy showers used to generate signal waveforms in simulation.","marker":"[48]"},{"why":"Supplies the GZK spectrum used to weight simulated neutrino interactions and to define the signal energy distribution.","marker":"[55]"},{"why":"Gives the neutrino-nucleon cross section used to compute the water-equivalent interaction length in the flux limit formula.","marker":"[61]"},{"why":"Provides the Feldman-Cousins 90% confidence prescription used to convert zero observed events into an upper limit.","marker":"[62]"},{"why":"Describes the proposed ARIA array whose projected five-year sensitivity is used as the benchmark for scaling up the technology.","marker":"[17]"}],"fun_headline_variants":["ARIANNA's 4.5-year neutrino search yields null result","Antarctic radio array sets new limit on ultra-high-energy neutrinos","No neutrinos seen in 4.5 years of Antarctic ice radio data","ARIANNA test bed places tenfold stronger limit on cosmic neutrinos","4.5 years of radio data: no cosmic neutrinos, tighter bound"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole limit is inversely proportional to the simulated effective volume, so the load-bearing premise is that the Monte Carlo correctly predicts how often neutrinos would make a station trigger; in particular it assumes a depth-averaged attenuation length of 500 m, a reflection coefficient of 0.9 at the ice-water interface, and ignores interactions in the shadow zone where ray bending blocks signals from reaching the surface.","fun_headline_variants_meta":{"raw":{"variants":["ARIANNA's 4.5-year neutrino search yields null result","Antarctic radio array sets new limit on ultra-high-energy neutrinos","No neutrinos seen in 4.5 years of Antarctic ice radio data","ARIANNA test bed places tenfold stronger limit on cosmic neutrinos","4.5 years of radio data: no cosmic neutrinos, tighter bound"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000641,"raw_usage":{"total_tokens":3020,"prompt_tokens":1087,"completion_tokens":1933,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":703,"completion_tokens_details":{"reasoning_tokens":1835}},"tokens_in":703,"tokens_out":1933,"duration_ms":16820,"temperature":1.0,"reasoning_tokens":1835,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T05:34:58.963813+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A calibration experiment would settle it: fire a pulser at a known depth—or use the stations' own heartbeat transmitter—repeatedly, and compare the measured trigger rate with the Monte Carlo prediction for the same geometry. If measurements fall systematically below prediction, the effective volume is overestimated and the flux limit should move upward; recomputing the limit with the measured basal reflection coefficient $\\sqrt{R}=0.82$ instead of 0.9 is a concrete immediate version of that test.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The previous ARIANNA search whose limit this work improves by an order of magnitude, and the origin of the template-matching approach."},{"cited_title":"Persichilli,Performance and Simulation of the ARIANNA Pilot Array, with Implications for Future Ultra-high Energy Neutrino Astronomy, Ph.D","cited_arxiv_id":null,"evidence_quote":"Defines the ShelfMC Monte Carlo code that produces the effective volume and simulated neutrino signals used throughout the analysis."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Supplies the measured ice thickness, depth-averaged attenuation length, and basal reflection coefficient that set the site parameters in Table 1."},{"cited_title":"Alvarez-Muñiz, R","cited_arxiv_id":null,"evidence_quote":"Provides the parametrization of Askaryan radio emission from high-energy showers used to generate signal waveforms in simulation."},{"cited_title":"Engel, D","cited_arxiv_id":null,"evidence_quote":"Supplies the GZK spectrum used to weight simulated neutrino interactions and to define the signal energy distribution."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Gives the neutrino-nucleon cross section used to compute the water-equivalent interaction length in the flux limit formula."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the Feldman-Cousins 90% confidence prescription used to convert zero observed events into an upper limit."}],"review_version":1}