{"id":"64372789-4dcd-47b9-8fd0-d7cdad87e93d","arxiv_id":"1908.07270","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Baikal-GVD developed and demonstrated a run-by-run data quality monitoring system that assigns quality ranks from optical module to cluster using exponential, uniformity, Poisson, and charge distribution fits.","lead":"The Baikal-GVD neutrino telescope team built a data quality monitoring system that checks whether detector signals follow expected statistical patterns on a run-by-run basis. It ranks data quality from individual optical modules up to the whole cluster, giving shifters and the multi-messenger network a fast read on detector health.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The quality-ranking thresholds in Section 5 are not validated against any ground-truth set of broken versus healthy channels, so the claimed 'very efficient' data-quality classification is not supported.","rationale":"The reader's weakest-assumption analysis focuses on the unvalidated quality-ranking thresholds in Section 5, and my independent review identifies the same issue as the single most load-bearing concern. The paper is an engineering report describing a monitoring system; its central claim is not a physics measurement but a practical assertion that the DQM system can classify data quality efficiently. For that claim to hold, the cut values must actually separate good from bad detector behavior. The manuscript provides illustrative examples—most notably the LED calibration runs, which clearly distort the exponential, uniformity, and Poisson distributions—but examples are not a validation of threshold values. The thresholds could be too tight (flagging healthy channels as bad due to normal seasonal variations) or too loose (missing a genuinely broken channel). The absence of any labeled ground-truth comparison, ROC analysis, or discussion of threshold selection leaves the reliability of the final quality codes unsupported. This does not, however, invalidate the system's design or its demonstrated responsiveness; it makes the 'very efficiently' conclusion conditional on threshold validation, which is exactly the reader's CONDITIONAL verdict. No additional concern of comparable weight emerged: the system is clearly described, the parameters are physically motivated, and the illustrative figures support the qualitative behavior. Therefore I recommend no change to the reader's verdict, and I agree that the weakest assumption is the untested threshold values.","tokens_in":4779,"tokens_out":1803,"duration_ms":21745,"concrete_test":"Take a set of runs from the 2016 and 2017 seasons for which detector experts have independently labeled the status of individual optical modules (or at least cluster-level run quality) based on known hardware interventions, LED/laser calibrations, and offline manual inspection. Run the Section 5 algorithm on these runs with thresholds held fixed, and compute a confusion matrix for the 'good'/'normal'/'bad' categories against the expert labels. If the sensitivity and specificity are not both high (e.g., AUC > 0.9 across a threshold sweep), the hand-picked thresholds are not validated.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that the DQM system can automatically estimate data quality from the optical-module level up to the whole cluster. The decision algorithm in Section 5 assigns quality codes 2–7 using a fixed, hand-picked set of thresholds: χ²/NDF cut values of 2 and 4, bin-deviation limits R1=25%/50% and R2=1%/5%, noise-rate deviations of 3σ and 5σ, and cumulative-integral limits of 15% and 25%. The manuscript demonstrates that these parameters respond to obvious external disturbances (LED calibration runs) and shows example distributions, but it never establishes that the chosen threshold values correctly separate genuinely malfunctioning detector configurations from benign environmental variations such as seasonal rate changes, temperature drift, or water-transparency fluctuations. Since the quality codes are literal binary decisions ('good', 'normal', 'bad'), the false-positive and false-negative rates are entirely determined by these cut values. Without a labeled data set of known-bad channels or runs, the claim that the system estimates quality 'very efficiently' rests on an untested assumption: that the thresholds were well chosen. There is also no discussion of how the thresholds were derived, whether they were tuned on the 2016 data shown, or whether they generalize to other seasons and detector configurations. This is the most load-bearing weakness because it directly affects the reliability of every quality code the system produces.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes a data quality monitoring (DQM) system for the Baikal-GVD neutrino telescope. The system operates on a run-by-run basis and monitors three statistical parameters (exponential time-difference distribution, uniformity of event rate over time, and Poissonian rate distribution), charge distributions of non-trigger and trigger signals, noise-rate stability, and shape comparison of charge distributions against a reference 'basic run'. Gantt plots are used to identify periods with the LED calibration system active. A quality estimation algorithm combines fit quality (χ²/NDF), bin deviations, noise-rate deviations, and integral-impact deviations into eight quality codes (0–7) at the channel, section, string, and cluster levels. The paper demonstrates the system on examples from the 2016 season, mainly showing that LED calibration runs are flagged as anomalous, and concludes that the system can estimate data quality from the optical-module level up to the whole cluster 'very efficiently'.","tokens_in":5182,"tokens_out":6771,"duration_ms":63988,"significance":"The manuscript addresses a practical and important problem: automated, run-by-run quality assessment for a large underwater neutrino telescope, with output usable by shifters and multi-messenger programs. The strength of the paper is that it presents a concrete, implementable algorithm: the monitored parameters are clearly defined, the expected functional forms are stated, and the quality codes are enumerated. The examples with LED calibration runs provide a useful sanity check that the system can detect an obvious external disturbance. However, the quantitative claim of 'very efficient' quality estimation is not supported by the presented evidence: the threshold values are not validated against labeled data, the aggregation rules for section/string/cluster codes are not specified, and no false-positive/false-negative rates or timing performance are given. Thus the paper is a useful system description, but the performance claim needs either additional validation or a substantial caveat.","major_comments":[{"comment":"The classification into codes 2-7 is fully determined by the hand-picked thresholds (χ²/NDF values 2 and 4; R1=25%/50% and R2=1%/5%; noise-rate deviations of 3σ and 5σ; integral-impact limits of 15% and 25%). The manuscript demonstrates that these parameters respond to LED calibration runs, but it never validates the threshold values against a ground-truth set of known-bad channels or runs, nor does it provide false-positive or false-negative rates. Because these cut values determine every quality code, the conclusion in Section 6 that the system 'allows to estimate the quality ... very efficiently' is not quantitatively supported. The authors should either perform a validation against labeled data or explicitly state that the thresholds are provisional and soften the efficiency claim.","section":"Section 5, Quality estimation algorithm"},{"comment":"The decision chain is stated as 'channel → section → string → cluster levels', but the described criteria apply only to 'channel, section and cluster' and note 'Charge (channel level only)'. The aggregation rule that converts per-channel quality codes into section-, string-, and cluster-level codes is not specified, and the derivation of the LED-related codes 5-7 from the Gantt-plot masks is not described. Without this information, the central claim that the system estimates quality 'from the level of the optical module and up to the whole cluster' cannot be reproduced or tested. Please provide the exact aggregation procedure and the explicit relation between the LED masks and the reported codes.","section":"Section 5, decision chain"},{"comment":"The shape-comparison parameter is defined relative to a 'basic run that is well known to have good channel performance', but the paper does not specify how the basic run is selected, whether it is the same run for all comparisons, or how the 15% and 25% integral-impact thresholds were tuned against this reference. Since this parameter enters directly into the channel quality code, the missing definition of the reference run is a gap in the reproducibility of the algorithm.","section":"Section 3.3"}],"minor_comments":[{"comment":"The quantities NBins, NBinsNormal, NBinsBad, and NBinsTotal used in the bin-deviation condition are never defined in the text, and the roles of R1 and R2 are not explained; please define them explicitly.","section":"Section 5, item 2"},{"comment":"The threshold conditions leave the boundary values undefined: for χ²/NDF the good/normal boundary is '<4' and the normal/bad boundary is '>4', leaving exactly 4 unassigned; similarly for exactly 5σ noise deviation and exactly 25% integral impact. Please state the convention for ties.","section":"Section 5, item 4"},{"comment":"The sentence 'Thus in standard conditions we expect maximum a doubling of the rate during the run' does not follow directly from the preceding statement about the 0.5 entries per bin threshold in the Gantt plot; please clarify the quantitative relationship.","section":"Section 4"},{"comment":"The criterion 'Number of events with charge value deposited in channel more than 100 p.e. should not exceed 100 events' is not referenced in the quality estimation algorithm; clarify whether it contributes to codes 5-7 or is a separate masking step.","section":"Section 3.3"},{"comment":"The abstract and Section 1 list participation in the global multi-messaging system as a design goal, but no latency, interface, or data-volume figures are reported; if this is outside the scope of the paper, this should be stated explicitly.","section":"Section 6 / Abstract"}],"recommendation":"major_revision","confidential_remarks":"The paper is a conference proceedings contribution and may be acceptable in that venue with a softened conclusion. For a journal submission, the missing validation and the unspecified aggregation rule are substantive gaps; I recommend major revision. Also, the authors do not discuss prior art on detector DQM systems, which would help place the contribution."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a system description paper for the Baikal-GVD run-by-run data quality monitoring system. The genuinely new part is the multi-level ranking algorithm (channel to section to string to cluster) and the Gantt-plot-based event selection, neither of which appears in the cited Baikal-GVD references. The paper does a solid job of laying out the monitored parameters -- exponential time-difference, uniformity, Poisson rate, charge distributions, and threshold/calibration ratios -- and it shows concrete examples where LED calibration runs are flagged as anomalous. That demonstration is credible; the system clearly responds to an obvious external disturbance.\n\nThe soft spot is exactly the one the stress-test flags: the quality thresholds in Section 5 are presented formulaically (chi2/NDF <2/<4, R1=25%/50%, R2=1%/5%, noise deviations <3sigma/<5sigma, integral impact <15%/<25%) with no derivation and no ground-truth validation. The paper never shows that these cut values separate genuinely broken channels from benign seasonal or environmental variation, and it offers no false-positive or false-negative rates. So the conclusion's claim that the system estimates data quality 'very efficiently' is simply not supported by the evidence shown. That is a real gap, but it is a gap common to many operational monitoring papers -- the system exists, it flags known problems, and the thresholds can in principle be tuned. The absence of validation does not make the described system implausible; it makes the efficiency claim unproven.\n\nWhat the paper does not do is also worth saying: it gives no comparison with DQM systems at other neutrino telescopes, no details on how thresholds were chosen or whether they were fit to the 2016 data shown, and no code or configuration file. That limits how much a reader outside the collaboration can transfer.\n\nThis paper is for people working on running water/ice neutrino detectors or building similar monitoring stacks. It is a legitimate engineering/operations contribution, not a scientific breakthrough. With a request for validation -- sensitivity/specificity on known-bad runs, threshold-tuning discussion, and a toned-down conclusion -- it deserves to go to peer review rather than being desk-rejected.","headline":"A solid, honest description of Baikal-GVD's monitoring system, but the quality-ranking thresholds are unvalidated and the 'very efficiently' conclusion overreaches.","tokens_in":5900,"tokens_out":2568,"would_cite":false,"duration_ms":24986,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["95.55.Vj"],"model":"deepseek-v4-flash","headline":"By fitting event and charge distributions, the Baikal-GVD DQM system assigns a quality code to every optical module and rolls it up to whole-cluster level.","keywords":["Baikal-GVD","data quality monitoring","neutrino telescope","run-by-run monitoring","charge distributions","optical module","quality codes","underwater Cherenkov detector"],"falsifier":"Run the algorithm on a set of 2016 runs that include known LED calibration periods and channels independently judged by a human expert, then check the agreement between the eight quality codes and the known status; if a known-broken channel is marked \"good\" or a known-good channel is marked \"bad\" often enough, the threshold choices are not separating the two populations.","tokens_in":4610,"feed_emoji":"🔭","tokens_out":3896,"duration_ms":38196,"temperature":0.7,"pith_summary":"This paper presents the Data Quality Monitoring system built for the Baikal-GVD neutrino telescope, and claims that it can estimate the quality of experimental data run-by-run, from a single optical module up to the entire cluster. The system fits three statistical distributions, exponential, uniformity, and Poisson, to event timing and rates, and analyzes PMT charge distributions and Gantt plots. Each channel is then assigned one of eight quality codes, and the decision chain aggregates these codes upward through sections, strings, and clusters. The authors show examples from the 2016 season, concluding that the parameters allow efficient, automatic quality estimation.","feed_headline":"Baikal-GVD monitors data quality from single sensor to full cluster","feed_subtitle":"A run-by-run DQM system fits event and charge distributions and assigns quality codes at every detector level.","key_machinery":"The load-bearing mechanism is a threshold-based decision chain over eight quality codes. The fit quality is judged by $\\chi^2/\\mathrm{NDF}$ with good/normal/bad boundaries at 2 and 4; per-bin deviations must satisfy $R_1 = 25\\%$ or $50\\%$ with $R_2 = 1\\%$ or $5\\%$; noise-rate deviations per depth level must be within $3\\sigma$ or $5\\sigma$; and the cumulative integral in the charge distribution must differ from the basic run by less than $15\\%$ or $25\\%$. These criteria are applied at channel, section, string, and cluster levels, so a single scalar code propagates upward.","core_discovery":"The central claim is that a compact set of automated fits and comparisons can reliably characterize detector health at every level of Baikal-GVD without manual inspection. For each run, the system checks whether time differences between events follow an exponential, whether event rate is uniform and Poisson-distributed, where the 1 p.e. charge peak sits, whether trigger thresholds are stable relative to the calibration, and whether charge-distribution shapes match a known-good \"basic run\". Quality is summarized in eight codes, from \"excluded by configuration\" and \"empty data\" through \"good\", \"normal\", and \"bad\", each optionally flagged as containing LED-calibration light. The paper's conclusion is that the 2016 examples show this multi-parameter analysis estimates data quality very efficiently from the optical-module level to the whole cluster.","pith_inferences":["The fixed thresholds would likely benefit from calibration against a labeled set of known-broken channels; without that, codes cannot guarantee the same false-positive rate across seasons.","The \"basic run\" comparison may drift as Lake Baikal's water transparency changes seasonally, so the system probably needs periodic re-baselining.","A similar channel-to-cluster quality ladder could be adapted to other large underwater or ice neutrino detectors with minimal changes.","The method could be tested live by running it simultaneously with human expert inspection during a future season and measuring agreement."],"forward_implications":["Detector shifters can see a per-run quality map of every optical module without manually opening event displays.","Analysis pipelines can automatically reject runs or channels whose quality code is \"bad\" or \"LED detected\".","The same run-by-run summary could feed the global multi-messenger alert system, since detector state is known for every time window.","Early detection of failing PMTs or calibration anomalies becomes possible within one run rather than after offline reprocessing.","Quality codes give an audit trail for the 2016 dataset and later seasons."],"supporting_citations":[{"why":"Supplies the detector layout and status that defines the cluster, string, section, and channel structure the DQM system monitors.","marker":"[1]"},{"why":"Supplies the trigger scheme and the event time and charge data that the DQM statistical and charge distributions analyze.","marker":"[2]"},{"why":"Supports the expected depth dependence of noise rates used in the noise-deviation criterion for channel quality.","marker":"[3]"}],"fun_headline_variants":["Baikal-GVD auto-fits data quality at every level","Eight quality codes flag Baikal-GVD run health","Run-by-run DQM for Baikal-GVD from module to cluster","Automated fits judge Baikal-GVD detector status per run","Baikal-GVD DQM: sensor-to-cluster quality codes"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The hand-picked threshold values in Section 5 are assumed to separate genuinely bad detector states from normal seasonal or environmental variation, but the paper does not validate them against a labeled set of broken versus healthy channels.","fun_headline_variants_meta":{"raw":{"variants":["Baikal-GVD auto-fits data quality at every level","Eight quality codes flag Baikal-GVD run health","Run-by-run DQM for Baikal-GVD from module to cluster","Automated fits judge Baikal-GVD detector status per run","Baikal-GVD DQM: sensor-to-cluster quality codes"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000217,"raw_usage":{"total_tokens":1338,"prompt_tokens":751,"completion_tokens":587,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":367,"completion_tokens_details":{"reasoning_tokens":500}},"tokens_in":367,"tokens_out":587,"duration_ms":5890,"temperature":1.0,"reasoning_tokens":500,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T12:20:12.081731+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the algorithm on a set of 2016 runs that include known LED calibration periods and channels independently judged by a human expert, then check the agreement between the eight quality codes and the known status; if a known-broken channel is marked \"good\" or a known-good channel is marked \"bad\" often enough, the threshold choices are not separating the two populations.","supporting_citations":[{"cited_title":"Avrorin et al, Baikal-GVD: status and prospects , in proceedings of XXth International Seminar on High Energy Physics (QUARKS-2018) , EPJ Web Conf","cited_arxiv_id":null,"evidence_quote":"Supplies the detector layout and status that defines the cluster, string, section, and channel structure the DQM system monitors."},{"cited_title":"Avrorin et al, Data acquisition system of the NT1000 Baikal neutrino teles cope, Instruments and Experimental Techniques 3 (2014) 262–273","cited_arxiv_id":null,"evidence_quote":"Supplies the trigger scheme and the event time and charge data that the DQM statistical and charge distributions analyze."},{"cited_title":"Avrorin, R","cited_arxiv_id":null,"evidence_quote":"Supports the expected depth dependence of noise rates used in the noise-deviation criterion for channel quality."}],"review_version":1}