{"id":"2c69bce7-67f6-45aa-bcb6-1303bfc0904f","arxiv_id":"2607.22477","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A full-scale four-cell WOM-based liquid-scintillator veto prototype for SHiP reconstructed muon tracks with about ±15 cm spatial and ±15° angular resolution.","lead":"A four-cell, full-size prototype of the veto detector planned for the SHiP particle-physics experiment was tested with a 5 GeV muon beam. It reconstructed where particles crossed the detector and at what angle, with centimetre- and degree-level precision, supporting the viability of the design.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Likelihood-reconstruction resolutions may be optimistic: the same four-cell data are used to tune the simulation/calibration (α=75%, Eq. 3.3) and to evaluate the quoted σ-values, with no described train/validation split.","rationale":"The reader identified the transferability/validation issue as the weakest assumption; I concur. I considered whether the low average fraction of correctly identified events (0.48 for Analysis 2, Sec. 3.6.2) is a separate fatal flaw, but the paper quotes continuous residual standard deviations, so a low exact-bin match is not by itself inconsistent with the stated σ-values; the decisive question is whether those σ-values survive on independent data. The reflectivity fit (Eq. 3.3) is a concrete channel for leakage between calibration and evaluation, and the absence of any described train/test split makes the headline reconstruction resolutions the least secure part of the paper. A hold-out test is the cleanest arbiter. This does not change the reader's CONDITIONAL verdict, but it sharpens the condition: the authors should supply an independent validation or explicitly state that the quoted resolutions are calibration-limited.","tokens_in":16262,"tokens_out":6883,"duration_ms":75735,"concrete_test":"Perform a leave-one-position-out cross-validation on the data of Fig. 17: build the likelihood templates using all measured positions/angles except one complete position (or one full angle series), then compute the residual σX, σY, σθX, σθY and the fraction of correctly assigned events on the held-out position. Repeat for all positions. If the average held-out resolutions are consistent with the quoted values (±15.5 cm, ±6.8 cm, ±15°, ±15°) within, say, 20–30%, the concern is resolved. If they degrade significantly, the quoted numbers should be reported as calibration-limited rather than as general performance.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline numbers (Sec. 3.6.2: ±15.5 cm in X, ±6.8 cm in Y, ±15° in θX and θY) come entirely from the likelihood-based reconstruction. The method is said to be adopted from [6] and 'the same correction function' is applied (Secs. 3.4.2 and 3.6.1), but the paper never describes a training/validation split or a set of held-out beam positions/angles. Moreover, the GEANT4 simulation that underpins the detector response is explicitly tuned to the same four-cell test-beam data: Eq. 3.3 minimizes χ² over the measured crossing points to fix α=75%. If the likelihood templates are derived from this simulation (or calibrated on these runs), evaluating the residuals in Fig. 18 on the same data biases σX, σY, σθ downward. If instead the templates are frozen from [6], their transfer to cells of different sizes is asserted without a dedicated validation. The final SBT will see tracks at arbitrary positions/angles, so the relevant quantity is generalization, not self-consistency. This concern is separate from the raw charge/timing measurements, which are direct, but it is precisely the reconstruction performance that supports the central claim.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports on a full-scale 2x2-cell liquid-scintillator prototype for the SHiP Surrounding Background Tagger, read out by wavelength-shifting optical modules (WOMs) coupled to SiPMs. The detector was exposed to 5 GeV muons at the CERN PS T9 beam line over a range of crossing positions and incident angles. The authors characterize the integrated light yield and arrival-time response, compare measurements with a GEANT4 simulation, and use a likelihood-based correction/reconstruction — inherited from the collaboration's single-cell study [6] — to correct the response and to reconstruct particle crossing coordinates and angles. The main quantitative claims are a timing variation corrected to better than 0.4 ns after likelihood correction, and, for the multi-cell reconstruction, spatial resolutions of ±15.5 cm in X and ±6.8 cm in Y and angular resolutions of about ±15° in both θX and θY for tracks crossing several cells. The raw waveforms, charge measurements, and timing calibrations are presented in detail and appear credible; the central question is whether the quoted reconstruction performance generalizes, since the simulation is tuned to the same data and the likelihood correction is not described with a validation split.","tokens_in":16588,"tokens_out":4931,"duration_ms":59663,"significance":"If the reconstruction claims hold, this is a valuable milestone for the SHiP SBT R&D: it is the first multi-cell prototype demonstration with full-size cells, and the timing correction to <0.4 ns using only ToT/timestamp observables is directly relevant to the final detector readout. The paper also gives useful quantitative information on light-yield uniformity, effective signal speed, and the need to model BaSO4 reflectivity below manufacturer specifications. The measurements are anchored to an external beam telescope, so the raw performance numbers are not circular by construction. The main risk is that the headline spatial/angular resolutions come from a likelihood method whose calibration and simulation input are connected to the same test-beam dataset used for the evaluation; if so, the quoted σ values are optimistic. The paper would be strengthened by a clear validation protocol or an explicit statement that all correction parameters were fixed before this dataset was analyzed.","major_comments":[{"comment":"The headline resolutions are obtained from the same four-cell dataset used to tune the simulation: Eq. (3.3) minimizes χ² to set the relative reflectivity α=75% against the measured light yields of this campaign, and the likelihood correction/reconstruction is described as the 'same correction function' as [6] without stating whether any parameters were re-estimated on these runs. If the likelihood templates or correction parameters were calibrated on the same positions/angles used to produce Fig. 18, the quoted σX≈15.5 cm, σY≈6.8 cm, σθ≈15° are in-sample estimates and likely optimistic. Please provide a validation protocol — e.g., leave-one-position/angle-out or a pre-registered split — or, if all parameters are frozen from [6], state that explicitly and show that the [6] calibration is statistically independent of the data in Fig. 17.","section":"Sec. 3.6.2 / Eq. (3.3)"},{"comment":"The paper reports angular resolutions of ~15°–19° while Fig. 15 shows only 52–69% of events assigned the correct incident angle (average fractions 0.57/0.69 for Position 1). In Sec. 3.6.2 the average fraction is given as 0.48/0.45, yet the text calls this 'correctly identified track crossing points and incident angles' even though Fig. 15 concerns angles only. A Gaussian σ of ~15° on a 15° grid can be consistent with ~50% exact-bin assignment, so this is not an internal contradiction, but the criterion for 'correctly identified' must be defined, and the quoted resolutions should be shown to apply to the full event sample, not a subset of well-reconstructed events. Please clarify how the residual distributions in Fig. 18 treat misassigned events.","section":"Sec. 3.6.1 / Sec. 3.6.2 (Figs. 15 and 18)"}],"minor_comments":[{"comment":"Notation errors: 'σY,1=10.5cm vs. σX,2=6.8cm' should presumably read 'σY,2=6.8cm', and 'σθY,1=14°' should likely be 'σθY,2=14°'.","section":"Sec. 3.6.2"},{"comment":"The upper light-yield threshold changes from 50 V·ns for perpendicular tracks to 100 V·ns for inclined multi-cell tracks. Please justify this difference and quantify the fraction of events removed by these cuts, especially since the final SBT efficiency requirement is >99%.","section":"Sec. 3.4.1 / Sec. 3.4.4"},{"comment":"The text attributes lower light yield in Cells 1 and 3 to imperfect optical coupling of one WOM in each cell. Since Analysis 2 uses all four cells, please state explicitly whether this known imperfection is included in the simulation and whether it affects the reported multi-cell reconstruction or only the absolute normalization.","section":"Sec. 3.4.1"},{"comment":"The y-axis label appears garbled (' X [mm]' on the ordinate); the units and quantity should be corrected.","section":"Fig. 14"},{"comment":"The 'Time-over-Threshold (ToT)' acronym is used before its full expansion; please define it at first use and specify the threshold value also in the text.","section":"Sec. 3.5.1"}],"recommendation":"major_revision","confidential_remarks":"This is a solid R&D paper with credible raw measurements and a useful prototype result. My recommendation is driven by the lack of a validation split for the likelihood-based reconstruction and the alpha=75% simulation tuning on the same dataset. If the authors add a leave-one-position/angle-out validation or demonstrate that all reconstruction parameters are fixed from [6], I would support acceptance; without that, the central resolution claims remain uncertain."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is the first full-scale, multi-cell WOM-based liquid-scintillator prototype for the SHiP SBT, and the first test-beam study of muon tracks crossing several cells with combined reconstruction. The raw measurements are credible and carefully reported. The headline resolutions — ±15.5 cm in X, ±6.8 cm in Y, ±15° in both angles — come from a likelihood reconstruction inherited from the collaboration's single-cell paper, applied to the same dataset that was also used to tune the simulation (α = 75%, Eq. 3.3). No training/validation split is described. I expect the exact numbers to shift after proper validation, but not the central claim.\n\nWhat's genuinely new: the 2×2 frustum-shaped full-scale cells; the multi-cell topology runs (e.g., θX = 90°, θY = 45° crossing three cells); and the first multi-cell track reconstruction, where including cells with no signal (Analysis 2) clearly helps, improving angular resolution at the cell center from ±17° to ±7° in θX. The timing work is also useful: position-dependent arrival-time variation is corrected to <0.4 ns, and the waveform-shape universality across cells, positions, and particle types supports the timestamp + time-over-threshold readout planned for the final SBT. The paper is honest where it matters: it shows the modest per-angle correct fractions (averages 0.52–0.69 in Fig. 15, 0.48 for position+angle in the full set), reports the bad optical coupling in two WOMs, and concludes the BaSO4 paint reached only ~75% of nominal reflectivity.\n\nSoft spots, in proportion. The load-bearing claim is the multi-cell reconstruction, and it is evaluated on the same data used for calibration. The likelihood correction is 'the same correction function' as [6], applied to all four cells with no described retraining or held-out sample; the GEANT4 response it relies on is tuned to the same test-beam data. So we can't tell whether the quoted σ's reflect in-sample tuning or an unvalidated transfer of [6] to different cell sizes; either way, no validation is shown. The fix is cheap — hold out some positions/angles — but without it, '±15 cm / ±15°' is an upper bound on capability, not verified performance. Second, the σ values are single Gaussians fit to residual distributions that visibly have multiple peaks (Fig. 16); with the low correct-fraction, resolution is not well described by one number per axis. The authors show the ingredients; the summary doesn't confront them. Third, the 99% efficiency benchmark from [3] is cited, not measured — fine as context, but easy to misread. And the summary calls the simulation agreement 'quantitative' when Section 3.4.3 frames it as qualitative with a fitted reflectivity; minor overstatement.\n\nWho this is for: the SHiP collaboration and detector R&D groups working on WOM/LS technology. It deserves a serious referee. Send it to peer review; the main thing to demand is a clear statement of whether the likelihood templates were trained on this dataset, a validation on held-out points, and a more honest summary of the residual distributions.","headline":"First full-scale multi-cell WOM-based liquid-scintillator prototype for the SHiP SBT: solid raw test-beam data and an honest write-up, but the headline reconstruction resolutions come from an in-sample likelihood evaluation — plausible, but not yet firmly established.","tokens_in":17141,"tokens_out":15706,"would_cite":true,"duration_ms":131097,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":["29.40.Mc"],"model":"deepseek-v4-flash","headline":"This paper reports that a full-scale four-cell liquid-scintillator prototype, read out by wavelength-shifting optical modules, reconstructs 5 GeV muons crossing several cells to ±15.5 cm horizontally, ±6.8 cm vertically, and ±15° in angle a","keywords":["SHiP","Surrounding Background Tagger","liquid scintillator","wavelength-shifting optical module","test beam","track reconstruction","time resolution","spatial resolution"],"falsifier":"Take an independent test-beam run at beam positions and angles not used to build the likelihood templates and evaluate the reconstruction residuals; genuine resolutions of ±15.5 cm in X, ±6.8 cm in Y, and ±15° in angle would reproduce, while template overfitting would show degraded residuals.","tokens_in":16189,"feed_emoji":"⚛️","tokens_out":5808,"duration_ms":63507,"temperature":0.7,"pith_summary":"The paper aims to show that the SHiP Surrounding Background Tagger concept—segmented liquid-scintillator cells read out by wavelength-shifting optical modules—works not just as a single cell but as a multi-cell veto that can locate and angle-track minimum-ionising muons crossing several cells. Using test-beam data from a full-scale 2×2-cell prototype, it demonstrates that a likelihood-based reconstruction recovers crossing coordinates to about ±15 cm and angles to about ±15°, and that the same correction reduces position-dependent timing variations to better than 0.4 ns. If correct, this establishes the core detector technology and reconstruction strategy for the SHiP background veto and provides a validated simulation for further optimisation.","feed_headline":"Four-cell veto prototype reconstructs muon tracks to ±15 cm","feed_subtitle":"Test-beam study shows the liquid-scintillator tagger meets SHiP's spatial and timing benchmarks.","key_machinery":"The central mechanism is a likelihood-based correction and reconstruction built from per-WOM fractional light yields, per-channel light-yield fractions of five-SiPM groups, and the difference in photon arrival times between the two WOMs of a cell. This machinery, carried over from the earlier single-cell study, turns the position- and angle-dependent detector response into an estimator of crossing coordinates and incident angles. Including cells that record no signal—Analysis 2—provides additional geometric constraints and improves angular resolution.","core_discovery":"A first multi-cell prototype of the SHiP Surrounding Background Tagger—four full-size liquid-scintillator cells, each read out by two wavelength-shifting optical modules coupled to silicon-photomultiplier arrays—was exposed to 5 GeV muons. The paper reports that a likelihood-based reconstruction, originally developed for a single cell, reconstructs muon trajectories crossing several cells with a horizontal spatial resolution of ±15.5 cm, a vertical resolution of ±6.8 cm, and angular resolutions of ±15° in both directions. The same method, using either charge fractions or time-over-threshold, equalises the position-dependent timing response to better than 0.4 ns. A detailed detector simulatio","pith_inferences":["A natural extension the paper leaves implicit: because including zero-activity cells improves angular resolution, a full SBT could use empty-cell information as a geometric constraint, potentially sharpening veto decisions without additional hardware.","The success of time-over-threshold in the timing correction suggests the final readout could rely on timestamps and time-over-threshold alone, dropping the need for full waveform digitisation and substantially reducing data volume.","The horizontal resolution (≈15.5 cm) being poorer than the vertical (≈6.8 cm) points to a testable hardware change: adding WOMs or side-mounted readout along the horizontal axis could improve X-localisation beyond the current benchmark."],"forward_implications":["The measured spatial resolutions meet the SBT's stated <20 cm benchmark, so the multi-cell concept is viable for vetoing shallow-angle muons entering from outside the decay volume.","Timing non-uniformity across a cell is corrected to below 0.4 ns, within the nanosecond-range requirement for distinguishing internal decays from external muons.","The likelihood-based reconstruction uses only detector observables such as charge fractions and arrival-time differences, so no external tracking information is needed in the final detector.","Simulation with 75% of nominal wall reflectivity reproduces the measured light-yield pattern, giving a predictive basis for optimising cell geometry, reflector materials, and WOM placement.","Including cells without signals in the reconstruction improves angular resolution, showing that hermetic multi-cell information helps constrain track angles."],"fun_headline_variants":["Four-cell SHiP veto prototype tracks muons to ±15 cm","First multi-cell veto prototype hits ±15 cm muon tracking","SHiP veto prototype localizes muons to ±15 cm in test beam","Multi-cell veto prototype nails muon positions to ±15 cm","Four-cell veto prototype reconstructs muon paths to ±15 cm"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The quoted resolutions rest on the assumption that the likelihood-based reconstruction developed for a single cell transfers to the four-cell prototype without retraining, and that the simulation tuned to a 75% wall reflectivity faithfully describes the test-beam detector.","fun_headline_variants_meta":{"raw":{"variants":["Four-cell SHiP veto prototype tracks muons to ±15 cm","First multi-cell veto prototype hits ±15 cm muon tracking","SHiP veto prototype localizes muons to ±15 cm in test beam","Multi-cell veto prototype nails muon positions to ±15 cm","Four-cell veto prototype reconstructs muon paths to ±15 cm"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001553,"raw_usage":{"total_tokens":6039,"prompt_tokens":733,"completion_tokens":5306,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":477,"completion_tokens_details":{"reasoning_tokens":5213}},"tokens_in":477,"tokens_out":5306,"duration_ms":37986,"temperature":1.0,"reasoning_tokens":5213,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-01T04:34:25.979732+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take an independent test-beam run at beam positions and angles not used to build the likelihood templates and evaluate the reconstruction residuals; genuine resolutions of ±15.5 cm in X, ±6.8 cm in Y, and ±15° in angle would reproduce, while template overfitting would show degraded residuals.","supporting_citations":[],"review_version":1}