{"id":"d31db03f-1447-4e73-9a55-944f5a90fdac","arxiv_id":"2506.06022","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"Resistance temperature detectors in liquid argon can be cross-calibrated to about 2.5 mK precision, sufficient to measure the 15 mK gradients that reveal poor mixing in large cryostats.","lead":"The paper describes a laboratory procedure for calibrating temperature probes in liquid argon and nitrogen so that pairs of probes agree to within a few thousandths of a degree. The work supports the DUNE neutrino detector, where tiny temperature differences reveal whether the liquid argon is being mixed and purified correctly.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The claimed mK-level calibration rests on an untested assumption of thermal homogeneity inside the capsule (Sec. 4.2) and rotational symmetry in the newer setup (Sec.","rationale":"The reader's weakest_assumption matches the key vulnerability: the reference-vs-tree cross-check cannot detect a capsule-wide thermal gradient, so the quoted error estimates are lower bounds on the systematic unless homogeneity is validated independently. The paper has strong positive aspects: four calibration campaigns, sub-mK repeatability, consistent mean offsets between campaigns, and no observed ageing drift. These support a conditional verdict rather than rejection. Secondary issues (the abstract's 2.5-mK statement not matching the body's 1.6-3.0 mK range, uncorrected readout offsets in the 2018 campaign, and the post-hoc choice of the averaging window) are real but would not by themselves overturn the method; they reinforce the need to phrase the headline precision carefully. A simple reversal test would settle the main systematic concern, and if it passes, the central claim becomes credible. Until then, the conditional verdict is the appropriate one.","tokens_in":15291,"tokens_out":3142,"duration_ms":33344,"concrete_test":"Perform a position-swap calibration: take a set of four sensors (or 12 corona sensors plus references) and calibrate them in the standard positions; then swap two sensors at different radii or positions (e.g., position 1 and position 2, or two corona sensors) and repeat the calibration. If the inferred offset between the two swapped sensors changes by more than the quoted repeatability (roughly 1-2 mK), a position-dependent thermal gradient is present and the calibration constants are biased. Repeat the swap for multiple pairs and for both the 2018 four-sensor capsule and the newer 14-sensor capsule; consistency under reversal would validate the homogeneity assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The calibration procedure assumes, in Sec. 4.2, that all sensors in the capsule are at the same temperature, and the post-2020 setup adds the assumption (Sec. 5.1) that convection is rotationally symmetric so the 12 corona sensors share temperature. These assumptions are load-bearing: if a few-mK spatial gradient exists inside the capsule during a calibration run, each sensor's measured offset absorbs a position-dependent bias. The internal consistency check of Sec. 4.3.5 compares reference-method and tree-method constants derived from the same capsule, so any common gradient cancels in the difference and the 2.4-mK spread under-estimates the true systematic. Repeatability distributions (mean 0.63-2.3 mK) measure random scatter, not this bias. The paper's own data show position-dependent behaviour: Fig. 10 displays different time patterns for positions 1 and 4 versus position 2, and Fig. 17 shows corona-reference offsets are larger and noisier than corona-corona offsets, indicating that geometry and convection create residual temperature differences. Without an independent test of intra-capsule uniformity, the central precision claim (2.5 mK in the abstract, 1.6-3.0 mK in the body) is not fully established.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper describes a cross-calibration technique for PT102 platinum RTDs used in the ProtoDUNE-SP temperature monitoring system. Sensors are calibrated in sets inside an aluminum capsule immersed in LAr (and later LN2), using either a reference method or a tree method to relate all sensors to a common reference. The authors report repeatabilities of 0.6–2.3 mK, apply a time-walk correction for reference-sensor fatigue, estimate calibration errors of 1.7 mK for the 2018 campaign and 1.6–3.0 mK for later campaigns, and compare four calibration campaigns to study ageing and liquid dependence. The central claim is that the calibration achieves millikelvin-level precision, stated in the abstract as 'an unprecedented precision of 2.5 mK'. The manuscript includes detailed hardware descriptions, multiple independent calibration campaigns, and internal cross-checks, but the precision claim rests on an untested assumption of thermal homogeneity inside the capsule.","tokens_in":15587,"tokens_out":2657,"duration_ms":24943,"significance":"If the claimed precision is robust, the technique is genuinely useful for DUNE and other LArTPCs: a dense grid of sensors cross-calibrated to a few mK would allow measuring the ~15 mK vertical temperature gradients predicted by CFD, providing a data-driven check of argon mixing and impurity maps without placing purity monitors in the active volume. The paper also provides useful information on RTD ageing, readout offset characterization, and the comparison of LAr versus LN2 calibration media. The work includes multiple calibration campaigns, quantitative repeatability distributions, and an explicit attempt to estimate systematic errors. However, the central precision claim is not fully supported by the presented evidence, because the dominant potential systematic—thermal gradients inside the calibration capsule—is assumed away rather than measured, and the internal cross-check between the two methods uses the same capsule and therefore cannot detect such a common bias.","major_comments":[{"comment":"The abstract claims 'an unprecedented precision of 2.5 mK', but the body does not present 2.5 mK as the central estimate. Section 4.3.5 concludes an upper limit of 1.7 mK for the tree method in 2018, and Section 5.4 estimates the single-calibration error in LAr to lie between 1.6 and 3.0 mK. The abstract's number appears nowhere in the quantitative error analysis; it should either be derived explicitly or the abstract should quote the body's actual range, otherwise the headline claim is misleading.","section":"Abstract and Sec. 4.3.5 / Sec. 5.4"},{"comment":"The calibration procedure assumes that 'all sensors in the capsule are at the same temperature' (Sec. 4.2). The paper provides no independent validation of this assumption, and its own data suggest position-dependent thermal structure: Fig. 10 shows different time patterns for positions 1 and 4 versus position 2, and Sec. 5.2 reports that corona-reference offsets are larger and noisier than corona-corona offsets. Since both the reference method and the tree method are calibrated in the same capsule with the same assumption, the cross-check of Sec. 4.3.5 (2.4 mK spread) cancels any common spatial gradient and therefore underestimates the true systematic uncertainty. The claimed 1.7 mK upper limit is thus not established.","section":"Sec. 4.2 and Sec. 4.3.5"},{"comment":"The post-2020 setup assumes rotational symmetry of convection inside the capsule so that the 12 corona sensors are at the same temperature. This assumption is load-bearing for the 1.6–3.0 mK error estimates of the 2022–2023 campaigns, yet it is only supported by the qualitative observation that corona-corona offsets are more stable than corona-reference offsets (Fig. 17). Fig. 20 shows that repeatability depends on corona position, which indicates that the assumed symmetry is not exact. Without a quantitative test of the symmetry assumption, the stated calibration error for the new setup is not fully justified.","section":"Sec. 5.1 and Sec. 5.2"},{"comment":"The readout offset correction is a key systematic: Sec. 3 shows channel offsets up to 2.5 mK and states that this correction was applied only in the post-2018 campaigns. The 2018 calibration therefore contains an uncorrected readout contribution, which is acknowledged in the LAr-2018 vs LAr-2023 comparison (std dev 4.3 mK). This means that the 2018 value of 1.7 mK is an underestimate of the total error, and the abstract's 2.5 mK cannot be taken as the precision achieved in 2018. The paper should explicitly state which error estimate, if any, corresponds to the abstract claim and whether the 2.5 mK includes the readout offset correction.","section":"Sec. 3 and Sec. 5.4"}],"minor_comments":[{"comment":"There is a typo: 'thse two measurements' should be 'these two measurements'.","section":"Sec. 3"},{"comment":"The word 'Resitance' is misspelled; it should be 'Resistance'.","section":"Sec. 2"},{"comment":"The phrase 'sligthly worst' should be 'slightly worse'.","section":"Sec. 5.3"},{"comment":"The axis label 'Inmersions' is a typo for 'Immersions', and the unit 'mk' should be 'mK' consistently throughout the figures.","section":"Fig. 12 and Sec. 4.3.2"},{"comment":"The sentence about the calibration run 'actually begins at that point and lasts for 40 minutes' is clear, but it would help to state explicitly how the 1000–2000 s stable interval (Sec. 4.3.1) relates to the 40-minute run, since the run start time is defined relative to capsule immersion.","section":"Sec. 4.2"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a solid engineering/metrology paper with real new data, but the headline precision number should be read as a self-consistency estimate, not an accuracy bound. The body is more careful than the abstract.\n\nWhat's new: four calibration campaigns over five years on the same set of PT102 RTDs, a time-walk ageing model for the reference sensors, the upgrade from 4 to 14 sensors in the capsule, and a LAr/LN2 comparison showing no large offset shift from the 10 K change. The readout offset measurement (Fig. 3) is a good piece of work; it explains why earlier campaigns were off and lets them correct later ones. The two independent calibration trees (reference vs. tree) and the 16-sensor cross-check in Sec. 4.3.5 are, as an internal consistency check, well designed. The ageing conclusion (no drift above the few-mK level over five years) is useful for DUNE.\n\nSoft spots: the abstract's \"unprecedented precision of 2.5 mK\" does not match the body's central estimate, which is 1.6–3.0 mK for the calibration error. The 2.5 mK number seems to come from the readout offset magnitude, not the calibration error. The stress-test concern is legitimate: the Sec. 4.2 assumption that all sensors in the capsule share a common temperature is untested, and the newer setup's rotational symmetry assumption is also untested. Because both calibration methods use the same capsule, any common-mode thermal gradient cancels in the Sec. 4.3.5 cross-check; the 2.4 mK spread therefore measures random scatter, not this systematic. Fig. 10 and Fig. 17 do show position-dependent behaviour that suggests small gradients inside the capsule exist. The 2018 campaign also lacks the readout-offset correction that later turned out to be up to 2.5 mK, so that campaign's error budget is incomplete.\n\nNone of this sinks the main conclusion. For the intended application—checking 15 mK gradients in ProtoDUNE—a calibration error of a few mK either way is tolerable. But the authors should either measure the intra-capsule gradient directly (e.g., swap sensor positions and see whether the offsets follow) or soften the precision claim. Raw data, or at least a table of calibration constants with uncertainties, would also make the numeric estimates independently checkable.\n\nWho it's for: anyone working on LArTPC cryogenics, precision temperature monitoring in large cryostats, or RTD calibration practice. It deserves serious peer review; the issues are addressable and the experimental work is valuable. I'd engage with it.","headline":"Solid metrology paper with genuinely useful long-term calibration data, but the mK precision claim should be read as an internal-consistency estimate until the thermal-homogeneity assumption inside the capsule is tested.","tokens_in":16161,"tokens_out":2366,"would_cite":true,"duration_ms":22463,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a laboratory cross-calibration procedure brings platinum resistance thermometers to millikelvin precision in liquid argon, enough to measure the 15 mK gradients that reveal mixing and purity in large cryogenic…","keywords":["temperature sensing","cryogenic detectors","liquid argon","resistance temperature detectors","RTD cross-calibration","temperature gradients","argon purity","ProtoDUNE-SP"],"falsifier":"Take a set of sensors, hold them in the calibration capsule, and impose a controlled asymmetric temperature difference across the capsule by warming one wall with a small heater; if the derived offsets between sensors shift systematically with position or with the applied heat, the isothermality premise fails and the quoted calibration errors are underestimates.","tokens_in":15065,"feed_emoji":"🌡️","tokens_out":7662,"duration_ms":68681,"temperature":0.7,"pith_summary":"This paper claims that resistance temperature sensors can be cross-calibrated against one another in liquid argon to a precision of about 2.5 mK, with a best-estimate calibration error of 1.7 mK for the 2018 campaign and 1.6 to 3.0 mK for later campaigns, using a nested-insulation laboratory capsule rather than expensive individual sensor calibration. The reason this matters is that the ProtoDUNE-SP and DUNE liquid-argon detectors need to measure vertical temperature gradients of about 15 mK predicted by computational fluid dynamics; those gradients are the experimental signature of whether purified argon is mixing uniformly through the cryostat. A dense grid of sensors calibrated this way can therefore act as a stand-in for purity monitors, which cannot be placed inside the active detector volume. The paper argues that the achieved precision is well below the 5 mK requirement for DUNE and that the calibration survives translation from liquid nitrogen to liquid argon and five years of sensor ageing.","feed_headline":"RTD cross-calibration reaches 2.5 mK for argon detectors","feed_subtitle":"A lab capsule lets a dense sensor grid resolve 15 mK gradients that reveal argon purity and mixing.","key_machinery":"The load-bearing object is the calibration capsule: a thin-walled aluminum cylinder suspended inside nested insulating volumes, assumed isothermal during a run because all sensors in it are supposed to share one temperature. On top of that assumption the paper introduces two cross-calibration topologies: the reference method, in which every batch includes one common reference sensor, and the tree method, in which a few promoted sensors link batches in later rounds; the tree method wins because it needs no time-walk correction. The time-walk correction itself is a second piece of machinery, a linear parametrization of how the reference sensor's offset drifts with the number of thermal immersions, caused by thermal fatigue. The third component is the readout: a shared 1 mA current source and a multiplexed single 24-bit ADC channel with four-wire sensing, which suppresses electronic offsets between channels to below about 0.5 mK repeatability.","core_discovery":"The central claim is that relative sensor offsets, not absolute temperatures, are what need to be calibrated, and that this can be done to millikelvin level by exposing batches of sensors to the same cryogenic bath. The authors built a calibration capsule with multiple insulating volumes and an aluminum inner vessel to make the thermal environment as homogeneous as possible, then cross-calibrated 48 sensors by two routes: a reference method that always compares against one sensor, and a tree method that chains offsets through promoted sensors. Although the tree method involves more intermediate steps, it avoids the drift of a single heavily cycled reference and turns out to be the more accurate route; after correcting the primary reference for a time-walk of 0.07 mK per immersion, growing to 0.22 mK per immersion, the two independent routes agree with a spread of 2.4 mK, giving an upper limit of 1.7 mK on the tree-method error. Re-calibrations after detector decommissioning, with an enlarged 14-sensor capsule, give a single-calibration error between 1.6 and 3.0 mK, no bias between liquid nitrogen and liquid argon, and no detectable ageing over five years.","pith_inferences":["If the capsule is not truly isothermal, both calibration routes share the same bias, because both use the same capsule; a direct test would be to impose a known asymmetric heat load and see whether derived offsets move with sensor position.","The calibration removes relative offsets, not absolute accuracy; vendor-level absolute accuracy of about 0.1 K remains, but for gradient and mixing diagnostics the relative precision is the physically relevant quantity.","The same cross-calibration logic should transfer to other cryogenic liquids or systems where mixing and purity correlate with temperature, such as liquid-hydrogen or liquid-oxygen targets.","The 4.3 mK spread between the 2018 and 2023 liquid-argon campaigns suggests that long-term reproducibility may be limited by uncorrected readout-channel offsets as much as by sensor ageing; future campaigns could isolate this by using identical readout channels for both measurements."],"forward_implications":["A grid of sensors calibrated this way can resolve the CFD-predicted 15 mK gradients in ProtoDUNE-SP, giving a data-driven check of argon mixing and purity.","The calibration error (1.7 mK for the 2018 tree method, 1.6 to 3.0 mK for later campaigns) is less than half the 5 mK precision required for DUNE far-detector monitoring.","Liquid nitrogen can be used for large-scale calibration instead of liquid argon: the 10 K difference in bath temperature does not bias the offsets.","The absence of measurable ageing over five years means a single laboratory calibration can remain valid across the lifetime of a detector module.","The scaled-up 14-sensor capsule makes it practical to calibrate the more than 500 sensors planned for the DUNE far detector."],"supporting_citations":[{"why":"It supplies the ProtoDUNE-SP context and the CFD prediction of about 15 mK vertical temperature gradients that the calibrated sensor grid must resolve.","marker":"[8]"},{"why":"It sets the DUNE far-detector requirements, including the more than 500 sensors and the precision target that motivate scaling up the calibration.","marker":"[9]"},{"why":"It provides prior 35-tonne LArTPC experience that guided the choice of PT102 sensors and cryogenic temperature-monitoring practice.","marker":"[14]"},{"why":"It defines the PT102 platinum RTD sensors whose 0.1 K vendor accuracy is the baseline the cross-calibration must improve on.","marker":"[17]"},{"why":"It supplies the four-wire resistance-thermometry method used to remove cable and connector resistance from the temperature readings.","marker":"[16]"},{"why":"It supplies the multiplexed, shared-current-source readout architecture that suppresses electronic offsets between channels.","marker":"[22]"},{"why":"It documents the static Temperature Gradient Monitor whose 48 sensors are the subject of the 2018 calibration campaign.","marker":"[21]"},{"why":"It describes the purity monitors placed outside the active volume, the limitation that motivates using a temperature grid as a proxy for purity.","marker":"[12]"}],"fun_headline_variants":["RTD cross-calibration reaches 2.5 mK for argon detectors","Cross-calibrated sensors give millikelvin precision in argon","2.5 mK temperature accuracy from sensor cross-calibration","Cryogenic sensor cross-calibration achieves 2.5 mK","Millikelvin sensing via cross-calibration in cryogenic liquids"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"During every calibration run, all sensors inside the capsule are at exactly the same temperature; if a spatial gradient exists inside the capsule, every measured offset is biased, and the internal cross-check between the two methods cannot reveal it because both methods use the same capsule.","fun_headline_variants_meta":{"raw":{"variants":["RTD cross-calibration reaches 2.5 mK for argon detectors","Cross-calibrated sensors give millikelvin precision in argon","2.5 mK temperature accuracy from sensor cross-calibration","Cryogenic sensor cross-calibration achieves 2.5 mK","Millikelvin sensing via cross-calibration in cryogenic liquids"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000247,"raw_usage":{"total_tokens":1525,"prompt_tokens":911,"completion_tokens":614,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":527,"completion_tokens_details":{"reasoning_tokens":523}},"tokens_in":527,"tokens_out":614,"duration_ms":5830,"temperature":1.0,"reasoning_tokens":523,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T06:02:42.059712+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of sensors, hold them in the calibration capsule, and impose a controlled asymmetric temperature difference across the capsule by warming one wall with a small heater; if the derived offsets between sensors shift systematically with position or with the applied heat, the isothermality premise fails and the quoted calibration errors are underestimates.","supporting_citations":[{"cited_title":"URLhttps://www.lakeshore.com/products/categories/overview/ temperature-products/cryogenic-temperature-sensors/platinum","cited_arxiv_id":null,"evidence_quote":"It defines the PT102 platinum RTD sensors whose 0.1 K vendor accuracy is the baseline the cross-calibration must improve on."},{"cited_title":"URLhttps://www.minco.com/wp-content/uploads/Resistance-Thermometry","cited_arxiv_id":null,"evidence_quote":"It supplies the four-wire resistance-thermometry method used to remove cable and connector resistance from the temperature readings."},{"cited_title":"Chojnacki, A Multiplexed RTD Temperature Map System for Multi-Cell SRF Cavities, in: SRF2009, Vol","cited_arxiv_id":null,"evidence_quote":"It supplies the multiplexed, shared-current-source readout architecture that suppresses electronic offsets between channels."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"It documents the static Temperature Gradient Monitor whose 48 sensors are the subject of the 2018 calibration campaign."},{"cited_title":"Xiao, Purity monitoring for ProtoDUNE-SP, J","cited_arxiv_id":null,"evidence_quote":"It describes the purity monitors placed outside the active volume, the limitation that motivates using a temperature grid as a proxy for purity."}],"review_version":1}