{"id":"caa81c7e-e8d3-48f6-b587-95413e4a5529","arxiv_id":"2607.13118","paper_version":1,"verdict":"ACCEPT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A digitized 198Au decay curve reproduces the published half-life under a no-offset fit, but uniform baseline offsets of ±0.05 cps shift the fitted half-life by up to 0.054 d.","lead":"This paper digitizes a published gold-198 decay plot and shows that a standard exponential fit reconstructs the reported half-life, while small artificial baseline shifts can move the result by more than its statistical error. It demonstrates how figure-level data can be stress-tested for reproducibility and sensitivity when raw spectra are unavailable.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Manual digitization may be confirmation-biased; independent blinded re-digitization of Fig. 2 is needed to validate the reproduced half-life.","rationale":"The reader's ACCEPT depends on the assumption that the digitized dataset faithfully represents the printed figure. That assumption is the one I would test first: an independent, blinded re-digitization would settle whether the headline reproduction is real. The paper's jitter test does not cover systematic selection bias. I also noticed a smaller, internal issue: Sec. 4.2 computes s = 1.201 but Table 1 and Sec. 4.7 quote the unscaled covariance error ±1477 s; applying the paper's own scale rule gives ±1774 s (±0.0205 d). This widens the interval but still overlaps the published value; it should be corrected in the final reporting. Neither issue invalidates the workflow, but the unverified digitization neutrality warrants conditional acceptance until the check is done.","tokens_in":30745,"tokens_out":12058,"duration_ms":127409,"concrete_test":"Have a second analyst, blinded to the published A(0), T1/2, and to the author's extracted values, independently digitize the same published figure using a different tool (e.g., Engauge Digitizer) and the same linear-axis calibration. Refit Eq. (1) to the independent dataset. If the independent T1/2 lies within ~0.01 d of 2.6687 d and A0 reproduces the published 3.68 ± 0.04 cps, the concern is resolved. If the independent fit differs by more than the combined statistical error or moves A0 by more than ~0.05 cps, the original point selection or axis calibration was not neutral and the headline claim must be re-evaluated.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central numerical claim—that the no-offset fit reproduces the Spillane et al. value to 0.0003 d (Sec. 4.1, Sec. 4.7)—is only meaningful if the digitized points are a neutral, unbiased sample of the printed figure. Digitization was manual (Sec. 2), and the analyst knew the published A(0) and T1/2. The point-picking jitter test (Sec. 4.7) perturbs the already-selected points; it cannot detect a systematic selection toward the known fitted curve, shared axis-calibration error, or a consistent vertical bias. The fitted A0 = 3.7333 ± 0.0103 cps is 1.3σ from the published 3.68 ± 0.04 cps, consistent but suggesting the vertical scale may not be perfectly reconstructed. Because every downstream diagnostic (offset scans, profile likelihood, MFV, toy MC) uses the same extracted dataset, a systematic extraction bias would make the headline 'reproduction' an artifact of expectation rather than evidence that the figure retains the decay scale. The paper acknowledges digitization limitations in Sec. 2 but does not validate this specific extraction against an independent digitization.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a reproducible reduced-data workflow for testing half-life estimates from digitized radiation-sensor decay plots, using the 198Au room-temperature curve of Spillane et al. as a case study. The author digitizes 78 count-rate points with error bars from the published figure, fits a weighted single-exponential no-offset model, and obtains T1/2 = (2.6687 ± 0.0171) d and A0 = (3.7333 ± 0.0103) cps, closely reproducing the published values T1/2 = (2.669 ± 0.017) d and A(0) = (3.68 ± 0.04) cps. The paper then applies a battery of diagnostics: profile likelihoods, uniform count-rate offsets, time-origin shifts, time-scale distortions, window/truncation and leave-one-out tests, pairwise-ratio and MFV summaries, MDR smoothing, an FFT residual-baseline check, and toy Monte Carlo controls. The main conclusion is that figure-level data can preserve the half-life scale sufficiently for a regression check, but that baseline-like offsets are the dominant figure-level sensitivity and should be reported as a separate diagnostic scale, not as a calibrated systematic uncertainty.","tokens_in":31087,"tokens_out":7889,"duration_ms":82252,"significance":"If the digitized data are unbiased, the paper provides a useful and unusually transparent case study in reduced-data analysis: it separates the primary fit from statistical uncertainty, figure-level sensitivity, and robustness diagnostics, and it makes data and scripts available on OSF. The external validation of the MFV implementation against the neutron-lifetime benchmark and the toy-MC controls that separate finite-window estimator behavior from digitization artifacts are genuine strengths. However, the central claim — that the digitized dataset reproduces the published half-life — rests on the accuracy of manual digitization, which is not independently validated. The fitted A0 excess and the lack of reported Δχ² values for the signed-offset profile are specific weaknesses that need to be addressed before the reproduction claim can be taken as established.","major_comments":[{"comment":"The load-bearing claim that the no-offset fit reproduces the published half-life is not validated against an independent digitization or a synthetic-figure calibration. The point-picking jitter test in Sec. 4.7 perturbs the already-selected points and cannot detect a systematic selection toward the known published curve, a shared axis-calibration error, or a uniform vertical bias. The fitted A0 = 3.7333 ± 0.0103 cps is 0.053 cps above the published A(0) = 3.68 ± 0.04 cps — about five times the fit standard error and roughly 1.3 combined standard deviations — and this difference is comparable to the ±0.05 cps uniform offset that changes T1/2 by 0.054 d. The paper should either test for a vertical extraction bias of this size, show why such a bias would not affect the half-life, or provide an independent blinded re-digitization / synthetic-figure calibration before asserting that the half-","section":"Sec. 2, Sec. 4.1, Sec. 4.7"},{"comment":"The signed-offset profile is used to support a non-identifiability conclusion, but the paper reports no Δχ² values for the profile minima at T1/2 ≈ 2.82–2.85 d. Without these values, the reader cannot distinguish a shallow valley that supports the 'diagnostic only' interpretation from a statistically significant three-parameter alternative. This matters because the reported shift (0.181 d) is larger than the published room-temperature vs 12 K difference (0.096 d). Please report Δχ²(T1/2 = 2.85 d) relative to the no-offset fit, and if the valley is broad, give the profile width at the usual thresholds (e.g., Δχ² = 1 or 2.71). Only then can the claim that the offset model is non-identifiable rather than preferred be evaluated.","section":"Sec. 4.6"},{"comment":"The manuscript states in Sec. 3.2 that when χ²/ndf > 1 the reported uncertainty should be scaled by s = sqrt(χ²/ndf), and Sec. 4.2 reports s = 1.201. However, the final statistical uncertainty in Sec. 4.7 and Sec. 6 is the unscaled value ±0.0171 d (and Table 1 reports the covariance-matrix standard error as ±0.0171 d). With s = 1.201, the scaled statistical uncertainty would be ±0.0205 d. The paper should either apply the scale factor consistently throughout the reporting hierarchy or explicitly state that all 'stat' entries are unscaled and that the scale factor is provided only as a diagnostic. As written, the internal inconsistency undercuts the paper's own stated reporting rule.","section":"Sec. 3.2, Sec. 4.7, Sec. 6"}],"minor_comments":[{"comment":"The phrase 'the uncertainties shown in the published plot were described as statistically significant' should presumably read 'statistical uncertainties'; please correct the wording.","section":"Sec. 2"},{"comment":"The row 'Constant uncertainty check, σ = median(σ_i)' appears in Table 1 but is not explained in the text of Sec. 4.5. Please describe this test in the main text and state what it is designed to probe.","section":"Table 1 / Sec. 4.5"},{"comment":"The percentile-bootstrap intervals for the pairwise median and MFV are obtained by resampling the 78 original points, then recomputing the pairwise distribution. Because pairwise ratios share points, these intervals are only approximate diagnostic intervals. The text acknowledges the shared-point structure, but the distinction between resampling points and resampling independent pairwise ratios should be stated more explicitly in the method description.","section":"Sec. 4.8"},{"comment":"The MDR fits report χ²/ndf values from 0.051 to 1.586; the text correctly notes that the MDR points are correlated and that the formal fit uncertainties are not meaningful. This is appropriate, but it would be clearer to state up front that the MDR χ²/ndf values are listed only as descriptive quantities, not as goodness-of-fit statistics.","section":"Appendix A"}],"recommendation":"major_revision","confidential_remarks":"The paper is a borderline accept after revision. The central claim — reproduction of the published half-life from digitized data — is plausible and the workflow is transparent, but it depends on the accuracy of a single manual digitization. The A0 discrepancy (0.053 cps above the published value) is, in my reading, the strongest internal red flag for a possible vertical extraction bias; it is close to the uniform-offset range that shifts T1/2 by 0.054 d. The paper should be asked to provide an independent digitization check or to substantially weaken the reproduction claim. The missing Δχ² values in Sec. 4.6 are also an important oversight, because they determine whether the offset model is merely poorly constrained or actually preferred. I would not reject, but I would require these points to be addressed before publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plainly: this paper is a methodological case study, not a new measurement. It digitizes one published 198Au decay curve, fits it with a weighted no-offset exponential, and reproduces the published half-life (2.6687 ± 0.0171 d vs 2.669 ± 0.017 d). The genuinely useful contribution is the reporting hierarchy: primary regression estimate, statistical uncertainty, figure-level sensitivity scale, and robustness diagnostics are computed together but kept separate. That is a good habit for anyone doing reduced-data QA.\n\nCredit where due: the internal consistency is strong. Profile-likelihood interval matches the covariance estimate; sensitivity scans are labeled as diagnostics rather than systematic uncertainties; the toy MC is used as a finite-window control, not as proof of the digitization; and the OSF repository appears to ship data, scripts, and logs. The MFV implementation is benchmarked against the neutron-lifetime dataset, so the robustness software is not black-box.\n\nSoft spots, in order. The biggest is the manual digitization. The analyst knew the target A(0) and T1/2 before picking points. The jitter test perturbs already-selected points and cannot catch a systematic selection toward the known curve, a shared axis-calibration error, or a consistent vertical bias. The extracted A0 is 1.3σ from the published 3.68 ± 0.04 cps, which is fine statistically but does not remove the worry. This does not sink the paper because the claim is a demonstration, not a new half-life, but an independent blinded re-digitization would materially raise confidence.\n\nSecond, the χ²/ndf = 1.444 mismatch between the fitted error bars and the reported χ²ν = 1.06 is never fully explained; it is plausible digitization extra scatter, but the paper does not resolve it. Minor.\n\nThird, the toy MC using the no-offset fit as truth means the 'expected' pairwise shifts are only about that model; that's fine as a control, and the paper says so.\n\nThe stress-test note overstates the threat. The reproduction could be expectation-driven, but the paper's own framing and the separate diagnostics (offset profiles, profile likelihood, pairwise/MFV) do not depend on the exact central value for their message. The central takeaway — small baseline-like offsets can move a figure-only half-life by more than its statistical uncertainty — holds even if the digitization is imperfect.\n\nI'd send this to peer review. It is a credible, reproducible methods paper for an audience doing legacy-data QA, figure reanalysis, or detector time-series work. The main revision should be a blinded re-digitization, or at least a clear acknowledgment that one digitizer's agreement is not evidence of unbiased extraction.","headline":"A careful, honest figure-level QA workflow whose real novelty is the reporting hierarchy; the digitized-data reproduction of the 198Au half-life is plausible but would be stronger with a blinded repeat digitization.","tokens_in":31516,"tokens_out":2396,"would_cite":true,"duration_ms":25381,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A weighted exponential fit to digitized points from a published 198Au decay plot reproduces the reported half-life, 2.6687 ± 0.0171 d, and shows that small baseline offsets—not statistical scatter—dominate figure-level reanalysis uncertaint","keywords":["radioactive decay","half-life estimation","198Au","plot digitization","profile likelihood","baseline sensitivity","exponential fitting","most frequent value"],"falsifier":"Obtain the original count-rate table for the same room-temperature decay curve and repeat the no-offset and signed-offset fits on the raw values. If the raw-data offset profile is narrow and a ±0.05 cps offset moves the half-life by much less than 0.054 d, the digitization-based sensitivity envelope overstates the true baseline sensitivity. A simpler independent check: have several operators digitize the same published figure and compare the spread in extracted low-rate points; if that spread exceeds roughly 0.03 cps, the reproduced half-life can move outside the quoted statistical band.","tokens_in":30666,"feed_emoji":"⏳","tokens_out":7045,"duration_ms":64703,"temperature":0.7,"pith_summary":"The paper asks how much of a published half-life measurement survives when only the plotted decay curve is available. It digitizes the room-temperature 198Au decay figure from an earlier experiment and fits a no-offset exponential, recovering T1/2 = 2.6687 ± 0.0171 d against the reported 2.669 ± 0.017 d. The main finding is that the fit is locally well constrained but globally sensitive to additive baseline offsets: a uniform ±0.05 cps shift moves the half-life by up to 0.054 d, roughly three times the statistical error. The paper packages this as a reporting hierarchy—primary estimate, statistical uncertainty, figure-level sensitivity scale, and diagnostic checks kept separate—so that secondary analysts do not overstate what plot-level data can certify. It matters because many legacy and modern experimental datasets are available only as reduced plots, and reproducibility tests need to know which parts of an inference are actually identifiable.","feed_headline":"Digitized decay plot yields 198Au half-life","feed_subtitle":"Baseline offsets, not statistical scatter, dominate figure-level reanalysis uncertainty","key_machinery":"The central device is the signed residual-offset exponential model A(t) = A0 exp(−ln 2 · t / T1/2) + Boff, where Boff is allowed to be positive or negative. At fixed T1/2 the model is linear in (A0, Boff), so the nuisance offset can be profiled out by weighted least squares without iterative nonlinear refits, yielding a smooth profile in T1/2 that exposes the offset–lifetime trade-off. Supporting machinery includes profile-likelihood scans that refit the remaining parameter at each fixed T1/2, pairwise count-rate ratios that cancel the normalization, a most-frequent-value robust summary for heavy-tailed lifetime distributions, and toy Monte Carlo controls that separate estimator behavior fro","core_discovery":"The central result is that the digitized 198Au room-temperature decay curve retains the decay scale of the original experiment: the weighted no-offset fit gives T1/2 = (2.6687 ± 0.0171) d, matching the published individual-curve value (2.669 ± 0.017) d, with a profile-likelihood interval of [2.6546, 2.6830] d and a strong A0–T1/2 correlation of −0.831. The same data, however, cannot separate a small constant residual offset from a change in the decay constant over the 3.2-day window. A uniform offset of ±0.05 cps changes the fitted half-life by up to 0.0540 d, and an unconstrained signed-offset profile moves the minimum to roughly 2.82–2.85 d, comparable to the original room-temperature vers","pith_inferences":["Inference: the measured baseline sensitivity is large enough that any reported systematic uncertainty smaller than about 0.05 d for this dataset cannot be validated from the figure alone; only raw-data access could certify such precision.","Inference: the offset–lifetime trade-off suggests a concrete test for the historical temperature-dependence question: if the original spectra were reanalyzed with an explicit constant-baseline nuisance, part of the apparent room-temperature versus low-temperature difference might be absorbable by baseline shifts.","Inference: analysts applying pairwise or most-frequent-value summaries to short-window legacy plots should run matching toy controls; without them, downward shifts in robust estimators could be mistaken for genuine half-life differences.","Inference: the same diagnostic hierarchy could be calibrated into a decision rule—if the offset-scan envelope exceeds the claimed total uncertainty of a published value, the plot-level data cannot support that precision claim."],"forward_implications":["Figure-level digitized data can reproduce a published half-life central value within its quoted statistical uncertainty, at least for well-resolved single-exponential curves.","Additive baseline offsets are the dominant analysis-level perturbation for figure-only reconstruction: a ±0.05 cps uniform shift moves the half-life by about 2%, so such reconstructions should carry a separate, non-statistical sensitivity scale.","Finite-window exponential data produce systematic estimator shifts—lower pairwise and most-frequent-value central values and broad offset profiles—that are expected and should not be read as evidence for a different physical half-life.","Lengthening the observation window reduces normalization–lifetime and offset–lifetime degeneracy, which means window length is a first-order design consideration for any half-life reanalysis.","The same reporting hierarchy can be applied to other radionuclides with published decay plots, such as argon-39, to test which parts of a half-life claim are reproducible from reduced data alone."],"fun_headline_variants":["Offsets shift 198Au half-life, not scatter","Baseline offsets, not scatter, shift 198Au half-life","198Au decay: baseline offsets dominate reanalysis","Digitized decay plot: half-life stable, but offsets sway","Reanalyzed 198Au decay data: offsets, not scatter, matter"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The entire reconstruction hangs on the digitized point coordinates and displayed error bars faithfully representing the plotted data; if the manual axis calibration or point placement is systematically off, especially near the low-count-rate end, the recovered half-life and the offset sensitivity both change.","fun_headline_variants_meta":{"raw":{"variants":["Offsets shift 198Au half-life, not scatter","Baseline offsets, not scatter, shift 198Au half-life","198Au decay: baseline offsets dominate reanalysis","Digitized decay plot: half-life stable, but offsets sway","Reanalyzed 198Au decay data: offsets, not scatter, matter"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001133,"raw_usage":{"total_tokens":4577,"prompt_tokens":813,"completion_tokens":3764,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":3678}},"tokens_in":557,"tokens_out":3764,"duration_ms":24497,"temperature":1.0,"reasoning_tokens":3678,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T06:12:51.633661+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Obtain the original count-rate table for the same room-temperature decay curve and repeat the no-offset and signed-offset fits on the raw values. If the raw-data offset profile is narrow and a ±0.05 cps offset moves the half-life by much less than 0.054 d, the digitization-based sensitivity envelope overstates the true baseline sensitivity. A simpler independent check: have several operators digitize the same published figure and compare the spread in extracted low-rate points; if that spread exceeds roughly 0.03 cps, the reproduced half-life can move outside the quoted statistical band.","supporting_citations":[],"review_version":1}