REVIEW 3 major objections 4 minor
Numerical Uncertainty in Linear Registration: An Experimental Study
T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read The paper claims that Monte-Carlo Arithmetic perturbation reveals stable and unstable linear registration tools, with SPM the most stable.
desk verdict A useful practical comparison of linear registration stability under Monte-Carlo Arithmetic, but the ranking's validity hinges on a perturbation model the abstract never calibrates. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Monte-Carlo Arithmetic (MCA) is the mechanism: instead of a single deterministic run, each floating-point operation's low-order bits are randomly perturbed, and the registration is repeated many times. The spread of the resulting transformation parameters estimates the numerical uncertainty of the tool and similarity measure. This perturbation-based dispersion is what the paper uses to compare tools and to demonstrate a quality-control signal.
What would settle it
Run the same SPM, FSL, and ANTs linear registrations on real hardware while varying compiler flags, rounding modes, or CPU architectures that change floating-point evaluation order, and check whether ANTs produces occasional failures and whether the same stability ranking appears. If the ranking and failure rate do not reproduce under genuine hardware-dependent rounding, the Monte-Carlo Arithmetic model is not faithful enough to transfer.
Extended reading notes
Core claim
The paper reports that, across 50 healthy and 50 Parkinson's disease subjects, two templates, and several similarity measures, Monte-Carlo Arithmetic perturbations of floating-point operations produce measurable output variability in all three major linear registration tools. With default similarity measures, SPM shows the smallest dispersion, FSL and ANTs show larger and similar dispersion, and ANTs occasionally fails entirely under perturbation. The study finds no significant difference in numerical stability between healthy and PD cohorts, and it shows that dispersion measures can flag registrations whose results should not be trusted.
Load-bearing premise
The whole ranking depends on the assumption that perturbing low-order arithmetic bits in a simulation faithfully reproduces the true numerical errors these tools make on real MRI scans.
Editorial extensions
If this is right
- Users of SPM's default settings can expect more reproducible linear registration results under numerical noise.
- ANTs users should treat registration outputs as potentially unstable and should consider reruns or alternative tools when perturbation tests show large dispersion.
- Automated QC pipelines could flag registrations with high numerical uncertainty without waiting for downstream artifacts.
- Numerical stability findings from healthy cohorts can inform clinical studies of Parkinson's disease, reducing the need to repeat such analyses on every patient group.
- Choice of similarity measure interacts with tool choice for numerical stability, so reproducible pipelines should pin both.
Reading between the lines
- A practical extension would be to convert MCA dispersion into a per-subject QC score with a threshold tuned on real failing registrations.
- The ANTs failures suggest a bifurcation in its optimization path under rounding; reproducing that bifurcation with reduced precision would identify the exact arithmetic step responsible.
- If the healthy-to-clinical generalization holds beyond PD, numerical stability audits could be run on open healthy datasets and safely reused for many clinical studies.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This manuscript reports an experimental study of numerical uncertainty in linear registration tools (SPM, FSL, ANTs) using Monte-Carlo Arithmetic. Based on the abstract, the authors claim that, with default similarity measures, SPM is the most stable, FSL and ANTs show greater and comparable variability with ANTs occasionally failing, that healthy controls (HC) and Parkinson's disease (PD) cohorts show no significant differences in numerical stability, and that numerical uncertainty measures can support automated quality control (QC) of linear registration. The study uses two brain templates, multiple similarity measures, and n=50 per cohort.
Significance. If the results are reliable, the paper would provide a useful empirical ranking of numerical robustness among widely used registration tools and a practical QC indicator. The use of a uniform external perturbation (MCA) across tools is a reasonable design choice, and the inclusion of a clinical cohort is a strength. The presented abstract alone, however, does not disclose the MCA calibration, the statistical power of the HC/PD comparison, or the validation of the proposed QC signal, so the significance currently depends on assumptions about the full text.
major comments (3)
- [Abstract, first sentence] The central claim of tool-specific stability is conditioned on the fidelity of Monte-Carlo Arithmetic. The abstract does not report the perturbation magnitude (e.g., fraction of mantissa bits perturbed), the number of Monte-Carlo runs per registration, or how the perturbation level was calibrated or validated against real hardware variability. Without this information, the SPM/FSL/ANTs ranking and the ANTs failure rate could be artifacts of the MCA hyperparameter rather than intrinsic numerical sensitivity.
- [Abstract, HC/PD comparison sentence] The generalization claim ('no significant differences were observed between healthy and PD cohorts') is a null result. The abstract reports neither effect sizes nor confidence intervals nor a power analysis for this comparison. With n=50 per cohort, the null could simply reflect insufficient statistical power, making the suggested generalization to clinical populations unsupported.
- [Abstract, final sentence] The QC demonstration is asserted without any reported validation metrics. The abstract says numerical uncertainty measures 'may support' automated QC but gives no sensitivity, specificity, or comparison to existing QC methods, so the claim is not substantiated at the level of the presented evidence.
minor comments (4)
- [Abstract, SPM/FSL/ANTs ranking] The phrase 'greater and similar ranges of variability' is ambiguous; please report the actual dispersion values (e.g., interquartile ranges or variance) for FSL and ANTs to support the comparison.
- [Abstract, methodology] The abstract does not state software versions, computing platform, or compiler/BLAS configuration, which are critical for a numerical reproducibility study.
- [Abstract, HC/PD comparison] The sentence 'no significant differences were observed' should be accompanied by a clear statement that this is a null result with its confidence interval, not evidence of equivalence.
- [Abstract, similarity measures] The term 'default similarity measures' is vague; please specify which measures were used for each tool in the abstract or clearly point to the full-text listing.
Circularity Check
No circularity found: the abstract reports an external perturbation experiment whose ranking and QC demonstration are not forced by construction from the inputs.
full rationale
This is an abstract-only review, and no equation, derivation, or fitted-parameter construction is available that would allow exhibiting a specific reduction of a claimed result to its own inputs. The central stability ranking is obtained by applying Monte-Carlo Arithmetic perturbation externally and uniformly to SPM, FSL, and ANTs; the outcome is not encoded in the perturbation model's definition. The HC/PD comparison uses an external clinical dataset and reports no significant differences, which is a contingent empirical finding rather than an analytic consequence of the methods. The QC demonstration is mentioned only as a demonstration, and without the methodology we cannot identify a success criterion defined from the same dispersion that defines the uncertainty measure; to flag circularity here would require quoting the specific construction, which the abstract does not provide. The load-bearing assumption that MCA perturbation faithfully reproduces true numerical uncertainty is a modeling assumption about fidelity and transferability, not a circularity: it does not make the ranking true by definition, and concerns about it belong to correctness risk rather than to the circularity score. There are no self-citations, no imported uniqueness theorems, and no renamed known results evident from the abstract. Accordingly, the honest finding is no significant circularity, score 0.
Assumptions & free parameters
free parameters (2)
- MCA perturbation magnitude (fraction of low-order mantissa bits perturbed)
- Number of Monte-Carlo runs per registration
assumptions (3)
- domain assumption Monte-Carlo Arithmetic perturbations faithfully simulate the true numerical uncertainty of the registration tools.
- domain assumption The two brain templates and the healthy/PD cohorts are representative enough to support the generalization claim.
- domain assumption The HC-versus-PD statistical comparison has adequate power to detect meaningful differences.
Cite this review
Pith. "Pith review of Numerical Uncertainty in Linear Registration: An Experimental Study." pith.science (2026). https://pith.science/paper/MV3FXLCU
@misc{pith2026250800781,
author = {Pith},
title = {Pith review of: Numerical Uncertainty in Linear Registration: An Experimental Study},
year = {2026},
howpublished = {\url{https://pith.science/paper/MV3FXLCU}},
note = {Machine review of arXiv:2508.00781}
}
read the original abstract
While linear registration is a critical step in MRI preprocessing pipelines, its numerical uncertainty is understudied. Using Monte-Carlo Arithmetic (MCA) simulations, we assessed the most commonly used linear registration tools within major software packages (SPM, FSL, and ANTs) across multiple image similarity measures, two brain templates, and both healthy control (HC, n=50) and Parkinson's Disease (PD, n=50) cohorts. Our findings highlight the influence of linear registration tools and similarity measures on numerical stability. Among the evaluated tools and with default similarity measures, SPM exhibited the highest stability. FSL and ANTs showed greater and similar ranges of variability, with ANTs demonstrating particular sensitivity to numerical perturbations that occasionally led to registration failure. Furthermore, no significant differences were observed between healthy and PD cohorts, suggesting that numerical stability analyses obtained with healthy subjects may generalise to clinical populations. Finally, we also demonstrated how numerical uncertainty measures may support automated quality control (QC) of linear registration results. Overall, our experimental results characterize the numerical stability of linear registration experimentally and can serve as a basis for future uncertainty analyses.
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.