Pith. sign in

REVIEW 3 major objections 4 minor

Numerical Uncertainty in Linear Registration: An Experimental Study

T0 review · 3 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash

Pith's one-line read The paper claims that Monte-Carlo Arithmetic perturbation reveals stable and unstable linear registration tools, with SPM the most stable.

desk verdict A useful practical comparison of linear registration stability under Monte-Carlo Arithmetic, but the ranking's validity hinges on a perturbation model the abstract never calibrates. read the letter →

arxiv 2508.00781 v2 pith:MV3FXLCU submitted 2025-08-01 q-bio.QM eess.IV

classification q-bio.QMeess.IV
keywords linearregistrationnumericaluncertaintyMonte-CarloArithmeticMRIpreprocessingSPMFSLANTsqualitycontrol
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to quantify numerical uncertainty in linear registration of brain images, a preprocessing step whose rounding errors are rarely studied. It claims that with default similarity measures SPM is the most stable tool, FSL and ANTs are more variable, and ANTs can fail under numerical perturbation. The authors argue that these uncertainty measures can support automated quality control, and that stability rankings observed in healthy subjects carry over to a Parkinson's disease cohort. A sympathetic reader would care because registration errors propagate through every downstream analysis in neuroimaging pipelines.

What carries the argument

Monte-Carlo Arithmetic (MCA) is the mechanism: instead of a single deterministic run, each floating-point operation's low-order bits are randomly perturbed, and the registration is repeated many times. The spread of the resulting transformation parameters estimates the numerical uncertainty of the tool and similarity measure. This perturbation-based dispersion is what the paper uses to compare tools and to demonstrate a quality-control signal.

What would settle it

Run the same SPM, FSL, and ANTs linear registrations on real hardware while varying compiler flags, rounding modes, or CPU architectures that change floating-point evaluation order, and check whether ANTs produces occasional failures and whether the same stability ranking appears. If the ranking and failure rate do not reproduce under genuine hardware-dependent rounding, the Monte-Carlo Arithmetic model is not faithful enough to transfer.

Watch

Extended reading notes

Core claim

The paper reports that, across 50 healthy and 50 Parkinson's disease subjects, two templates, and several similarity measures, Monte-Carlo Arithmetic perturbations of floating-point operations produce measurable output variability in all three major linear registration tools. With default similarity measures, SPM shows the smallest dispersion, FSL and ANTs show larger and similar dispersion, and ANTs occasionally fails entirely under perturbation. The study finds no significant difference in numerical stability between healthy and PD cohorts, and it shows that dispersion measures can flag registrations whose results should not be trusted.

Load-bearing premise

The whole ranking depends on the assumption that perturbing low-order arithmetic bits in a simulation faithfully reproduces the true numerical errors these tools make on real MRI scans.

Editorial extensions

If this is right

  • Users of SPM's default settings can expect more reproducible linear registration results under numerical noise.
  • ANTs users should treat registration outputs as potentially unstable and should consider reruns or alternative tools when perturbation tests show large dispersion.
  • Automated QC pipelines could flag registrations with high numerical uncertainty without waiting for downstream artifacts.
  • Numerical stability findings from healthy cohorts can inform clinical studies of Parkinson's disease, reducing the need to repeat such analyses on every patient group.
  • Choice of similarity measure interacts with tool choice for numerical stability, so reproducible pipelines should pin both.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A practical extension would be to convert MCA dispersion into a per-subject QC score with a threshold tuned on real failing registrations.
  • The ANTs failures suggest a bifurcation in its optimization path under rounding; reproducing that bifurcation with reduced precision would identify the exact arithmetic step responsible.
  • If the healthy-to-clinical generalization holds beyond PD, numerical stability audits could be run on open healthy datasets and safely reused for many clinical studies.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. This manuscript reports an experimental study of numerical uncertainty in linear registration tools (SPM, FSL, ANTs) using Monte-Carlo Arithmetic. Based on the abstract, the authors claim that, with default similarity measures, SPM is the most stable, FSL and ANTs show greater and comparable variability with ANTs occasionally failing, that healthy controls (HC) and Parkinson's disease (PD) cohorts show no significant differences in numerical stability, and that numerical uncertainty measures can support automated quality control (QC) of linear registration. The study uses two brain templates, multiple similarity measures, and n=50 per cohort.

Significance. If the results are reliable, the paper would provide a useful empirical ranking of numerical robustness among widely used registration tools and a practical QC indicator. The use of a uniform external perturbation (MCA) across tools is a reasonable design choice, and the inclusion of a clinical cohort is a strength. The presented abstract alone, however, does not disclose the MCA calibration, the statistical power of the HC/PD comparison, or the validation of the proposed QC signal, so the significance currently depends on assumptions about the full text.

major comments (3)
  1. [Abstract, first sentence] The central claim of tool-specific stability is conditioned on the fidelity of Monte-Carlo Arithmetic. The abstract does not report the perturbation magnitude (e.g., fraction of mantissa bits perturbed), the number of Monte-Carlo runs per registration, or how the perturbation level was calibrated or validated against real hardware variability. Without this information, the SPM/FSL/ANTs ranking and the ANTs failure rate could be artifacts of the MCA hyperparameter rather than intrinsic numerical sensitivity.
  2. [Abstract, HC/PD comparison sentence] The generalization claim ('no significant differences were observed between healthy and PD cohorts') is a null result. The abstract reports neither effect sizes nor confidence intervals nor a power analysis for this comparison. With n=50 per cohort, the null could simply reflect insufficient statistical power, making the suggested generalization to clinical populations unsupported.
  3. [Abstract, final sentence] The QC demonstration is asserted without any reported validation metrics. The abstract says numerical uncertainty measures 'may support' automated QC but gives no sensitivity, specificity, or comparison to existing QC methods, so the claim is not substantiated at the level of the presented evidence.
minor comments (4)
  1. [Abstract, SPM/FSL/ANTs ranking] The phrase 'greater and similar ranges of variability' is ambiguous; please report the actual dispersion values (e.g., interquartile ranges or variance) for FSL and ANTs to support the comparison.
  2. [Abstract, methodology] The abstract does not state software versions, computing platform, or compiler/BLAS configuration, which are critical for a numerical reproducibility study.
  3. [Abstract, HC/PD comparison] The sentence 'no significant differences were observed' should be accompanied by a clear statement that this is a null result with its confidence interval, not evidence of equivalence.
  4. [Abstract, similarity measures] The term 'default similarity measures' is vague; please specify which measures were used for each tool in the abstract or clearly point to the full-text listing.

Circularity Check

0 steps flagged · score 0.0 of 10

No circularity found: the abstract reports an external perturbation experiment whose ranking and QC demonstration are not forced by construction from the inputs.

full rationale

This is an abstract-only review, and no equation, derivation, or fitted-parameter construction is available that would allow exhibiting a specific reduction of a claimed result to its own inputs. The central stability ranking is obtained by applying Monte-Carlo Arithmetic perturbation externally and uniformly to SPM, FSL, and ANTs; the outcome is not encoded in the perturbation model's definition. The HC/PD comparison uses an external clinical dataset and reports no significant differences, which is a contingent empirical finding rather than an analytic consequence of the methods. The QC demonstration is mentioned only as a demonstration, and without the methodology we cannot identify a success criterion defined from the same dispersion that defines the uncertainty measure; to flag circularity here would require quoting the specific construction, which the abstract does not provide. The load-bearing assumption that MCA perturbation faithfully reproduces true numerical uncertainty is a modeling assumption about fidelity and transferability, not a circularity: it does not make the ranking true by definition, and concerns about it belong to correctness risk rather than to the circularity score. There are no self-citations, no imported uniqueness theorems, and no renamed known results evident from the abstract. Accordingly, the honest finding is no significant circularity, score 0.

Assumptions & free parameters 2 free parameters · 3 assumptions · 0 invented entities

All entries here are simulation design choices and domain assumptions rather than fitted physical constants or new theoretical objects. The MCA perturbation magnitude and run count are chosen by the authors and are not stated in the abstract; together with the similarity-measure selection they determine the measured variability. The axioms concern the fidelity of the perturbation model, the representativeness of the templates and cohorts, and the power of the HC/PD comparison, each load-bearing for the generalization and QC claims. No invented entities are introduced.

free parameters (2)
  • MCA perturbation magnitude (fraction of low-order mantissa bits perturbed)
    Controls how strongly arithmetic is perturbed and directly determines the measured variability; the value is a methodological choice, not stated in the abstract, and the ranking depends on it.
  • Number of Monte-Carlo runs per registration
    Determines the resolution and statistical stability of the dispersion estimate; unstated in the abstract, and the width of the reported variability depends on it.
assumptions (3)
  • domain assumption Monte-Carlo Arithmetic perturbations faithfully simulate the true numerical uncertainty of the registration tools.
    The entire stability ranking is measured under MCA; if the perturbation model does not match real floating-point rounding behavior, the ranking would not transfer to practice. Invoked by the abstract's description of the MCA simulation setup.
  • domain assumption The two brain templates and the healthy/PD cohorts are representative enough to support the generalization claim.
    The paper generalizes from one healthy and one Parkinson's cohort (n=50 each) and two templates to a broader conclusion about clinical populations; this representativeness is assumed, not demonstrated in the abstract.
  • domain assumption The HC-versus-PD statistical comparison has adequate power to detect meaningful differences.
    The abstract converts a null result into a generalization statement; a null result from an underpowered test would not justify the claim, and no power analysis is reported.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Numerical Uncertainty in Linear Registration: An Experimental Study." pith.science (2026). https://pith.science/paper/MV3FXLCU

@misc{pith2026250800781,
  author       = {Pith},
  title        = {Pith review of: Numerical Uncertainty in Linear Registration: An Experimental Study},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/MV3FXLCU}},
  note         = {Machine review of arXiv:2508.00781}
}
read the original abstract

While linear registration is a critical step in MRI preprocessing pipelines, its numerical uncertainty is understudied. Using Monte-Carlo Arithmetic (MCA) simulations, we assessed the most commonly used linear registration tools within major software packages (SPM, FSL, and ANTs) across multiple image similarity measures, two brain templates, and both healthy control (HC, n=50) and Parkinson's Disease (PD, n=50) cohorts. Our findings highlight the influence of linear registration tools and similarity measures on numerical stability. Among the evaluated tools and with default similarity measures, SPM exhibited the highest stability. FSL and ANTs showed greater and similar ranges of variability, with ANTs demonstrating particular sensitivity to numerical perturbations that occasionally led to registration failure. Furthermore, no significant differences were observed between healthy and PD cohorts, suggesting that numerical stability analyses obtained with healthy subjects may generalise to clinical populations. Finally, we also demonstrated how numerical uncertainty measures may support automated quality control (QC) of linear registration results. Overall, our experimental results characterize the numerical stability of linear registration experimentally and can serve as a basis for future uncertainty analyses.

Discussion (0). Continue with ORCID to comment.

Pith tools

Reviewed August 6, 2026 · model on record in the stance chip above.