REVIEW 5 major objections 5 minor 59 references
Interpretable deformable image registration: A geometric deep learning perspective
T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read GeoReg models image registration as continuous-coordinate cross-attention, avoiding feature resampling between refinement steps.
desk verdict Genuinely novel, parameter-efficient registration architecture with strong synthetic-deformation results, but the abstract overclaims brain improvements that the paper's own body reports as 'on par' and no significance tests support. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the deformation function $\tau$, implemented as position-aware cross-attention. For a source point with feature $f$, the query is $q = f W_Q$, and keys and values are built from neighboring source or target features $F_N$ added to a Fourier positional embedding $E(X_N - \phi_n(x))$ of their coordinates relative to the point's current transformed position $\phi_n(x)$; a softmax-weighted combination yields the displacement update. Because neighbor contributions are weighted by continuous relative coordinates rather than fixed kernel positions, the function can be applied repeatedly at floating-point locations without resampling to a regular grid. A second learned function $\delta$ interpolates deformations across resolutions by letting child points cross-attend to parent control points in the coarser resolution, which the paper argues preserves sharp transformation boundaries that naive interpolation would smooth.
What would settle it
A controlled ablation would settle the central claim: replace the round-to-nearest neighborhood lookup with true continuous sampling, or conversely replace the learned interpolation $\delta$ with bilinear feature warping while holding everything else fixed, and measure the same Dice and Hausdorff metrics on the brain and OCT splits. If accuracy does not change, the continuous-coordinate mechanism is not the source of the gains.
Extended reading notes
Core claim
On its own terms, the paper's discovery is that deformation refinement can be cast as repeated position-aware cross-attention over continuous coordinates, so that successive refinements do not require resampling features to a regular grid. The authors show that a small model can capture the majority of a transformation at coarse resolutions, with finer decoder levels acting largely as interpolation, and that this scale-separated design remains end-to-end trainable with supervision at every resolution. They report that this formulation matches or exceeds larger baselines on inter-subject brain registration and outperforms them on longitudinal retinal OCT, while also recovering large synthetic affine-plus-noise deformations that other learned methods fail on.
Load-bearing premise
The load-bearing premise is that evaluating the deformation function at continuous coordinates actually provides the target information it is supposed to, since neighborhoods are gathered by rounding the current floating-point coordinate to the nearest grid index; if sub-voxel motion does not change the gathered neighbors, the claimed advantage over feature warping would rest on the learned cross-resolution interpolation and supervision scheme instead.
Editorial extensions
If this is right
- Registration models can refine transformations without warping feature vectors or re-encoding images between steps, removing a source of interpolation error and repeated encoding cost.
- Feature vectors do not need to be carried across decoder levels; only transformation vectors are passed between resolutions, simplifying the architecture.
- Because the majority of the deformation is assigned to coarse resolutions, finer decoder levels can be cheap interpolation-only layers.
- Supervising every refinement and interpolation step propagates meaningful gradients through the whole pipeline and acts as an implicit regularizer on featureless regions.
- Encoding positions as continuous coordinates means the same formulation can in principle adapt to variable grid spacing, which the authors suggest as relevant to anisotropic data.
Reading between the lines
- One consequence the authors leave implicit is that their implementation gathers target neighborhoods by rounding the current floating-point coordinate to the nearest grid index, so the 'continuous' advantage is only approximate; isolating true continuous sampling from round-to-nearest would clarify which design choice drives the reported gains.
- A natural testable extension is to apply the same $\tau$ and $\delta$ formulation to non-grid domains such as cortical surfaces or point clouds, where rounding to a grid is impossible and the geometric deep learning framing is native.
- The observed scale separation suggests the multi-resolution objective could be tuned with different similarity or regularization weights per resolution, a knob the paper discusses qualitatively but does not investigate quantitatively.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. GeoReg is a deformable image registration framework built on geometric deep learning principles. It uses a dual-stream encoder to extract independent source/target feature pyramids, then refines deformations in a coarse-to-fine manner with a position-aware cross-attention deformation function τ evaluated on continuous coordinates, together with a learned cross-resolution interpolation function δ. The paper claims that this spatially continuous formulation avoids resampling to a regular grid between successive refinements, and it evaluates the method on CamCAN T1w/T1w-T2w inter-subject brain registration, longitudinal retinal OCT registration, and synthetic large-deformation tasks, with code released publicly.
Significance. The geometric-deep-learning framing is timely and the architecture is genuinely parameter-efficient (741k parameters versus 46.8M for Transmorph). The qualitative scale-separation analysis in Figures 3-4 is a useful interpretability contribution, and the synthetic large-deformation experiments with ground-truth transformations in Table 3 are a valuable stress test in which GeoReg is consistently strong and nearly folding-free. If the empirical claims are properly supported, the paper would make a meaningful contribution to interpretable and data-efficient registration. However, as written, the headline claim in the abstract is not supported by the paper's own reported brain results, and the continuous-neighborhood claim is only approximately realized in the implementation; both need to be addressed before publication.
major comments (5)
- [Abstract, Section 4.2, Table 1] The abstract claims "significant improvement in performance metrics over state-of-the-art approaches for both mono- and multi-modal inter-subject brain registration," but Section 4.2 states that on both brain tasks the method "shows on par performance," and Table 1 shows overlapping standard deviations on every brain metric. For T1w brain, GeoReg's DSC is 0.838±0.06 versus Transmorph's 0.822±0.01, its HD95 is 2.05±0.82 versus 1.92±0.52, and its folding is 0.09 versus 0.02; on T1w-T2w the DSC values are identical at 0.788. No paired significance tests, confidence intervals, or effect sizes are reported. Please either add a proper statistical comparison (paired tests across subjects, with test-set sizes and confidence intervals) or revise the abstract and conclusion to state that brain performance is on par with, not significantly better than, the state of the art.
- [Appendix 10, Section 3.2] The "spatially continuous" claim is weakened by the implementation detail in Appendix 10: target neighborhoods are obtained by "mapping the current source coordinates into the index space of the target grid and rounding to the closest integer." This means that sub-voxel updates of the transformed position φ_n(x) do not change the set of target features seen by the deformation function; the continuous relative-coordinate embedding does not eliminate the discretization introduced by nearest-neighbor rounding. Please clarify how much of the claimed advantage over feature warping remains under this rounding, and ideally provide an ablation comparing against trilinear feature interpolation or soft/continuous neighborhood weighting.
- [Section 4.2, Table 1 (OCT)] The OCT improvement is also not statistically supported. GeoReg's DSC is 0.581±0.09 versus Transmorph's 0.577±0.12 and HD95 is 1.72±1.72 versus 1.95±2.54, all from a single 80/10/10 split with no significance tests. Given the large standard deviations, the reported differences may be within noise. Please report the number of test volumes, paired significance tests, and confidence intervals for all datasets.
- [Section 3.3, Table 1 (feat. warp ablation)] The architectural axiom that features need not be propagated across decoder levels is not isolated by the "feat. warp" ablation. Replacing δ with bilinear feature warping changes several factors at once: how features are transported, whether neighborhoods are integer-rounded, and what information is available to the interpolation function. To support the claim in Section 3.3 that local features at each resolution are sufficient, please add an ablation that propagates or concatenates feature vectors across decoder levels while keeping the rest of the design fixed.
- [Table 3] LapIRN results are entirely missing from Table 3 (all entries are "− ± −") with no explanation, despite LapIRN being a key multi-resolution baseline in Table 1. Without these entries, or an explicit statement of why LapIRN could not be evaluated, the claim that GeoReg "consistently outperforms other baselines" on the synthetic large-deformation tasks is not fully supported.
minor comments (5)
- [Section 3.4] The word "correspondances" should be "correspondences".
- [Appendix 10] The phrase "utilize bending energy as a regularize" should read "as a regularizer".
- [Table 1] Table 1 contains stray spaces in several numerical entries (e.g., "5 .46", "2 .55", "2 .05") and inconsistent decimal formatting; please clean up the table.
- [Equation (10), Appendix 10] The resolution-specific weights α_r and the regularization weight λ are not specified in the paper; please report their values for reproducibility.
- [Figure 16] The caption of Figure 16 lists method names without a clear correspondence to panels; please label the subfigures and describe the layout.
Circularity Check
No significant circularity: the architecture is not derived from its own outputs and the evaluation relies on held-out external labels.
full rationale
The paper's derivation chain is architectural and empirical, not definitional. The method is motivated by a stated design foundation (separated feature extraction, dynamic receptive fields, continuous-coordinate deformation modeling) and implemented through explicit equations: the generalized continuous convolution (Eq. 3), position-aware cross-attention with Fourier positional embeddings (Eqs. 4-7), the learned cross-resolution interpolation (Section 3.3), and the multi-resolution supervision objective (Eqs. 8-10). None of these quantities is defined in terms of the reported outcomes, and no fitted parameter is later renamed as a prediction. The evaluation is external: test performance in Table 1 is measured against held-out 10% splits with independently produced label sets (MALPEM for brain, manual retinal layer labels for OCT), using Dice, HD95, and folding metrics. The self-citations that appear (e.g., Refs. 22, 29, 30, 32, 38, involving the authors) are dataset, baseline, regularization, or preprocessing provenance and are not load-bearing supports for the novelty claims; they do not carry the central argument. The abstract's wording 'significant improvement' is in tension with the body's own 'on par' conclusion and with overlapping standard deviations, but that is a statistical reporting weakness, not a circularity: the claim is not forced by the construction of the method. Similarly, Appendix 10's use of integer rounding when indexing target neighborhoods qualifies the 'no resampling' claim but does not make the continuous-coordinate formulation equivalent to its own input; the position embeddings still consume continuous relative coordinates. No step in the paper reduces, by its own equations or by a self-citation chain, to the quantity it purports to predict or explain.
Assumptions & free parameters
free parameters (5)
- Number of refinement iterations N =
4 iterations on the two coarsest resolutions; 0 on finest layer (interpolation-only)
- Target neighborhood size for tau =
3x3 or 5x5 depending on resolution
- Resolution-specific loss weights alpha_r =
not reported
- Regularization weight for bending energy =
not reported
- Neighborhood size for delta =
closest 3x3 points of previous coarser resolution
assumptions (5)
- standard math Distributive property of convolution over concatenated source-target channels
- standard math Fourier feature positional encoding provides a suitable continuous relative-position embedding
- domain assumption Local cross-attention over small neighborhoods is sufficient to model deformations at each resolution
- ad hoc to paper Features do not need to be propagated across decoder resolutions; only transformations are propagated
- ad hoc to paper Rounding continuous coordinates to nearest integer grid indices does not harm deformation accuracy
Cite this review
Pith. "Pith review of Interpretable deformable image registration: A geometric deep learning perspective." pith.science (2026). https://pith.science/paper/AR5PAHE2
@misc{pith2026241213294,
author = {Pith},
title = {Pith review of: Interpretable deformable image registration: A geometric deep learning perspective},
year = {2026},
howpublished = {\url{https://pith.science/paper/AR5PAHE2}},
note = {Machine review of arXiv:2412.13294}
}
read the original abstract
Deformable image registration poses a challenging problem where, unlike most deep learning tasks, a complex relationship between multiple coordinate systems has to be considered. Although data-driven methods have shown promising capabilities to model complex non-linear transformations, existing works employ standard deep learning architectures assuming they are general black-box solvers. We argue that understanding how learned operations perform pattern-matching between the features in the source and target domains is the key to building robust, data-efficient, and interpretable architectures. We present a theoretical foundation for designing an interpretable registration framework: separated feature extraction and deformation modeling, dynamic receptive fields, and a data-driven deformation functions awareness of the relationship between both spatial domains. Based on this foundation, we formulate an end-to-end process that refines transformations in a coarse-to-fine fashion. Our architecture employs spatially continuous deformation modeling functions that use geometric deep-learning principles, therefore avoiding the problematic approach of resampling to a regular grid between successive refinements of the transformation. We perform a qualitative investigation to highlight interesting interpretability properties of our architecture. We conclude by showing significant improvement in performance metrics over state-of-the-art approaches for both mono- and multi-modal inter-subject brain registration, as well as the challenging task of longitudinal retinal intra-subject registration. We make our code publicly available
Figures
Figures from the paper (13 more)
Reference graph
Works this paper leans on
-
[1]
B. B. Avants, C. L. Epstein, M. Grossman, and J. C. Gee. Symmetric diffeomorphic image registration with cross- correlation: evaluating automated labeling of elderly and neurodegenerative brain. Medical image analysis, 12(1):26– 41, 2008
work page 2008
-
[2]
B. B. Avants, N. Tustison, G. Song, et al. Advanced normal- ization tools (ANTS). Insight j, 2(365):1–35, 2009
work page 2009
-
[3]
G. Balakrishnan, A. Zhao, M. R. Sabuncu, J. Guttag, and A. V . Dalca. V oxelmorph: a learning framework for de- formable medical image registration. IEEE Transactions on Medical Imaging, 38(8):1788–1800, 2019
work page 2019
-
[4]
M. M. Bronstein, J. Bruna, T. Cohen, and P. Veliˇckovi´c. Ge- ometric deep learning: Grids, groups, graphs, geodesics, and gauges. arXiv preprint arXiv:2104.13478, 2021
arXiv 2021
-
[5]
J. Chen, E. C. Frey, Y . He, W. P. Segars, Y . Li, and Y . Du. Transmorph: Transformer for unsupervised medical image registration. Medical image analysis, 82:102615, 2022
work page 2022
-
[6]
Z. Chen, Y . Zheng, and J. C. Gee. Transmatch: a transformer-based multilevel dual-stream feature matching network for unsupervised deformable image registration. IEEE Transactions on Medical Imaging, 2023
work page 2023
-
[7]
A. V . Dalca, G. Balakrishnan, J. V . Guttag, and M. R. Sabuncu. Unsupervised learning for fast probabilistic diffeo- morphic registration. In International Conference on Med- ical Image Computing and Computer-Assisted Intervention , 2018
work page 2018
-
[8]
L. Deng. The mnist database of handwritten digit images for machine learning research [best of the web]. IEEE signal processing magazine, 29(6):141–142, 2012
2012
Show all 59 references
-
[9]
Farahani, A.and Vitay and F
J. Farahani, A.and Vitay and F. H. Hamker. Deep neural net- works for geometric shape deformation. In German Confer- ence on Artificial Intelligence (K¨unstliche Intelligenz), pages 90–95. Springer, 2022
2022
-
[10]
Fuchs, D
F. Fuchs, D. Worrall, V . Fischer, and M. Welling. Se (3)- transformers: 3d roto-translation equivariant attention net- works. Advances in neural information processing systems , 33:1970–1981, 2020
1970
-
[11]
H-ViT: A hierarchical vision transformer for deformable im- age registration
Morteza Ghahremani, Mohammad Khateri, Bailiang Jian, Benedikt Wiestler, Ehsan Adeli, and Christian Wachinger. H-ViT: A hierarchical vision transformer for deformable im- age registration. In Proceedings of the IEEE/CVF Confer- 8 Figure 5. Qualitative results of all compared me...
2024
-
[12]
Hansen and M
L. Hansen and M. P. Heinrich. Deep learning based geo- metric registration for medical images: How accurate can we get without visual features? In Information Processing in Medical Imaging: 27th International Conference, IPMI 2021, Virtual Event, June 28–June 30, 2021, Proceed...
2021
-
[13]
Hansen and M
L. Hansen and M. P. Heinrich. Graphregnet: Deep graph regularisation networks on sparse keypoints for dense reg- istration of 3D lung CTs. IEEE Transactions on Medical Imaging, 40(9):2246–2257, 2021
2021
-
[14]
Haskins, U
G. Haskins, U. Kruger, and P. Yan. Deep learning in medical image registration: a survey. Machine Vision and Applica- tions, 31:1–18, 2020
2020
-
[15]
Hoopes, J
A. Hoopes, J. E. Iglesias, B. Fischl, D. Greve, and A. V . Dalca. TopoFit: rapid reconstruction of topologically-correct cortical surfaces. Proceedings of machine learning research, 172:508, 2022
2022
-
[16]
A. Horn. MNI T1 6thGen NLIN to MNI 2009b NLIN ANTs transform. 2016
2016
-
[17]
B. Hu, S. Zhou, Z. Xiong, and F. Wu. Recursive decom- position network for deformable image registration. IEEE Journal of Biomedical and Health Informatics, 26(10):5130– 5141, 2022
2022
-
[18]
Iglesias, C
J. Iglesias, C. Liu, P. Thompson, and Z. Tu. Robust brain ex- traction across datasets and comparison with publicly avail- able methods. IEEE Transactions on Medical Imaging , 30(9):1617–1634, 2011
2011
-
[19]
X. Jia, J. Bartlett, T. Zhang, W. Lu, Z. Qiu, and J. Duan. U- net vs transformer: Is u-net outdated in medical image regis- tration? In International Workshop on Machine Learning in Medical Imaging, pages 151–160. Springer, 2022
2022
-
[20]
M. Kang, X. Hu, W. Huang, M. R. Scott, and M. Reyes. Dual-stream pyramid registration network. Medical image analysis, 78:102379, 2022
2022
-
[21]
Kashefi, D
A. Kashefi, D. Rempe, and L.s J. Guibas. A point-cloud deep learning framework for prediction of fluid flow fields on irregular geometries. Physics of Fluids, 33(2), 2021
2021
-
[22]
Ledig, R
C. Ledig, R. Heckemann, A. Hammers, J. L ´opez, V . New- 9 combe, A. Makropoulos, J. L ¨otj¨onen, D. Menon, and D. Rueckert. Robust whole-brain segmentation: Application to traumatic brain injury. Medical image analysis, 21 1:40–58, 2015
2015
-
[23]
Y . Liu, L. Zuo, S. Han, Y . Xue, J. L. Prince, and A. Carass. Coordinate translator for learning deformable medical image registration. In International Workshop on Multiscale Multi- modal Medical Imaging, pages 98–109. Springer, 2022
2022
-
[24]
Lowekamp, D
B. Lowekamp, D. Chen, L. Ib ´a˜nez, and D. Blezek. The de- sign of simpleitk. Frontiers in Neuroinformatics, 7, 2013
2013
-
[25]
IIRP-Net: iterative inference residual pyramid network for enhanced image registration
Tai Ma, Suwei Zhang, Jiafeng Li, and Ying Wen. IIRP-Net: iterative inference residual pyramid network for enhanced image registration. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern Recognition, pages 11546–11555, 2024
2024
-
[26]
M. Meng, L. Bi, D. Feng, and J. Kim. Non-iterative coarse- to-fine registration based on single-pass deep cumulative learning. In International Conference on Medical Image Computing and Computer-Assisted Intervention , pages 88–
-
[27]
Mok and A.C.S
T.C.W. Mok and A.C.S. Chung. Conditional deformable im- age registration with convolutional neural network. In Med- ical Image Computing and Computer Assisted Intervention– MICCAI 2021: 24th International Conference, Strasbourg, France, September 27–October 1, 2021, Proceeding...
2021
-
[28]
T. C. W. Mok and A. C. S. Chung. Large deformation dif- feomorphic image registration with Laplacian pyramid net- works. In Medical Image Computing and Computer Assisted Intervention–MICCAI 2020: 23rd International Conference, Lima, Peru, October 4–8, 2020, Proceedings, Part I...
2020
-
[29]
H. Qiu, C. Qin, A. Schuh, K. Hammernik, and D. Rueck- ert. Learning diffeomorphic and modality-invariant registra- tion using B-splines. InInternational Conference on Medical Imaging with Deep Learning, 2021
2021
-
[30]
Rueckert, L
D. Rueckert, L. I. Sonoda, C. Hayes, D. L. G. Hill, M. O. Leach, and D. J. Hawkes. Nonrigid registration using free- form deformations: application to breast MR images. IEEE Transactions on Medical Imaging, 18:712–721, 1999
1999
-
[31]
Sandk ¨uhler, S
R. Sandk ¨uhler, S. Andermatt, G. Bauman, S. Nyilas, C. Jud, and P. C. Cattin. Recurrent registration neural networks for deformable image registration. Advances in Neural Informa- tion Processing Systems, 32, 2019
2019
-
[32]
Schuh, M
A. Schuh, M. Murgasova, A. Makropoulos, C. Ledig, S. Counsell, J. Hajnal, P. Aljabar, and D. Rueckert. Construc- tion of a 4D brain atlas and growth model using diffeomor- phic registration. In STIA, 2014
2014
-
[33]
Shafto, L
M. Shafto, L. Tyler, M. Dixon, Jason R. Taylor, J. Rowe, R. Cusack, A. Calder, W. D. Marslen-Wilson, J. Duncan, T. Dalgleish, R. Henson, C. Brayne, and F. Matthews. The Cambridge centre for ageing and neuroscience (Cam-CAN) study protocol: a cross-sectional, lifespan, multidis...
2014
-
[34]
Z. Shen, J. Feydy, P. Liu, A. H. Curiale, R. San Jose Estepar, R. San Jose Estepar, and M. Niethammer. Accurate point cloud registration with robust optimal transport. Advances in Neural Information Processing Systems , 34:5373–5389, 2021
2021
-
[35]
Sotiras, C
A. Sotiras, C. Davatzikos, and N. Paragios. Deformable med- ical image registration: A survey. IEEE Transactions on Medical Imaging, 32:1153–1190, 2013
2013
-
[36]
M. A. Suliman, L. Z. J. Williams, A. Fawaz, and E. C. Robin- son. Geomorph: Geometric deep learning for cortical surface registration. In Geometric Deep Learning in Medical Image Analysis, 2022
2022
-
[37]
D. Sun, X. Yang, M.-Y . Liu, and J. Kautz. Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume. In Proceedings of the IEEE conference on computer vision and pattern recognition, pages 8934–8943, 2018
2018
-
[38]
Sutton, M
J. Sutton, M. J. Menten, S. Riedl, H. Bogunovi ´c, O. Lein- gang, P. Anders, A. M. Hagag, S. Waldstein, A. Wilson, A. J. Cree, et al. Developing and validating a multivariable pre- diction model which predicts progression of intermediate to late age-related macular degeneratio...
2023
-
[39]
Tancik, P
M. Tancik, P. Srinivasan, B. Mildenhall, S. Fridovich-Keil, N. Raghavan, U. Singhal, R. Ramamoorthi, J. Barron, and R. Ng. Fourier features let networks learn high frequency functions in low dimensional domains. NeurIPS, 2020
2020
-
[40]
Taylor, N
J. Taylor, N. Williams, R. Cusack, T. Auer, M. Shafto, M. Dixon, L. Tyler, Cam-CAN Group, and R. Henson. The cambridge centre for ageing and neuroscience (Cam-CAN) data repository: Structural and functional MRI, MEG, and cognitive data from a cross-sectional adult lifespan sam...
2017
-
[41]
Verleysen and D
M. Verleysen and D. Franc ¸ois. The curse of dimensionality in data mining and time series prediction. In International work-conference on artificial neural networks , pages 758–
-
[42]
H. Wang, D. Ni, and Y . Wang. Modet: Learning deformable image registration via motion decomposition transformer. In International Conference on Medical Image Computing and Computer-Assisted Intervention , pages 740–749. Springer, 2023
2023
-
[43]
H. Xiao, X. Teng, C. Liu, T. Li, G. Ren, R. Yang, D. Shen, and J. Cai. A review of deep learning-based three- dimensional medical image registration methods. Quantita- tive Imaging in Medicine and Surgery, 11(12):4895, 2021
2021
-
[44]
S. Zhao, Y . Dong, E. I. Chang, Y . Xu, et al. Recursive cas- caded networks for unsupervised medical image registration. In Proceedings of the IEEE/CVF international conference on computer vision, pages 10600–10610, 2019
2019
-
[45]
Zhu and S
Y . Zhu and S. Lu. Swin-voxelmorph: A symmetric unsuper- vised learning model for deformable medical image registra- tion using swin transformer. In International Conference on Medical Image Computing and Computer-Assisted Interven- tion, pages 78–87. Springer, 2022. 10 Interp...
2022
-
[46]
Distributive property over neighbors Assume a feature grid F ∈ R din×H×W ×D where d is the feature dimension and H, W, Dare spatial dimensions
Proof on distributive property of convolu- tions over source-target domains 6.1. Distributive property over neighbors Assume a feature grid F ∈ R din×H×W ×D where d is the feature dimension and H, W, Dare spatial dimensions. A convolution over the domain F would utilize a weig...
-
[47]
First a neighbor-wise feature projection is performed us- ing a projection matrix W[:,j] ∈ R (din)×dout on each element j in the k × k × k neighborhood
-
[48]
Next, we perform a uniformly weighted addition of all projection vectors. 6.2. Distributive property over channels As established, each neighbor f j ∈ FNi around cen- tral node i has an independent projection matrix W[:,j] ∈ R din×dout. Because of the distributive property of ...
-
[49]
masking out
Learned interpolation connection to para- metric interpolation 7.1. Parametric interpolation A commonly adopted technique in parametric image regis- tration involves predicting deformations at a coarse spacing and interpolating to the desired resolution via continuous 1 mappin...
-
[50]
CamCAN dataset consists of 310 T1w and T2w MR 3D images (160 × 180 × 160, 1mm3 isotropic resolution)
Data pre-processing Dataset, pre-processing, and label information of the Cam- CAN [33, 40] and the challenging retinal optical coher- ence tomography (OCT) images [38] datasets. CamCAN dataset consists of 310 T1w and T2w MR 3D images (160 × 180 × 160, 1mm3 isotropic resolutio...
-
[51]
feat. warp
Baselines We compare our method (GeoReg) against several conven- tional iterative methods and learning-based image registra- tion models. Regarding the iterative optimization methods, we choose from the Medical Image Registration ToolKit 2 (MIRTK) [32], a widely-used free-form...
-
[52]
Implementation details Our approach utilizes a lightweight dual-stream encoder de- sign to independently extract features for the source S and target T images. The encoder consists of two convolutional residual blocks per layer followed by pooling layers, which allow hierarchi...
-
[53]
VRAM requirements per model in GigaBytes (GB) under a batch size of 1
Training memory footprints Table 2. VRAM requirements per model in GigaBytes (GB) under a batch size of 1. Models VRAM V oxelMorph 3.55 GB LapIRN 6.67 GB Transmorph 7.09 GB D-PRNet 11.11 GB RCN 6.21 GB LK-UNet 4.18 GB Ours (feat. warp) 6.75 GB Ours (GeoReg) 9.08 GB
-
[54]
Architectural overview of the proposed method
Architectural overview Figure 6. Architectural overview of the proposed method. A dual- stream encoders extracts features independently from source and target images. The decoder does not carry features across resolu- tions, only transformations are passed across resolutions i...
-
[55]
Learnable parameters θτ
Pseudocode of function τ Algorithm 1: Pseudocode of implementation of deforma- tion function τ for a given decoder resolution 1 Function τ (F S, X S, F T , X T , θτ ): Input: Features and coordinates of source points S and target points T . Learnable parameters θτ . Output: De...
-
[56]
Pseudocode of function δ Algorithm 2: Pseudocode of implementation of deforma- tion interpolation function δ between decoder resolution layers r and r − 1 1 Function δ (F r+1, X r+1, U r+1, F r, X r, θδ): Input: Features F r+1, starting coordinates X r+1, and deformation U r+1...
-
[57]
Learnable parameters θδ
Features F r and coordinates X r source points S in layer r. Learnable parameters θδ. Output: Deformations U r for source points Sr in reso- lution r // Initialize deformations as zeros U r ← 0Sr×3 // Perform node-wise deformation interpolation on layer r for s ∈ Sr do f ← F r...
-
[58]
We create a dataset of intra-subject brain pairs with varying ranges of non-rigid deformations com- prised of a combination of an affine and Brownian noise components
Qualitative results on large synthetic affine transformations In this section we investigate varying kinds of large syn- thetic deformations without any form of affine registration preprocessing. We create a dataset of intra-subject brain pairs with varying ranges of non-rigid...
-
[59]
Qualitative results of all compared methods for the CamCAN T1w-T1w inter-subject deformable registration experiment
Qualitative results Figure 10. Qualitative results of all compared methods for the CamCAN T1w-T1w inter-subject deformable registration experiment. Figure 11. Qualitative results of all compared methods for the CamCAN T1w-T1w inter-subject deformable registration experiment. F...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.