REVIEW 3 major objections 6 minor 2 cited by
3D Wavelet Latent Diffusion Model for Whole-Body MR-to-CT Modality Translation
T0 review · 3 major / 6 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A 3D wavelet latent diffusion model synthesizes whole-body CT from MR scans, reporting PSNR gains up to 3.98 dB over GAN and diffusion baselines and downstream organ segmentation Dice gains of 4.1–17.2%.
desk verdict Solid, well-engineered incremental advance in whole-body MR-to-CT synthesis, but the SOTA claim is undercut by missing 3D diffusion baselines and a private evaluation. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The framework is a two-stage 3D wavelet latent diffusion model. Stage one pretrains shared-parameter encoders and decoders with the disentanglement loss $L_D = L_{stru} + L_{modal}$: the latent vector is split channel-wise into structural and modality halves, paired samples are pushed to share structure, and unpaired samples are pushed to share modality. Wavelet Residual Modules inside the wavelet blocks decompose features with 3D wavelet transforms and re-fuse low- and high-frequency streams to preserve fine detail. Stage two runs the CT latent through a 1000-step forward diffusion and a reverse denoising process conditioned on the MR latent, with Dual Skip Connection Attention (DSCA) using a structure-emphasis cross-attention and a modality-filtering subtraction at each skip connection. This set of mechanisms carries the claim that anatomy stays anchored while modality appearance is transferred.
What would settle it
Take a set of whole-body pairs, intentionally perturb the registration (e.g., simulate 5–10 mm misalignment in the abdomen), retrain or fine-tune 3D-WLDM, and measure PSNR/SSIM plus a fixed landmark error; if the structural anchor cannot correct or flag the misalignment, the reported spatial-consistency gain should disappear or invert. A simpler check: run the released code on a public paired whole-body dataset with independent registration and see whether the 1.04 dB margin over the best baseline survives.
Extended reading notes
Core claim
The paper establishes that whole-body MR-to-CT synthesis can be made spatially consistent by performing the translation in a learned latent space with explicit separation of structure and modality. The structural component of the latent code is aligned between paired MR and CT volumes and anchored during denoising, while the modality component is pushed to encode only MR- or CT-specific appearance. Wavelet residual modules add frequency-resolved detail during encoding and decoding, and dual skip connection attention suppresses MR-only textures when conditioning denoising on the MR latent. The claimed outcome is synthetic CT with better bony structure and soft-tissue contrast than CycleGAN, 2D/3D latent diffusion, and pixel-space diffusion baselines, and enough fidelity to improve automated organ and spine segmentation.
Load-bearing premise
The load-bearing premise is that each MR and CT pair is accurately aligned before training; if the alignment is off, the model will treat mismatched anatomy as the 'structure' to preserve, and the synthesized CT inherits the registration error.
Editorial extensions
If this is right
- If 3D-WLDM generalizes beyond the study population, MR-only radiotherapy workflows could compute dose directly from synthetic CT, removing the CT acquisition and its registration step.
- PET/MR scanners could use the synthesized CT as an attenuation map, replacing the need for a separate low-dose CT or atlas-based pseudo-CT.
- Because the synthetic CT is aligned to the input MR, cross-modal registration tasks become intra-modal alignment problems, simplifying longitudinal and multi-modal tracking.
- CT-trained analysis tools such as organ and bone segmentation can be applied to MR-acquired images, with the reported Dice gains implying clinically usable anatomy.
Reading between the lines
- Editorial inference: the disentanglement idea should transfer to other paired cross-modality synthesis tasks (e.g., MR-to-PET or T1-to-T2 MRI), since it only assumes shared anatomy plus distinct modality appearance.
- Editorial inference: the registration sensitivity flagged in the conclusion suggests a testable boundary — performance should degrade gracefully as MR-CT misalignment increases, and a version trained with realistic motion augmentation would reveal how much of the gain is registration-dependent.
- Editorial inference: the reported gains are on a private single-scanner dataset; a multi-site public evaluation would be needed to know whether the 0.02 SSIM margin over CycleGAN is clinically meaningful.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes 3D-WLDM, a 3D wavelet latent diffusion model for whole-body MR-to-CT synthesis. The method has two stages: a pretraining stage that learns a shared latent space with wavelet residual modules and structure-modality disentanglement, and a translation stage that uses a latent diffusion model conditioned on MR latent features with a dual skip connection attention mechanism. The model is evaluated on a private multi-centre dataset of 268 subjects, with quantitative metrics (PSNR, SSIM, MAE, NCC), qualitative comparisons, and a downstream organ segmentation task. The authors report consistent improvements over several GAN- and diffusion-based baselines and release the source code.
Significance. If the reported results hold, 3D-WLDM addresses a clinically relevant problem—MR-only radiotherapy and PET/MR attenuation correction—and introduces a technically coherent architecture. The release of source code and the use of a downstream segmentation evaluation are strengths. However, the evidence for the central superiority claim is incomplete: the three closest 3D diffusion baselines (MC-DDPM, WDM, Tapp et al.) are cited but not compared, and one baseline result (DDIM's SSIM) is internally inconsistent, which directly affects a headline claim in the abstract. These issues must be resolved before the state-of-the-art claim can be accepted.
major comments (3)
- [Section II.B and Section IV.B (Table I)] The three most relevant 3D diffusion baselines—MC-DDPM [10], WDM [35], and Tapp et al. [36]—are described in Related Work but are absent from the quantitative comparison in Table I. Without these comparisons, the conclusion in Section V that 3D-WLDM shows "clear advantages over the state-of-the-art GAN-based and diffusion-based methods" is not supported. Please add these baselines, or provide a concrete justification for their omission, and temper the state-of-the-art claim accordingly.
- [Section IV.B, Table I] The reported DDIM metrics are internally inconsistent: DDIM achieves higher PSNR (24.29 vs 21.35 dB) and lower MAE (76.64 vs 101.04) than DDPM, but its SSIM is substantially lower (0.45 vs 0.60). This pattern is atypical and is not discussed in the text. Because the abstract's headline "SSIM improvements of up to 0.36" is computed with respect to DDIM, this anomaly directly affects a central claim. Please verify the metric calculation or provide a specific, data-based explanation for the discrepancy.
- [Section IV.B, Tables I and III] No statistical significance tests or confidence intervals are reported for any of the quantitative comparisons. This is particularly important because the advantage over the best baseline is only 0.02 in SSIM (0.81 vs 0.79) and 0.01 in NCC (0.97 vs 0.96). Please report per-subject results and appropriate significance tests (e.g., paired t-test or Wilcoxon signed-rank test) to support the claim that 3D-WLDM is consistently superior.
minor comments (6)
- [Section III.C, Eqs. (9) and (10)] The sentences after Eqs. (9) and (10) are incomplete: "The output processed via captures consistent structural features" and "The output processed via filters out modality-specific noise" should be reworded to describe the computation fully.
- [Section IV.C.2] The text says "The addition of WDM and the DSCA module" but this should read "WRM" (Wavelet Residual Module), not WDM, to avoid confusion with the WDM baseline in Related Work.
- [Section IV.A] The voxel spacing is written as "2mm 3"; this should be typeset as 2 mm³.
- [Section IV.B.2] The sentence "NCC increases of 0.04" is ambiguous because the gain over the best baseline (CycleGAN, 0.96) is only 0.01; please clearly specify the baseline for each reported improvement range (PSNR, SSIM, MAE, NCC).
- [Section III.B, Eq. (8)] The values of the hyperparameters α, β, and γ in the pretraining loss L_pre are not specified. Please provide the values and, if possible, a brief sensitivity analysis.
- [Section IV] The paper does not report training time, inference time, or GPU memory consumption. Since whole-body imaging is computationally demanding, these practical details would strengthen the clinical applicability discussion.
Circularity Check
No significant circularity: the method's components are explicit regularizers and the main claims rest on held-out comparisons, not on fitted inputs or load-bearing self-citations.
full rationale
3D-WLDM is an empirical architecture paper whose central claims are evaluated by comparing generated CT volumes against ground-truth CT on a held-out test split. No parameter is fitted to the evaluation metric in a way that forces the reported PSNR/SSIM/MAE/NCC gains by construction. The Structure-Modality Disentanglement losses (Eqs. 1-7) are explicit pretraining regularizers: they encourage structural latent components of CT and MR to be similar and modality components to be dissimilar, but this constraint does not by itself determine the image-space synthesis metrics, and the diffusion objective (Eq. 12) is a standard noise-prediction loss conditioned on MR latents. The paper's acknowledged dependence on Elastix co-registration is an assumption about input data quality rather than a circular step. Self-citations to prior work by coauthors (e.g., [13], [38]) appear only in related-work context and do not carry the argument. Concerns such as the omission of MC-DDPM, WDM, and Tapp et al. from the comparison table, and the atypical DDIM SSIM value, are about completeness and internal consistency of the evidence, not about circular derivation. No equation or claimed prediction reduces to its own input by construction.
Assumptions & free parameters
free parameters (3)
- alpha, beta, gamma in total pretraining loss L_pre (Eq. 8) =
not reported
- Diffusion timesteps T and beta scheduler =
T=1000, beta_1=1e-4, beta_T=0.2
- Latent channel split ratio between structural and modality components =
1:1 split
assumptions (3)
- domain assumption MR and CT volumes from the same subject are accurately aligned after rigid and non-rigid Elastix registration using mutual information.
- domain assumption Cosine similarity on latent vectors can separate anatomical structure from modality-specific appearance.
- standard math The standard DDPM noise-prediction objective is sufficient for cross-modal translation in the latent space when conditioned on the MR latent.
Cite this review
Pith. "Pith review of 3D Wavelet Latent Diffusion Model for Whole-Body MR-to-CT Modality Translation." pith.science (2026). https://pith.science/paper/EWMI52U5
@misc{pith2026250711557,
author = {Pith},
title = {Pith review of: 3D Wavelet Latent Diffusion Model for Whole-Body MR-to-CT Modality Translation},
year = {2026},
howpublished = {\url{https://pith.science/paper/EWMI52U5}},
note = {Machine review of arXiv:2507.11557}
}
read the original abstract
Magnetic Resonance (MR) imaging plays an essential role in contemporary clinical diagnostics. It is increasingly integrated into advanced therapeutic workflows, such as hybrid Positron Emission Tomography/Magnetic Resonance (PET/MR) imaging and MR-only radiation therapy. These integrated approaches are critically dependent on accurate estimation of radiation attenuation, which is typically facilitated by synthesizing Computed Tomography (CT) images from MR scans to generate attenuation maps. However, existing MR-to-CT synthesis methods for whole-body imaging often suffer from poor spatial alignment between the generated CT and input MR images, and insufficient image quality for reliable use in downstream clinical tasks. In this paper, we present a novel 3D Wavelet Latent Diffusion Model (3D-WLDM) that addresses these limitations by performing modality translation in a learned latent space. By incorporating a Wavelet Residual Module into the encoder-decoder architecture, we enhance the capture and reconstruction of fine-scale features across image and latent spaces. To preserve anatomical integrity during the diffusion process, we disentangle structural and modality-specific characteristics and anchor the structural component to prevent warping. We also introduce a Dual Skip Connection Attention mechanism within the diffusion model, enabling the generation of high-resolution CT images with improved representation of bony structures and soft-tissue contrast.
Figures
Figures from the paper (5 more)
Forward citations
Cited by 2 Pith papers
-
Knowledge-Guided 3D CT Generation: A Conditioning-Centric Taxonomy
A three-axis taxonomy (knowledge type, integration paradigm, architecture) for knowledge-guided 3D CT generation maps 25 methods and identifies geometric-mask-conditioned latent diffusion as the dominant paradigm.
-
Heterogeneity-Adaptive Diffusion Schrodinger Bridge for PET-Guided Whole-Body MRI Translation
HA-DSB uses a diffusion Schrödinger bridge with vision-language model region embeddings and PET-guided noise modulation to translate whole-body MRI while preserving lesion fidelity.
Reference graph
Works this paper leans on
-
[10]
Synthetic ct generation from mri using 3d transformer-based denoising diffusion model,
S. Pan, E. Abouei, J. Wynne, C.-W. Chang, T. Wang, R. L. Qiu, Y . Li, J. Peng, J. Roper, P. Patel et al. , “Synthetic ct generation from mri using 3d transformer-based denoising diffusion model,” Medical Physics, vol. 51, no. 4, pp. 2538–2548, 2024
work page 2024
-
[35]
Wdm: 3d wavelet diffusion models for high-resolution medical image synthesis,
P. Friedrich, J. Wolleb, F. Bieder, A. Durrer, and P. C. Cattin, “Wdm: 3d wavelet diffusion models for high-resolution medical image synthesis,” in MICCAI Workshop on Deep Generative Models. Springer, 2024, pp. 11–21
work page 2024
-
[36]
Mr to ct synthesis using 3d latent diffusion,
A. Tapp, A. Parida, C. Zhao, V . Lam, N. Lepore, S. M. Anwar, and M. G. Linguraru, “Mr to ct synthesis using 3d latent diffusion,” in 2024 IEEE International Symposium on Biomedical Imaging (ISBI) . IEEE, 2024, pp. 1–5
work page 2024
-
[1]
W. B. Overcast, K. M. Davis, C. Y . Ho, G. D. Hutchins, M. A. Green, B. D. Graner, and M. C. Veronesi, “Advanced imaging techniques for neuro-oncologic tumor diagnosis, with an emphasis on pet-mri imaging of malignant brain tumors,” Current Oncology Reports , vol. 23, pp. 1– 15, 2021
work page 2021
-
[2]
R. Ranjbarzadeh, A. Caputo, E. B. Tirkolaee, S. J. Ghoushchi, and M. Bendechache, “Brain tumor segmentation of mri images: A com- prehensive review on the application of artificial intelligence tools,” Computers in Biology and Medicine , vol. 152, p. 106405, 2023
work page 2023
-
[3]
Attenuation correction for human pet/mri studies,
C. Catana, “Attenuation correction for human pet/mri studies,” Physics in Medicine & Biology , vol. 65, no. 23, p. 23TR02, 2020
work page 2020
-
[4]
Mri-only treatment planning: benefits and challenges,
A. M. Owrangi, P. B. Greer, and C. K. Glide-Hurst, “Mri-only treatment planning: benefits and challenges,” Physics in Medicine & Biology , vol. 63, no. 5, p. 05TR01, 2018
work page 2018
-
[5]
S. K. Kang, H. J. An, H. Jin, J.-i. Kim, E. K. Chie, J. M. Park, and J. S. Lee, “Synthetic ct generation from weakly paired mr images using cycle- consistent gan for mr-guided radiotherapy,” Biomedical Engineering Letters, vol. 11, no. 3, pp. 263–271, 2021
work page 2021
Show all 43 references
-
[6]
Unsupervised mr-to-ct synthesis using structure-constrained cyclegan,
H. Yang, J. Sun, A. Carass, C. Zhao, J. Lee, J. L. Prince, and Z. Xu, “Unsupervised mr-to-ct synthesis using structure-constrained cyclegan,” IEEE Transactions on Medical Imaging, vol. 39, no. 12, pp. 4249–4261, 2020
2020
-
[7]
Deformable mr- ct image registration using an unsupervised, dual-channel network for neurosurgical guidance,
R. Han, C. K. Jones, J. Lee, P. Wu, P. Vagdargi, A. Uneri, P. A. Helm, M. Luciano, W. S. Anderson, and J. H. Siewerdsen, “Deformable mr- ct image registration using an unsupervised, dual-channel network for neurosurgical guidance,” Medical Image Analysis , vol. 75, p. 102292, 2022
2022
-
[8]
Ct synthesis from mri with an improved multi-scale learning network,
Y . Li, S. Xu, Y . Lu, and Z. Qi, “Ct synthesis from mri with an improved multi-scale learning network,” Frontiers in Physics, vol. 11, p. 1088899, 2023
2023
-
[9]
Liver synthetic ct generation based on a dense- cyclegan for mri-only treatment planning,
Y . Liu, Y . Lei, T. Wang, J. Zhou, L. Lin, T. Liu, P. Patel, W. J. Curran, L. Ren, and X. Yang, “Liver synthetic ct generation based on a dense- cyclegan for mri-only treatment planning,” in Medical Imaging 2020: Image Processing, vol. 11313. SPIE, 2020, pp. 659–664
2020
-
[11]
Freqgan: Infrared and visible image fusion via unified frequency adversarial learning,
Z. Wang, Z. Zhang, W. Qi, F. Yang, and J. Xu, “Freqgan: Infrared and visible image fusion via unified frequency adversarial learning,” IEEE Transactions on Circuits and Systems for Video Technology , 2024
2024
-
[12]
Low-light image enhancement with multi-scale attention and frequency-domain optimization,
Z. He, W. Ran, S. Liu, K. Li, J. Lu, C. Xie, Y . Liu, and H. Lu, “Low-light image enhancement with multi-scale attention and frequency-domain optimization,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 34, no. 4, pp. 2861–2875, 2023
2023
-
[13]
Unpaired whole-body mr to ct synthesis with correlation coefficient constrained adversarial learning,
Y . Ge, Z. Xue, T. Cao, and S. Liao, “Unpaired whole-body mr to ct synthesis with correlation coefficient constrained adversarial learning,” in Medical Imaging 2019: Image Processing , vol. 10949. SPIE, 2019, pp. 28–35
2019
-
[14]
Mutual information guided diffusion for zero-shot cross- modality medical image translation,
Z. Wang, Y . Yang, Y . Chen, T. Yuan, M. Sermesant, H. Delingette, and O. Wu, “Mutual information guided diffusion for zero-shot cross- modality medical image translation,” IEEE Transactions on Medical Imaging, vol. 43, no. 8, pp. 2825–2838, 2024
2024
-
[15]
Multi-modal modality- masked diffusion network for brain mri synthesis with random modality missing,
X. Meng, K. Sun, J. Xu, X. He, and D. Shen, “Multi-modal modality- masked diffusion network for brain mri synthesis with random modality missing,” IEEE Transactions on Medical Imaging , vol. 43, no. 7, pp. 2587–2598, 2024
2024
-
[16]
Synthesizing pet images from high-field and ultra-high- field mr images using joint diffusion attention model,
T. Xie, C. Cao, Z.-x. Cui, Y . Guo, C. Wu, X. Wang, Q. Li, Z. Hu, T. Sun, Z. Sang et al., “Synthesizing pet images from high-field and ultra-high- field mr images using joint diffusion attention model,” Medical Physics, vol. 51, no. 8, pp. 5250–5269, 2024
2024
-
[17]
Towards extreme image compression with latent feature guidance and diffusion prior,
Z. Li, Y . Zhou, H. Wei, C. Ge, and J. Jiang, “Towards extreme image compression with latent feature guidance and diffusion prior,” IEEE Transactions on Circuits and Systems for Video Technology , 2024
2024
-
[18]
Synthetic ct-aided mri-ct image registration for head and neck radiotherapy,
Y . Fu, Y . Lei, J. Zhou, T. Wang, S. Y . David, J. J. Beitler, W. J. Curran, T. Liu, and X. Yang, “Synthetic ct-aided mri-ct image registration for head and neck radiotherapy,” in Medical Imaging 2020: Biomedical Applications in Molecular, Structural, and Functional Imaging ,...
2020
-
[19]
Resvit: residual vision transformers for multimodal medical image synthesis,
O. Dalmaz, M. Yurt, and T. C ¸ ukur, “Resvit: residual vision transformers for multimodal medical image synthesis,” IEEE Transactions on Medical Imaging, vol. 41, no. 10, pp. 2598–2614, 2022
2022
-
[20]
Deep learning-enabled mri-only photon and proton therapy treatment planning for paediatric abdominal tumours,
M. C. Florkow, F. Guerreiro, F. Zijlstra, E. Seravalli, G. O. Janssens, J. H. Maduro, A. C. Knopf, R. M. Castelein, M. van Stralen, B. W. Raaymakers et al., “Deep learning-enabled mri-only photon and proton therapy treatment planning for paediatric abdominal tumours,” Radio- t...
2020
-
[21]
Imitation learning for improved 3d pet/mr attenuation correction,
K. Kl ¨aser, T. Varsavsky, P. Markiewicz, T. Vercauteren, A. Hammers, D. Atkinson, K. Thielemans, B. Hutton, M. J. Cardoso, and S. Ourselin, “Imitation learning for improved 3d pet/mr attenuation correction,” Medical Image Analysis , vol. 71, p. 102079, 2021
2021
-
[22]
Synthesis of pseudo-ct images from pelvic mri images based on an md-cyclegan model for radiotherapy,
H. Sun, Q. Xi, R. Fan, J. Sun, K. Xie, X. Ni, and J. Yang, “Synthesis of pseudo-ct images from pelvic mri images based on an md-cyclegan model for radiotherapy,” Physics in Medicine & Biology , vol. 67, no. 3, p. 035006, 2022
2022
-
[23]
A deep learning approach to generate synthetic ct in low field mr-guided adaptive radiotherapy for abdominal and pelvic cases,
D. Cusumano, J. Lenkowicz, C. V otta, L. Boldrini, L. Placidi, F. Catucci, N. Dinapoli, M. V . Antonelli, A. Romano, V . De Luca et al. , “A deep learning approach to generate synthetic ct in low field mr-guided adaptive radiotherapy for abdominal and pelvic cases,” Radiothera...
2020
-
[24]
Improving generalization in mr-to-ct synthesis in radiotherapy by using an augmented cycle generative adversarial network with unpaired data,
K. N. Brou Boni, J. Klein, A. Gulyban, N. Reynaert, and D. Pasquier, “Improving generalization in mr-to-ct synthesis in radiotherapy by using an augmented cycle generative adversarial network with unpaired data,” Medical physics, vol. 48, no. 6, pp. 3003–3010, 2021
2021
-
[25]
Medical image synthesis with deep convolutional adversarial networks,
D. Nie, R. Trullo, J. Lian, L. Wang, C. Petitjean, S. Ruan, Q. Wang, and D. Shen, “Medical image synthesis with deep convolutional adversarial networks,” IEEE Transactions on Biomedical Engineering , vol. 65, no. 12, pp. 2720–2730, 2018
2018
-
[26]
Frequency-supervised mr- to-ct image synthesis,
Z. Shi, P. Mettes, G. Zheng, and C. Snoek, “Frequency-supervised mr- to-ct image synthesis,” in Deep Generative Models, and Data Augmen- tation, Labelling, and Imperfections . Springer, 2021, pp. 3–13
2021
-
[27]
Grad-cam guided u-net for mri-based pseudo-ct synthesis,
G. Dovletov, D. D. Pham, S. L ¨orcks, J. Pauli, M. Gratz, and H. H. Quick, “Grad-cam guided u-net for mri-based pseudo-ct synthesis,” in 2022 44th Annual International Conference of the IEEE Engineering in Medicine & Biology Society (EMBC) . IEEE, 2022, pp. 2071–2075. 12
2022
-
[28]
Double grad-cam guidance for improved mri-based pseudo-ct synthesis,
G. Dovletov, S. L ¨orcks, J. Pauli, M. Gratz, and H. H. Quick, “Double grad-cam guidance for improved mri-based pseudo-ct synthesis,” inBVM Workshop. Springer, 2023, pp. 45–50
2023
-
[29]
Ct synthesis from mr in the pelvic area using residual transformer conditional gan,
B. Zhao, T. Cheng, X. Zhang, J. Wang, H. Zhu, R. Zhao, D. Li, Z. Zhang, and G. Yu, “Ct synthesis from mr in the pelvic area using residual transformer conditional gan,” Computerized Medical Imaging and Graphics, vol. 103, p. 102150, 2023
2023
-
[30]
Multi-scale tokens-aware transformer network for multi-region and multi-sequence mr-to-ct synthesis in a single model,
L. Zhong, Z. Chen, H. Shu, K. Zheng, Y . Li, W. Chen, Y . Wu, J. Ma, Q. Feng, and W. Yang, “Multi-scale tokens-aware transformer network for multi-region and multi-sequence mr-to-ct synthesis in a single model,” IEEE Transactions on Medical Imaging , vol. 43, no. 2, pp. 794–806, 2023
2023
-
[31]
Palette: Image-to-image diffusion models,
C. Saharia, W. Chan, H. Chang, C. Lee, J. Ho, T. Salimans, D. Fleet, and M. Norouzi, “Palette: Image-to-image diffusion models,” in ACM SIGGRAPH 2022 Conference Proceedings , 2022, pp. 1–10
2022
-
[32]
Toward real-world blind face restoration with generative diffusion prior,
X. Chen, J. Tan, T. Wang, K. Zhang, W. Luo, and X. Cao, “Toward real-world blind face restoration with generative diffusion prior,” IEEE Transactions on Circuits and Systems for Video Technology , vol. 34, no. 9, pp. 8494–8508, 2024
2024
-
[33]
Stylediffusion: Controllable dis- entangled style transfer via diffusion models,
Z. Wang, L. Zhao, and W. Xing, “Stylediffusion: Controllable dis- entangled style transfer via diffusion models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 7677–7689
2023
-
[34]
Adding conditional control to text-to-image diffusion models,
L. Zhang, A. Rao, and M. Agrawala, “Adding conditional control to text-to-image diffusion models,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , 2023, pp. 3836–3847
2023
-
[37]
High- resolution image synthesis with latent diffusion models,
R. Rombach, A. Blattmann, D. Lorenz, P. Esser, and B. Ommer, “High- resolution image synthesis with latent diffusion models,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogni- tion, 2022, pp. 10 684–10 695
2022
-
[38]
Cola-diff: Conditional latent diffusion model for multi-modal mri synthesis,
L. Jiang, Y . Mao, X. Wang, X. Chen, and C. Li, “Cola-diff: Conditional latent diffusion model for multi-modal mri synthesis,” in International Conference on Medical Image Computing and Computer-Assisted Inter- vention. Springer, 2023, pp. 398–408
2023
-
[39]
Denoising diffusion implicit models,
J. Song, C. Meng, and S. Ermon, “Denoising diffusion implicit models,” in International Conference on Learning Representations , 2021. [Online]. Available: https://openreview.net/forum?id=St1giarCHLP
2021
-
[40]
Elastix: a toolbox for intensity-based medical image registration,
S. Klein, M. Staring, K. Murphy, M. A. Viergever, and J. P. Pluim, “Elastix: a toolbox for intensity-based medical image registration,” IEEE Transactions on Medical Imaging , vol. 29, no. 1, pp. 196–205, 2009
2009
-
[41]
Denoising diffusion probabilistic models,
J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in Neural Information Processing Systems , vol. 33, pp. 6840– 6851, 2020
2020
-
[42]
Unpaired image-to-image translation using cycle-consistent adversarial networks,
J.-Y . Zhu, T. Park, P. Isola, and A. A. Efros, “Unpaired image-to-image translation using cycle-consistent adversarial networks,” in Proceedings of the IEEE International Conference on Computer Vision , 2017, pp. 2223–2232
2017
-
[43]
Totalsegmentator: robust segmentation of 104 anatomic structures in ct images,
J. Wasserthal, H.-C. Breit, M. T. Meyer, M. Pradella, D. Hinck, A. W. Sauter, T. Heye, D. T. Boll, J. Cyriac, S. Yang et al., “Totalsegmentator: robust segmentation of 104 anatomic structures in ct images,”Radiology: Artificial Intelligence, vol. 5, no. 5, 2023
2023
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.