{"id":"e247f606-e4ea-4dd1-849d-77f24c9e5301","arxiv_id":"2608.08819","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"A diffusion bridge model, instantiated from the residual diffusion bridge framework, reconstructs high-resolution MRI in ten deterministic sampling steps and outperforms nine baselines on PSNR, SSIM, and GMSD on 7T brain and prostate data.","lead":"MRI super-resolution usually needs many diffusion sampling steps and starts from noise. This paper applies an existing diffusion-bridge framework to MRI, pinning the reconstruction to the low-resolution input and recovering high-resolution images in ten sampling steps, and reports the best fidelity scores on 7T brain and prostate datasets against nine baselines.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Statistical significance may rest on slice-level Wilcoxon tests ignoring within-patient correlation; patient-level analysis could overturn the claimed superiority over all baselines.","rationale":"After checking the derivation of the bridge marginal (Eq. 3) and the deterministic sampler (Eq. 8), I find the mathematical construction internally consistent; the forward bridge mean and variance match a standard h-transformed OU process, and the update rule is a valid DDIM-type inversion. The paper's core risk is not the math but the statistical evidence for the headline \"significant gains over every comparison method.\" The Methods section omits the analysis unit, and the claimed contribution of \"per-patient paired statistics and effect sizes\" is not delivered. If the tests were performed per slice, correlated observations violate the signed-rank test's independence assumption, making p-values unreliable. This is a concrete, checkable threat to the central claim. The reader's weakest assumption (synthetic degradation) is a real limitation for clinical generalization, but it does not speak to whether the benchmark results are correctly established; I therefore focus on the statistical unit as the more load-bearing concern. Verdict should remain conditional: acceptance requires clarifying the analysis unit and demonstrating robustness at the patient level.","tokens_in":13565,"tokens_out":10490,"duration_ms":101159,"concrete_test":"Recompute the Wilcoxon signed-rank tests with the patient as the statistical unit: for each patient, average or median the slice-level PSNR, SSIM, and GMSD for each method, then run paired Wilcoxon tests on the 21 brain patients and 66 prostate patients, applying Holm correction within each metric. Also report matched rank-biserial correlation or median difference with confidence intervals as an effect size. If any of the nine comparisons loses significance at p<0.05, the claim of significant gains over every comparison method is not supported as stated. The authors should also disclose whether the original tests were per-slice or per-patient and justify the independence assumption.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claim—statistically significant gains over every comparison method on PSNR, SSIM, and GMSD—depends on the unit of the Wilcoxon signed-rank test, which Section 2.7 never specifies. The contribution bullet promises \"per-patient paired statistics and effect sizes,\" but no effect sizes appear in the paper and the reported means and standard deviations span thousands of slices from only 21 brain patients (2,552 slices) and 66 prostate patients (2,668 slices). If the test was run per slice, adjacent slices from the same patient are treated as independent replicates, inflating the effective sample size by roughly 120-fold (brain) and 40-fold (prostate). With n≈2,500, a mean PSNR gap of less than 0.1 dB can reach p<0.05, so the reported significance against the closest baseline SR-EMamba (0.76 dB brain, 0.72 dB prostate) may not survive a patient-level test. If the test was actually per patient, the paper should state this and report effect sizes; as written, the abstract's \"statistically significant gains over every comparison method\" is not verifiable. This concern is independent of the external-validity question about synthetic downsampling and is more load-bearing because it directly supports the headline claim.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes SR-DBM, a diffusion bridge model for MRI super-resolution. The forward process is a mean-reverting Ornstein-Uhlenbeck SDE whose stationary mean is the low-resolution image, conditioned via a Doob h-transform to start at the high-resolution image and end at the low-resolution image. The reverse process is trained with a combined L1 and LPIPS objective and reconstructs the HR image with a deterministic ten-step sampler. The method is evaluated on 7T brain T1 MP2RAGE maps and ProstateX T2-weighted prostate images against nine baselines, reporting PSNR, SSIM, GMSD, and LPIPS. The reported results show the best PSNR, SSIM, and GMSD on both datasets, while LPIPS is not the best on either dataset. The central claim is that the PSNR, SSIM, and GMSD gains are statistically significant relative to all comparison methods.","tokens_in":13843,"tokens_out":13628,"duration_ms":146603,"significance":"If the quantitative claims hold, SR-DBM is a useful and practical instantiation of a continuous diffusion bridge for MRI super-resolution: it requires only ten network evaluations, the bridge coefficients are given in closed form, and the training and sampling algorithms are clearly specified. The evaluation on two anatomies and contrasts against nine diverse baselines is a solid empirical contribution to the medical imaging super-resolution literature. At the same time, the significance is currently constrained by the unspecified statistical unit of the Wilcoxon tests, the absence of baseline training protocols and effect sizes, the unfulfilled promised analysis of step-count trade-offs, and the gap between the clinical significance statement and the synthetic downsampling evaluation. The theoretical construction largely follows an existing external framework ([17]); this is acceptable for an application paper but should be framed accordingly.","major_comments":[{"comment":"The unit of the paired Wilcoxon signed-rank test is never specified, despite the contribution bullet promising 'per-patient paired statistics and effect sizes.' The test set comprises 2,552 brain slices from 21 patients and 2,668 prostate slices from 66 patients; if slices are treated as independent observations, the effective sample size is inflated by roughly 120-fold (brain) and 40-fold (prostate), and the reported mean PSNR gap of 0.76 dB (brain) and 0.72 dB (prostate) over SR-EMamba may become statistically significant even when patient-level differences are not. Since the abstract's headline claim that SR-DBM achieved 'statistically significant gains over every comparison method' rests entirely on these tests, the authors must state the unit of analysis, report patient-level paired statistics (e.g., per-patient median slice metrics) and effect sizes, or justify slice-level independence. As written, the central quantitative claim is not verifiable.","section":"Section 2.7, Tables 2-3, and Section 1 contribution bullet"},{"comment":"No training or implementation details are given for the nine comparison methods. The paper does not state which official implementations were used, how hyperparameters were chosen, whether each baseline was retrained on the same downsampled data and the same train/test split, or for how many epochs. Unequal training effort is a standard confound in head-to-head super-resolution comparisons; without these details, the reported superiority over baselines such as SR-EMamba, SwinIR, and Res-SRDiff cannot be independently reproduced or assessed as fair.","section":"Section 2.7 and Section 3"},{"comment":"The low-resolution inputs are produced by synthetic downsampling of fully sampled high-resolution images (4x in-plane for the brain data; 9x in-plane and 2x through-plane for the prostate data). This models only a limited form of resolution loss and does not include the noise, motion, slice profile, or aliasing that characterize true rapid MRI acquisitions. The Discussion and Significance statements that SR-DBM may 'reduce scan time' and 'support improved lesion delineation' therefore extrapolate beyond the evaluation presented. The authors should temper these claims or add validation on prospectively acquired low-resolution data.","section":"Section 2.6 and Section 4"},{"comment":"The introduction claims the paper will 'characterise reconstruction quality and inference cost as a function of the number of sampling steps,' but the manuscript reports only T=100 and S=10 and contains no ablation or cost analysis over S. The Discussion lists such a characterization as future work, which contradicts the contribution statement. The authors should either add this analysis or remove the claim from the contribution list.","section":"Section 1 contribution bullet 2 and Section 4"}],"minor_comments":[{"comment":"Table 3 reports SSIM values in percentage-like units (e.g., 79.51±4.74 for the proposed method), whereas Table 2 reports SSIM as fractions (0.96±0.02). The scale should be unified or explicitly labeled to avoid confusion.","section":"Tables 2 and 3"},{"comment":"The phrase 'statistically significant gains over every comparison method' is ambiguous, since on the brain LPIPS comparison the difference from Pix2pix is not significant (p=0.71) and several methods have significantly lower LPIPS than SR-DBM. Please state explicitly that the significance claim applies to PSNR, SSIM, and GMSD rather than to all reported metrics.","section":"Abstract and Section 3.1"},{"comment":"The downsampling procedure is not fully specified: the paper does not state the interpolation kernel, whether anti-aliasing filtering was applied, or whether the same kernel was used for both spatial axes. This information is needed to reproduce the exact low-resolution inputs.","section":"Section 2.6"},{"comment":"The reduction of the KL divergence to the L1 plus LPIPS objective in Eq. (7) is stated in a single sentence ('after dropping the time-dependent weights'); a short derivation or a reference to the corresponding derivation in [17] would allow readers to verify the objective.","section":"Section 2.2.2"},{"comment":"Although the title emphasizes ten sampling steps, no wall-clock inference time or computational cost comparison with the baselines is reported. The efficiency claim would be strengthened by reporting runtime per volume or per slice.","section":"Section 2.5 and Section 3"}],"recommendation":"major_revision","confidential_remarks":"The main barrier to acceptance is the statistical analysis unit, which is fixable by re-analysis. I would encourage the editor to ask for a code/data-sharing statement and the exact list of baseline implementations, since the institutional dataset is not public. The reference list contains many self-citations, but the closest external prior work (RDBM [17]) is cited appropriately."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"First thing you should know: this is a competent application of the residual diffusion bridge framework to MRI super-resolution, and the 10-step deterministic sampler is a nice efficiency contribution. But the paper's central claim—statistically significant gains over every baseline on PSNR/SSIM/GMSD—rests on a statistical test whose unit of analysis is never stated. If the Wilcoxon was run per slice, the significance is likely an artifact of within-patient correlation, and the headline will not survive a patient-level analysis.\n\nWhat is genuinely good: the math is clean. The bridge coefficients from RDBM are restated correctly, the deterministic sampler follows the usual DDIM update, and the training objective is simple. The empirical comparison includes nine baselines on two datasets, and the reported gains over SR-EMamba are plausible on synthetic downsampling. The authors are also honest that they do not win on LPIPS. Hyperparameters are fixed, so there is no sense of fitting to produce the reported metrics.\n\nThe soft spots are real but not all equally load-bearing. The statistical unit is the big one. The contributions bullet promises 'per-patient paired statistics and effect sizes,' but the methods section never says per-patient and no effect sizes appear. Slices are not independent: 2,552 brain slices come from 21 patients. A per-slice Wilcoxon with n≈2,500 will turn a 0.76 dB PSNR gap into p<0.05 even if the true per-patient effect is negligible. This needs to be fixed before the claim is credible. The step-sweep characterization promised in the contributions is absent; the discussion even lists it as future work. Baseline training protocols are not described, so we don't know whether all methods got equivalent training budgets. And the evaluation is entirely on synthetically downsampled images; real fast acquisitions may have different noise and slice-profile effects. Code and the institutional data are unavailable, so reproducibility is limited.\n\nWho is this for? Researchers working on efficient diffusion-based SR for medical images. The method itself is worth knowing about. I would send this to peer review, but I would ask for major revision: specify the test unit, report patient-level statistics and effect sizes, describe baseline training, and either add the step sweep or delete that contribution claim. If the patient-level analysis overturns the significance, the paper can still be a useful efficiency study, but the abstract needs to be toned down.","headline":"A clean application of RDBM to MRI super-resolution with a promising 10-step deterministic sampler, but the headline significance claim rests on an unspecified statistical unit that could invalidate it.","tokens_in":14395,"tokens_out":3318,"would_cite":false,"duration_ms":33994,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"SR-DBM, a diffusion bridge model, reconstructs high-resolution MRI from low-resolution input in ten sampling steps, and on 7T brain T1 and prostate T2-weighted images it reports the best PSNR, SSIM, and GMSD among ten methods.","keywords":["MRI super-resolution","diffusion bridge model","Doob's h-transform","mean-reverting Ornstein-Uhlenbeck process","ultra-high-field 7T brain T1 mapping","prostate T2-weighted imaging","ten-step sampling","image restoration"],"falsifier":"Scan the same patients twice, once with a fast low-resolution protocol and once with a slow high-resolution protocol, and check whether SR-DBM still beats the nine baselines on PSNR, SSIM, and GMSD on the real paired data; if the advantage shrinks or reverses, the central claim is an artifact of synthetic downsampling.","tokens_in":13404,"feed_emoji":"🧠","tokens_out":6876,"duration_ms":65230,"temperature":0.7,"pith_summary":"This paper tries to show that MRI super-resolution can be made both fast and accurate by treating it as a stochastic bridge from the low-resolution measurement to the high-resolution target, rather than as denoising from Gaussian noise. The proposed model, SR-DBM, anchors a mean-reverting diffusion process at the paired low- and high-resolution images through Doob's h-transform, so reconstruction starts from measured anatomy and needs only ten sampling steps. On 7T brain T1 and pelvic T2-weighted prostate images, the paper reports the highest PSNR and SSIM and the lowest GMSD among ten methods, with statistically significant gains over every comparison; SR-EMamba is the closest runner-up. If this holds, the practical payoff is that faster low-resolution MRI acquisitions could be converted to high-resolution images without the usual computational cost of diffusion models.","feed_headline":"Ten-step diffusion bridge beats nine MRI super-resolution baselines","feed_subtitle":"Bridge-pinned diffusion beats nine baselines on 7T brain and prostate MRI in ten steps.","key_machinery":"The central object is the h-transformed mean-reverting OU bridge with coefficients $\\psi_t=\\sinh(S_T-S_t)/\\sinh(S_T)$ and $\\omega_t^2=2\\nu\\sinh(S_t)\\sinh(S_T-S_t)/\\sinh(S_T)$, using the cumulative drift $S_t$ from a cosine schedule. The mean coefficient interpolates from 1 to 0, and the noise scale vanishes at both endpoints; this is what lets the reverse trajectory start at $x_{\\mathrm{LR}}$ and end at the predicted $\\hat{x}_{\\mathrm{HR}}$ in closed-form updates. The network's job is to estimate the residual $\\hat{x}_{\\mathrm{HR}}-x_{\\mathrm{LR}}$ rather than the absolute image.","core_discovery":"SR-DBM formulates super-resolution as continuous-time stochastic transport between the HR image at time zero and the LR image at the terminal time. The forward process is a mean-reverting Ornstein-Uhlenbeck SDE whose stationary mean is the LR image, conditioned by Doob's h-transform to end exactly at the LR image; its marginals are Gaussian with mean $x_{\\mathrm{LR}}+\\psi_t(x_{\\mathrm{HR}}-x_{\\mathrm{LR}})$ and variance $\\omega_t^2 I$, so intermediate states are convex interpolations plus controlled noise. A network trained to predict the clean HR image under an $\\ell_1$ plus LPIPS objective drives a deterministic reverse trajectory over ten strided steps. On the two datasets the paper reports top means for PSNR (brain $27.66\\pm1.52$ dB, prostate $27.87\\pm2.29$ dB), SSIM (0.96, 0.80), and GMSD (7.96%, 8.38%), each statistically significant versus all nine baselines, while on LPIPS the model ranks behind SR-EMamba and a few others; the paper attributes this to a perception-distortion trade-off.","pith_inferences":["Our inference: the same h-transformed bridge construction should transfer to other paired inverse problems in medical imaging (for example, PET-to-CT synthesis or artifact removal) whenever a low-quality input can be pinned as the terminal state, although the paper only tests MRI super-resolution.","Our inference: because the method ranks lower on LPIPS while winning on PSNR, SSIM, and GMSD, combining the bridge with a perceptual or adversarial term could yield a single model that leads on all four metrics; the paper mentions adding perceptual objectives but does not test this.","Our inference: the deterministic ten-step sampler is the key practical claim; comparing the reconstruction at S=1, 2, 5, 10, and 20 steps would directly map the cost-quality curve, something the paper only partially characterizes."],"forward_implications":["Diffusion-based MRI super-resolution can be run in ten network evaluations, cutting sampling cost relative to methods that need hundreds of steps.","Because reconstruction starts from the measured low-resolution image rather than Gaussian noise, fidelity-oriented restoration is more direct.","The continuous-time bridge generalizes residual-shifting diffusion, so improvements carry over to that family of image-restoration models.","Clinically, the results support using faster low-resolution acquisitions followed by SR to preserve fine structures and lesions in brain and prostate imaging, with reduced scan time.","Perceptual quality remains a separate axis: LPIPS scores trail some baselines, and the paper's own discussion points to perceptual objectives as future work."],"supporting_citations":[{"why":"Supplies the generalised residual diffusion bridge that SR-DBM instantiates for MRI super-resolution.","marker":"[17]"},{"why":"Defines residual-shifting diffusion, which the paper treats as a discrete-time special case of its continuous bridge and uses as background.","marker":"[15]"},{"why":"Prior MRI super-resolution diffusion baseline whose residual-shifting scheme and preprocessing the paper extends.","marker":"[14]"},{"why":"Image-to-image Schrödinger bridge baseline that also connects LR and HR distributions but is outperformed on the reported metrics.","marker":"[16]"},{"why":"SR-EMamba, the strongest competing baseline on PSNR, SSIM, and GMSD, and source of the two-cohort preprocessing pipeline.","marker":"[23]"},{"why":"Provides the cosine noise schedule reused to set the OU drift schedule in SR-DBM.","marker":"[19]"},{"why":"Institutional 7T brain T1 MP2RAGE dataset on which the brain results are measured.","marker":"[20]"},{"why":"Public ProstateX T2-weighted dataset on which the prostate results are measured.","marker":"[21]"},{"why":"Defines SSIM, one of the three metrics on which the central claim of superiority rests.","marker":"[26]"},{"why":"Defines GMSD, the gradient-structure metric on which SR-DBM ranks first.","marker":"[27]"}],"fun_headline_variants":["Ten-step diffusion bridge beats nine MRI SR baselines","Diffusion bridge super-resolves MRI in ten steps","Bridge-pinned diffusion tops MRI SR on PSNR and SSIM","Ten-step MRI super-resolution via diffusion bridge"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The argument leans on the assumption that low-resolution images produced by downsampling full-resolution scans capture what a real fast MRI acquisition produces, including its noise and artifacts.","fun_headline_variants_meta":{"raw":{"variants":["Ten-step diffusion bridge beats nine MRI SR baselines","Diffusion bridge super-resolves MRI in ten steps","Bridge-pinned diffusion tops MRI SR on PSNR and SSIM","Ten-step MRI super-resolution via diffusion bridge"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000929,"raw_usage":{"total_tokens":4079,"prompt_tokens":1147,"completion_tokens":2932,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":763,"completion_tokens_details":{"reasoning_tokens":2867}},"tokens_in":763,"tokens_out":2932,"duration_ms":27244,"temperature":1.0,"reasoning_tokens":2867,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:23:33.186603+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Scan the same patients twice, once with a fast low-resolution protocol and once with a slow high-resolution protocol, and check whether SR-DBM still beats the nine baselines on PSNR, SSIM, and GMSD on the real paired data; if the advantage shrinks or reverses, the central claim is an artifact of synthetic downsampling.","supporting_citations":[{"cited_title":"Residual diffusion bridge model for image restoration","cited_arxiv_id":null,"evidence_quote":"Supplies the generalised residual diffusion bridge that SR-DBM instantiates for MRI super-resolution."},{"cited_title":"Eﬀicient diffusion model for image restoration by residual shifting","cited_arxiv_id":null,"evidence_quote":"Defines residual-shifting diffusion, which the paper treats as a discrete-time special case of its continuous bridge and uses as background."},{"cited_title":"Mri super-resolution reconstruc- 29 tion using eﬀicient diffusion probabilistic model with residual shifting","cited_arxiv_id":null,"evidence_quote":"Prior MRI super-resolution diffusion baseline whose residual-shifting scheme and preprocessing the paper extends."},{"cited_title":"Eﬀicient vision mamba for mri super-resolution via hybrid selective scanning","cited_arxiv_id":null,"evidence_quote":"SR-EMamba, the strongest competing baseline on PSNR, SSIM, and GMSD, and source of the two-cohort preprocessing pipeline."},{"cited_title":"Improved denoising diffusion probabilistic models","cited_arxiv_id":null,"evidence_quote":"Provides the cosine noise schedule reused to set the OU drift schedule in SR-DBM."},{"cited_title":"7 t lesion-attenuated magnetization- 30 prepared gradient echo acquisition for detection of posterior fossa demyelinating lesions in multiple sclerosis","cited_arxiv_id":null,"evidence_quote":"Institutional 7T brain T1 MP2RAGE dataset on which the brain results are measured."},{"cited_title":"Prostatex challenges for computerized classification of prostate lesions from multiparametric magnetic resonance images","cited_arxiv_id":null,"evidence_quote":"Public ProstateX T2-weighted dataset on which the prostate results are measured."},{"cited_title":"Image qual- ity assessment: from error visibility to structural similarity","cited_arxiv_id":null,"evidence_quote":"Defines SSIM, one of the three metrics on which the central claim of superiority rests."},{"cited_title":"Gradient magnitude similarity deviation: A highly eﬀicient perceptual image quality index","cited_arxiv_id":null,"evidence_quote":"Defines GMSD, the gradient-structure metric on which SR-DBM ranks first."}],"review_version":1}