Pith. sign in

REVIEW 5 major objections 5 minor 49 references

MeanSR: Restoration Trajectory Learning for One-Step Perceptual Super-Resolution

T0 review · 5 major / 5 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read A one-step super-resolution model that learns the average velocity of the LR-to-HR restoration path claims better perceptual quality than its closest one-step rival at a fraction of the compute.

desk verdict A plausible one-step SR method whose claimed win over CTMSR rests on a single run and no-reference metrics; the method itself is coherent and worth a careful look. read the letter →

arxiv 2608.09405 v1 pith:XXIVZEO2 submitted 2026-08-10 cs.CV

classification cs.CV
keywords one-stepsuper-resolutionperceptualaveragevelocityfieldMeanFlowdistributiontrajectorymatchingstage-awaretemporalsamplingreal-worldimage
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

MeanSR tries to make single-image super-resolution both perceptual and fast by learning, in a single network, the average velocity field that carries a degraded and noisy latent image to a credible high-resolution output. The authors argue that this explicit modeling of the restoration trajectory is what one-step methods like CTMSR lack, and they show that adding distribution trajectory matching on top of the learned field, with stage-specific temporal sampling, improves no-reference perceptual scores (CLIPIQA, MUSIQ, MANIQA) on synthetic and real-world benchmarks while cutting FLOPs by about six times and inference latency by roughly half compared with CTMSR. If the claims hold, MeanSR would be a state-of-the-art one-step perceptual SR method that does not need a pretrained diffusion teacher.

What carries the argument

The central object is the learned LR-conditioned average velocity field $u(z_t, c_{LR}, r, t)$, the time-averaged transport velocity between two latent states on a flow-matching path; it is the quantity that collapses the whole iterative denoising chain into one network evaluation. The paper inherits it from MeanFlow, which expresses the average velocity from an instantaneous velocity $v(z_\tau, c_{LR}, \tau) = \epsilon - z_0$, and the training objective regresses the network output to a stop-gradient target that includes a Jacobian-vector-product term $\frac{d}{dt}u_\theta$, computed with automatic differentiation. This same field drives the second stage, where generated and ground-truth trajectories are perturbed in tandem and compared through a delayed target network, with the discrepancy used to build a pseudo-target for LPIPS supervision.

What would settle it

Re-run both MeanSR and CTMSR with several random seeds and compare the distributions of CLIPIQA/MUSIQ/MANIQA on RealSR, RealSet65, and a held-out real-world set; if the inter-seed spread covers the reported gaps, the claimed advantage is not reproducible. Also run a blind human preference study between single-step outputs; if raters do not prefer MeanSR, the metric gains are not perceptual gains.

Watch

Extended reading notes

Core claim

MeanSR's central claim is that the finite-time transition from a low-resolution input to a high-resolution image can be captured explicitly as an LR-conditioned average velocity, so that a single step $z_{SR} = z_1 - u(z_1, c_{LR}, 0, 1)$ produces the SR latent directly from Gaussian noise. The paper derives this velocity from the MeanFlow identity $u(z_t, c_{LR}, r, t) = v(z_t, c_{LR}, t) - (t-r)\frac{d}{dt}u(\cdot)$, learns it with a stop-gradient regression, and then reuses the same field to align the generated trajectory with the ground-truth HR trajectory through Distribution Trajectory Matching in image space. With a logit-normal temporal distribution for velocity estimation and uniform sampling for DTM, the method claims state-of-the-art perceptual quality among one-step SR methods on RealSR, RealSet65, and ImageNet-Test, improving MANIQA over CTMSR by 19.9% and 50.0% on RealSR and RealSet65 respectively while using roughly 6x less FLOPs and 2x less runtime.

Load-bearing premise

The entire superiority claim rests on no-reference perceptual metrics (CLIPIQA, MUSIQ, MANIQA) and a single unrepeated training run, with temporal-sampling choices tuned on the same ImageNet-Test benchmark; if these metrics do not track human perception or the reported gaps are within run-to-run noise, the advantage over CTMSR collapses.

Editorial extensions

If this is right

  • One-step SR with a single network evaluation: sampling is $z_{SR} = z_1 - u(z_1, c_{LR}, 0, 1)$, so no iterative denoising or numerical ODE solving is needed.
  • No dependence on pretrained diffusion teachers or distillation, unlike SinSR or OSEDiffR; training starts from synthetic LR-HR pairs and a VAE.
  • Substantially lower compute: about 6x fewer FLOPs and 2x lower latency than CTMSR while improving all three no-reference perceptual metrics on both real-world benchmarks.
  • Faster training: roughly 13x fewer iterations to reach the same CLIPIQA level as CTMSR on ImageNet.
  • Stage-aware temporal sampling is a general training recipe: logit-normal sampling for velocity estimation and uniform sampling for trajectory alignment.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The same LR-conditioned average-velocity objective could transfer to other restoration tasks (denoising, deblurring, face enhancement) because it only requires defining a condition and a target distribution, though the paper does not test this.
  • Because the main comparison is entirely metric-based, a forced-choice human preference study on RealSR and RealSet65 outputs would be a direct check of whether the higher CLIPIQA/MUSIQ/MANIQA numbers translate into visible perceptual gains.
  • The stage-2 hyper-parameters (10k iterations, logit-normal vs uniform sampling) were selected on the same ImageNet-Test used for final numbers, so the reported gains could partly reflect benchmark overfitting; evaluating on a freshly degraded hold-out set would quantify this.
  • The large MANIQA improvements (19.9% on RealSR, 50.0% on RealSet65) may be inflated by metric sensitivity; pairwise perceptual tests with human raters would show whether the perceived gain is as large as the numbers suggest.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper proposes MeanSR, a one-step perceptual super-resolution method that learns an LR-conditioned average velocity field using a MeanFlow-style formulation, combines it with a reformulated Distribution Trajectory Matching (DTM) objective, and introduces a Stage-Aware Temporal Sampling (SATS) strategy. The method is trained in two stages: first to estimate the average velocity from degraded/noisy latents to HR latents, and second to align the generated trajectory with the target HR trajectory. The authors evaluate on synthetic ImageNet-Test and real-world RealSR/RealSet65, reporting better CLIPIQA, MUSIQ, and MANIQA scores than CTMSR and other one-step methods, with substantially lower FLOPs and inference latency.

Significance. If the empirical claims hold, MeanSR would be a valuable contribution: a distillation-free one-step SR method that explicitly models restoration dynamics and achieves state-of-the-art perceptual quality with a 6x FLOPs reduction relative to CTMSR. The conceptual shift from consistency training to average-velocity trajectory learning is interesting, and the two-stage design with stage-specific temporal sampling is a reasonable idea. However, the current evidence is weakened by the absence of variance information, the use of the test set for hyperparameter selection, contradictory perceptual metrics (worse LPIPS), missing fidelity metrics, and an incompletely explained efficiency comparison. The central mathematical derivation in Eq. (4) is correct, and the paper clearly describes the method, but the empirical claims are not yet fully supported.

major comments (5)
  1. [Section 3, Eq. (5) and Algorithm 1] The MeanSR training target is self-referential: u_tgt = v - (t-r) * d/dt u_theta uses the network's own Jacobian (via the JVP in Algorithm 1). This makes the loss a fixed-point equation rather than a regression to an external target. For any function satisfying the ODE u = v - (t-r)u' along the sampled path, the loss vanishes, so without a boundary condition (e.g., enforcing u(r) = v(r)) the objective may have many degenerate solutions. While the additional l1 and VGG losses on x_hat_0 provide some grounding, the interpretation of loss_mf as matching the true average velocity is not established. The authors should either provide a convergence/stability argument for this fixed-point objective or replace the target with a network-independent estimate, such as the Monte Carlo average of the instantaneous velocity.
  2. [Appendix A, Table 4] All quantitative results are reported from a single training run, and the second-stage fine-tuning length is selected by inspecting ImageNet-Test, the same benchmark used for the final tables. Table 4 shows a knife-edge result: 10k iterations give the headline scores, 20k slightly degrade, and 30k collapse catastrophically (LPIPS 0.689, CLIPIQA 0.294, MANIQA 0.2822). Without repeated runs or a proper validation split, the reported superiority over CTMSR could easily be within run-to-run variance or an artifact of early stopping on the test set. The authors should rerun with multiple seeds and report mean and standard deviation, and should hold out a validation set for hyperparameter selection.
  3. [Table 2 and Section 4, 'Evaluation Metrics'] The paper claims 'perceptual' superiority over CTMSR but reports worse LPIPS (0.228 vs 0.197), and it omits reference-based PSNR/SSIM entirely despite evaluating on a synthetic benchmark with ground-truth HR images. Since LPIPS is itself a widely used perceptual metric, reporting only CLIPIQA/MUSIQ/MANIQA is selective. The authors should include PSNR/SSIM (or other fidelity metrics) in Table 2 and discuss the LPIPS trade-off explicitly; ideally, they should also include a human reader study to substantiate the claim of 'fewer perceptual artifacts'.
  4. [Table 1 and Section 4.1] The efficiency comparison is not sufficiently explained to be convincing. According to Table 1, MeanSR has about 1.3x fewer parameters than CTMSR (131.3 vs 171.5) yet reports 6.6x fewer FLOPs (46.18 vs 305.50 G). Such a large discrepancy cannot be explained by a modest parameter reduction alone and suggests differences in latent resolution, output channels, or architecture choices between the models. The authors need to describe the architecture of MeanSR, the latent resolution used for FLOPs measurement, and the exact hardware/protocol for the runtime comparison; without this, the 'substantial FLOPs reduction' claim may not be comparing like with like.
  5. [Section 3, Stage 2, Eq. (8)] The second-stage pseudo-target tilde z0 = sg(hat z0 - Delta) with Delta = hat z'_0 - z'_0 is computed through a delayed target network, which makes the alignment objective a self-correcting fixed-point procedure. The paper does not analyze the convergence, bias, or stability of this target, nor does it compare with alternative trajectory-alignment objectives beyond the two variants in Table 3. This matters because the authors themselves observe a catastrophic degradation after 30k second-stage iterations (Table 4), suggesting the objective can become unstable. A formal or empirical analysis of the target's behavior would strengthen the method's credibility.
minor comments (5)
  1. [Table 1] The MeanSR row of Table 1 is garbled ('0.65611 131.346.181 26.392'), mixing metric values with parameter/FLOP/runtime numbers; fix the table formatting so that each column is readable.
  2. [Section 4.2] The phrase 'We alos analyze' contains a typo; it should be 'We also analyze'.
  3. [Algorithm 1] The loss is written as 'metric(u - stopgrad(u_tgt))' but the metric is not specified; if it is the L2 norm used in Eq. (5), state this explicitly.
  4. [Table 3 caption] The abbreviations 'DTM-u' and 'DTM-x0' are not defined in the caption or the text; add definitions to clarify what is being ablated.
  5. [Figure 1 caption] The caption states the x-axis is CLIPIQA and the y-axis is runtime, but the figure is not shown clearly in the text; if the axes are correct, ensure the figure is legible and includes labeled axes.

Circularity Check

1 steps flagged · score 2.0 of 10

One self-referential training target in Stage 1, but the headline benchmark claims rest on external evaluation and are not circular.

  1. self definitional [Algorithm 1 (Appendix B); Eq. (4)-(5), Sec. 3, Stage 1]
    "u, dudt = jvp(fn, (x_t, y, r, t), (v, 0, 0, 1)) u_tgt = v - (t - r) * dudt x_hat_0 = x_t - (t - r) * u loss_mf = metric(u - stopgrad(u_tgt))"

    In Eq. (4)-(5) and Algorithm 1, the target for the average-velocity network u_theta is u_tgt = v - (t-r)*dudt, where dudt is the Jacobian-vector product of fn, i.e., of u_theta itself, along the trajectory. The loss then regresses u onto this stop-grad target. Thus the 'target average velocity' is not an externally defined field; it is a function of the very network being trained. The optimization is a fixed-point consistency condition on u_theta — the target shifts as u_theta changes — rather than a supervised fit to an independent restoration velocity. Any claim that the model 'learns the average velocity' from first principles is therefore partly self-referential by construction.

full rationale

MeanSR's central claim is empirical: on RealSR, RealSet65, and ImageNet-Test it reports higher CLIPIQA/MUSIQ/MANIQA than CTMSR and lower FLOPs/latency. These numbers come from training and external evaluation, not from the self-referential velocity objective alone. The one genuine self-referential point is Stage 1: the regression target for u_theta is built from u_theta's own Jacobian (Algorithm 1), making the objective a fixed-point consistency condition. This is a reasonable training formulation (it follows MeanFlow) and does not make the benchmark superiority circular. The remaining weaknesses — a single training run (Appendix A), SATS and second-stage length selected on ImageNet-Test (Tables 4-5), and reliance on no-reference metrics — are evaluation-risk issues, not derivation circularity. The only self-citation is ACDMSR in related work, which is not load-bearing. Overall circularity is therefore minor.

Assumptions & free parameters 4 free parameters · 5 assumptions · 0 invented entities

The central claim rests on several domain assumptions (latent-space transport, transfer of synthetic degradations, metric validity) and empirically fitted hyperparameters (temporal sampling, loss weights, training length). The free parameters are tuned on the evaluation benchmark, which weakens the generality of the reported gains.

free parameters (4)
  • Stage 1 temporal sampling distribution logN(-0.4, 1.0) = mu=-0.4, sigma=1.0
    Selected as the best among uniform and logit-normal alternatives on the ImageNet-Test benchmark (Table 5).
  • Stage 2 temporal sampling distribution U(0,1) = uniform on [0,1]
    Selected as the best among several interval samplers on ImageNet-Test (Table 5).
  • Second-stage fine-tuning length = 10k iterations
    Best among 10k/20k/30k evaluated on ImageNet-Test (Table 4); longer training collapses perceptual scores.
  • Loss weights lambda_l1, lambda_vgg, lambda_dtm = 0.5, 0.1, 1.6
    Set empirically according to Appendix A with no sensitivity analysis.
assumptions (5)
  • domain assumption The interpolation path z_t = (1 - t) z_0 + t epsilon with epsilon ~ N(0,I) in VAE latent space is a valid transport from latent noise to latent HR images.
    Used in Eq. (1) for flow matching; inherited from latent diffusion literature but unverified for the SR setting.
  • domain assumption RealESRGAN's synthetic degradation pipeline is representative of real-world degradations, so training on it transfers to RealSR and RealSet65.
    Section 4 training details; all real-world claims rely on this transfer.
  • domain assumption No-reference metrics CLIPIQA, MUSIQ, MANIQA are reliable proxies for human perceptual quality.
    All headline results (Tables 1 and 2) use these metrics; the paper provides no human study.
  • ad hoc to paper The self-consistency target u_tgt = v - (t - r) d/dt u_theta converges to the true average velocity via fixed-point optimization.
    Algorithm 1; this is the MeanFlow training objective, whose stability is assumed rather than proven in this setting.
  • ad hoc to paper The second-stage pseudo-target tilde z0 = sg(hat z0 - Delta) with Delta = hat z'_0 - z'_0 provides valid perceptual supervision.
    Eqs. (7)-(9); the correctness of this self-correcting target is assumed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of MeanSR: Restoration Trajectory Learning for One-Step Perceptual Super-Resolution." pith.science (2026). https://pith.science/paper/XXIVZEO2

@misc{pith2026260809405,
  author       = {Pith},
  title        = {Pith review of: MeanSR: Restoration Trajectory Learning for One-Step Perceptual Super-Resolution},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XXIVZEO2}},
  note         = {Machine review of arXiv:2608.09405}
}
read the original abstract

Diffusion-based super-resolution (SR) achieves strong perceptual quality but requires costly iterative denoising. Existing one-step distillation methods reduce inference time but depend on expensive pretrained teachers, whereas CTMSR avoids distillation through PF-ODE consistency training yet does not explicitly model the restoration dynamics from low-resolution (LR) inputs to high-resolution (HR) images. We propose MeanSR, a one-step perceptual SR method that learns an LR-conditioned average velocity field to directly capture the finite-time transition from degraded or noisy inputs to plausible HR outputs. We further reformulate distribution trajectory matching for average-velocity generation and introduce a Stage-Aware Temporal Sampling strategy to improve trajectory learning. Experiments on synthetic and real-world benchmarks show that MeanSR outperforms CTMSR on CLIPIQA, MUSIQ, and MANIQA while substantially reducing FLOPs and inference latency. MeanSR also reconstructs sharper structures and more realistic textures with fewer perceptual artifacts.

Figures

Figures reproduced from arXiv: 2608.09405 by the authors.

Figure 1
Figure 1. Comparison of representative one-step super [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of the proposed MeanSR framework. Stage 1 estimates an LR-conditioned restoration trajectory by predicting [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Qualitative comparison on the RealSet65 dataset. Zoom in for a better view. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (5 more)
Figure 4
Figure 4. Figure 4: Qualitative comparison on synthetic ImageNet examples. Zoom in for a better view. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 5
Figure 5. Figure 5: Training convergence comparison on ImageNet [PITH_FULL_IMAGE:figures/full_fig_p007_5.png]
Figure 6
Figure 6. Figure 6: Qualitative comparison on the RealSet65 dataset. [PITH_FULL_IMAGE:figures/full_fig_p012_6.png]
Figure 7
Figure 7. Figure 7: Qualitative comparison on the RealSR dataset. [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Qualitative comparison on the ImageNet-Test dataset. [PITH_FULL_IMAGE:figures/full_fig_p014_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

49 extracted references · 22 canonical work pages

  1. [1]

    Advances in Neural Information Processing Systems , volume=

    Denoising Diffusion Probabilistic Models , author=. Advances in Neural Information Processing Systems , volume=

  2. [2]

    International Conference on Learning Representations , year=

    Score-Based Generative Modeling through Stochastic Differential Equations , author=. International Conference on Learning Representations , year=

  3. [3]

    Advances in Neural Information Processing Systems , year=

    Elucidating the Design Space of Diffusion-Based Generative Models , author=. Advances in Neural Information Processing Systems , year=

  4. [4]

    Consistency models , author=

  5. [5]

    International Conference on Learning Representations , year=

    Flow Matching for Generative Modeling , author=. International Conference on Learning Representations , year=

  6. [6]

    International Conference on Learning Representations , year=

    Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow , author=. International Conference on Learning Representations , year=

  7. [7]

    arXiv preprint arXiv:2505.13447 , year=

    Mean Flows for One-Step Generative Modeling , author=. arXiv preprint arXiv:2505.13447 , year=

  8. [8]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

    Consistency Trajectory Matching for One-Step Generative Super-Resolution , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

Show all 49 references
  1. [9]

    Wang, Xintao and Yu, Ke and Wu, Shixiang and Gu, Jinjin and Liu, Yihao and Dong, Chao and Qiao, Yu and Loy, Chen Change , booktitle=

  2. [10]

    Zhang, Jingyun and Zeng, Huan and Guo, Yuchao and Zhang, Lei , booktitle=

  3. [11]

    Wang, Xintao and Xie, Liangbin and Dong, Chao and Shan, Ying , booktitle=. Real-

  4. [12]

    Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

    Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network , author=. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

  5. [13]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

    High-Resolution Image Synthesis with Latent Diffusion Models , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

  6. [14]

    Advances in Neural Information Processing Systems , year=

    Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding , author=. Advances in Neural Information Processing Systems , year=

  7. [15]

    Analyzing and Improving the Image Quality of

    Karras, Tero and Laine, Samuli and Aittala, Miika and Hellsten, Janne and Lehtinen, Jaakko and Aila, Timo , booktitle=. Analyzing and Improving the Image Quality of

  8. [16]

    Proceedings of the European Conference on Computer Vision , pages=

    Image Super-Resolution Using Very Deep Residual Channel Attention Networks , author=. Proceedings of the European Conference on Computer Vision , pages=

  9. [17]

    Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops , pages=

    Enhanced Deep Residual Networks for Single Image Super-Resolution , author=. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops , pages=

  10. [18]

    International Conference on Learning Representations , year=

    Diffusion Posterior Sampling for General Noisy Inverse Problems , author=. International Conference on Learning Representations , year=

  11. [19]

    International Journal of Computer Vision , volume=

    Exploiting diffusion prior for real-world image super-resolution , author=. International Journal of Computer Vision , volume=. 2024 , publisher=

  12. [20]

    NeurIPS , year=

    ResShift: Efficient Diffusion Model for Image Super-Resolution by Residual Shifting , author=. NeurIPS , year=

  13. [21]

    Proceedings of the IEEE/CVF international conference on computer vision , pages=

    Designing a Practical Degradation Model for Deep Blind Image Super-Resolution , author=. Proceedings of the IEEE/CVF international conference on computer vision , pages=

  14. [22]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

    SinSR: diffusion-based image super-resolution in a single step , author=. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , pages=

  15. [23]

    IEEE transactions on pattern analysis and machine intelligence , volume=

    Image super-resolution using deep convolutional networks , author=. IEEE transactions on pattern analysis and machine intelligence , volume=. 2015 , publisher=

  16. [24]

    Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

    Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network , author=. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

  17. [25]

    IEEE Transactions on Broadcasting , volume=

    ACDMSR: Accelerated conditional diffusion models for single image super-resolution , author=. IEEE Transactions on Broadcasting , volume=. 2024 , publisher=

  18. [26]

    Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

    Deep Back-Projection Networks for Super-Resolution , author=. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition , pages=

  19. [27]

    IEEE transactions on pattern analysis and machine intelligence , volume=

    Image super-resolution via iterative refinement , author=. IEEE transactions on pattern analysis and machine intelligence , volume=. 2022 , publisher=

  20. [28]

    European conference on computer vision , pages=

    Diffbir: Toward blind image restoration with generative diffusion prior , author=. European conference on computer vision , pages=. 2024 , organization=

  21. [29]

    Advances in Neural Information Processing Systems , year=

    One-Step Effective Diffusion Network for Real-World Image Super-Resolution , author=. Advances in Neural Information Processing Systems , year=

  22. [30]

    Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

    Tsd-sr: One-step diffusion with target score distillation for real-world image super-resolution , author=. Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

  23. [31]

    arXiv preprint arXiv:2502.01993 , year=

    One diffusion step to real-world super-resolution via flow trajectory distillation , author=. arXiv preprint arXiv:2502.01993 , year=

  24. [32]

    Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

    Seesr: Towards semantics-aware real-world image super-resolution , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

  25. [33]

    Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , year =

    Yu, Xiaolong and Zhang, Kai and others , title =. Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition , year =

  26. [34]

    arXiv preprint arXiv:2409.17778 , year=

    Taming diffusion prior for image super-resolution with domain shift sdes , author=. arXiv preprint arXiv:2409.17778 , year=

  27. [35]

    Neurocomputing , volume=

    Srdiff: Single image super-resolution with diffusion probabilistic models , author=. Neurocomputing , volume=. 2022 , publisher=

  28. [36]

    arXiv preprint arXiv:2302.00482 , year=

    Improving and generalizing flow-based generative models with minibatch optimal transport , author=. arXiv preprint arXiv:2302.00482 , year=

  29. [37]

    Journal of Machine Learning Research , volume=

    Stochastic interpolants: A unifying framework for flows and diffusions , author=. Journal of Machine Learning Research , volume=

  30. [38]

    Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

    Scaling up to excellence: Practicing model scaling for photo-realistic image restoration in the wild , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

  31. [39]

    Proceedings of the AAAI conference on artificial intelligence , volume=

    Resdiff: Combining cnn and diffusion model for image super-resolution , author=. Proceedings of the AAAI conference on artificial intelligence , volume=

  32. [40]

    Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

    Dit4sr: Taming diffusion transformer for real-world image super-resolution , author=. Proceedings of the IEEE/CVF International Conference on Computer Vision , pages=

  33. [41]

    Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

    Uncertainty-guided perturbation for image super-resolution diffusion model , author=. Proceedings of the Computer Vision and Pattern Recognition Conference , pages=

  34. [42]

    2009 IEEE conference on computer vision and pattern recognition , pages=

    Imagenet: A large-scale hierarchical image database , author=. 2009 IEEE conference on computer vision and pattern recognition , pages=. 2009 , organization=

  35. [43]

    Proceedings of the IEEE/CVF international conference on computer vision , pages=

    Toward real-world single image super-resolution: A new benchmark and a new model , author=. Proceedings of the IEEE/CVF international conference on computer vision , pages=

  36. [44]

    Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

    High-resolution image synthesis with latent diffusion models , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

  37. [45]

    Proceedings of the AAAI conference on artificial intelligence , volume=

    Exploring clip for assessing the look and feel of images , author=. Proceedings of the AAAI conference on artificial intelligence , volume=

  38. [46]

    Proceedings of the IEEE/CVF international conference on computer vision , pages=

    Musiq: Multi-scale image quality transformer , author=. Proceedings of the IEEE/CVF international conference on computer vision , pages=

  39. [47]

    Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

    Maniqa: Multi-dimension attention network for no-reference image quality assessment , author=. Proceedings of the IEEE/CVF conference on computer vision and pattern recognition , pages=

  40. [48]

    Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

    The unreasonable effectiveness of deep features as a perceptual metric , author=. Proceedings of the IEEE conference on computer vision and pattern recognition , pages=

  41. [49]

    Proceedings of the AAAI conference on artificial intelligence , volume=

    Effective diffusion transformer architecture for image super-resolution , author=. Proceedings of the AAAI conference on artificial intelligence , volume=

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.