{"id":"2937abe0-1878-4145-bb75-28f57c03085e","arxiv_id":"2412.03268","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"RFSR fine-tunes diffusion super-resolution models with early wavelet low-frequency constraints and late CLIP-IQA/ImageReward feedback, boosting non-reference metrics at a small LPIPS cost.","lead":"This paper introduces a fine-tuning recipe for diffusion-based image super-resolution that applies low-frequency structure constraints early and reward feedback learning late in the denoising process. It reports substantial gains on learned quality metrics across three SR models, but the evaluation overlaps with the training objectives and no human study is provided.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The subjective-quality claim is supported only by reward-metric gains; the paper itself shows CLIPIQA diverging from human quality under reward hacking, so human evaluation is required.","rationale":"I read the paper as proposing a plug-and-play fine-tuning method for diffusion-based image super-resolution, with the strongest claim being that it improves perceptual and aesthetic quality to the point of 'excellent subjective results.' For that claim to hold, the reward models used during training must be aligned with human perception, and the reported metrics must not simply reflect optimization of the same rewards used as losses. Both conditions are insecure. The in-paper reward-hacking evidence (Figure 2 and Table 4) directly shows CLIPIQA diverging from subjective quality, so using CLIPIQA as the main quantitative evidence for 'subjective results' is not safe without human evaluation. This is the same weakest assumption the reader identified, and I agree with the conditional verdict. I do not see an internal inconsistency severe enough to warrant rejection: the method is described with clear ablations, the code is released, and the plug-and-play framing is plausible. But the missing human evaluation is the decisive gap, and the released code makes the proposed test feasible. Because the reader already recommended a conditional verdict, my analysis does not change that verdict.","tokens_in":11941,"tokens_out":3697,"duration_ms":38358,"concrete_test":"Conduct a pairwise human preference study with at least 20 raters on roughly 50 images from DIV2K-val, RealSR, and DRealSR, comparing each base ISR model (DiffBIR, PASD, SeeSR) against its RFSR fine-tune, and include the no-regularization reward-hacked checkpoint from Table 4 as a control. Compute human win rates and correlate them with CLIPIQA, MUSIQ, and Aesthetic deltas. If RFSR wins significantly over the base model while the no-regularization control does not, the subjective-quality claim is supported; if the metric deltas do not predict human preferences, the load-bearing assumption fails.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that RFSR yields 'excellent subjective results' depends on the assumption that the reward models used for training track human perceptual quality. This assumption is both circular and weakened by evidence inside the paper. CLIP-IQA is used as a training reward (Eq. 5) and then reported as a headline evaluation metric (Table 1), so part of the reported CLIPIQA gain is by construction. More seriously, Table 4 shows the no-regularization baseline reaches CLIPIQA 0.8964 and Aesthetic 5.3612, which are higher than the proposed Gram-KL method's 0.7944 and 5.2683, yet the authors describe that baseline as reward hacking with visibly degraded images. Thus CLIPIQA gains can be inversely related to subjective quality. The only evidence for the 'subjective' part of the claim is a few qualitative figures and non-reference metrics; there is no human study, no error bars, and only single-seed results. The conclusion even acknowledges that the reward model 'lacks robustness when confronted with larger-scale real-world data and diffusion-generated data.' Therefore the central claim is not yet established: the reported metric improvements could be reward overoptimization rather than genuine perceptual improvement.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes RFSR, a plug-and-play fine-tuning method for diffusion-based image super-resolution (ISR) models. During early denoising steps the method imposes a low-frequency DWT constraint against the ground truth to preserve structure; during later steps it trains with reward feedback from CLIP-IQA and ImageReward, plus a Gram-KL regularization intended to mitigate reward hacking. Experiments on DIV2K-val, DRealSR, and RealSR report improved MANIQA, MUSIQ, CLIPIQA, and Aesthetic scores for DiffBIR, PASD, and SeeSR after RFSR fine-tuning. The paper claims 'excellent subjective results' based on these metrics and selected visual comparisons, and releases code.","tokens_in":12170,"tokens_out":3149,"duration_ms":31892,"significance":"If the subjective-quality claim were convincingly established, RFSR would be a useful general recipe: it is model-agnostic, simple to implement, and the code is released. The authors also deserve credit for identifying and attempting to mitigate reward hacking in the ISR setting via Gram-KL regularization. However, the central evidence is weakened by a circularity: CLIP-IQA is both a training reward (Eq. 5) and a headline evaluation metric (Table 1), and Table 4 explicitly shows a reward-hacked baseline with higher CLIPIQA and Aesthetic scores than the proposed method. The claim of 'excellent subjective results' therefore currently rests on selected images and non-reference metrics rather than on demonstrated human preference. The significance is conditional on adding a human evaluation and more rigorous statistical reporting.","major_comments":[{"comment":"CLIP-IQA is used as a training reward in Eq. (5) and then reported as the primary evaluation metric in Tables 1, 2, 3, and 4. Part of the reported CLIPIQA gain is therefore by construction, and the same applies, to a smaller extent, to ImageReward-aligned aesthetic judgments. To support the claim that RFSR improves perceptual quality, the paper should either report a human preference study, or evaluate on metrics that are not directly or indirectly optimized by the training rewards, or both.","section":"Sec. 3.3, Eq. (5); Sec. 4.2, Metrics"},{"comment":"Table 4 shows that the 'w/o regularization' baseline achieves CLIPIQA 0.8964 and Aesthetic 5.3612, which are substantially higher than the proposed Gram-KL method's 0.7944 and 5.2683, yet the authors describe that baseline as reward hacking with visibly degraded images. This is direct evidence that the reported reward metrics can be inversely related to subjective quality. The paper's own data therefore undermine the inference from Table 1's metric gains to 'excellent subjective results'; a human evaluation is needed to establish the central claim.","section":"Sec. 4.4, Table 4; Sec. 1, Fig. 2"},{"comment":"All quantitative results are reported from a single training run with no error bars, no multiple seeds, and no statistical significance tests. Some differences between configurations are small (e.g., Aesthetic 5.2683 vs. 5.2669 in Table 4), so it is not possible to tell whether the reported improvements are robust. The authors should provide at least 3 seeds, report mean and standard deviation, and discuss checkpoint selection or early stopping, especially since Fig. 2 shows metric behavior varying strongly with training iterations.","section":"Sec. 4, Tables 1-5"},{"comment":"The qualitative claim that RFSR 'excels at enhancing high-quality texture details' is supported only by selected crops in Figure 4. Since the paper explicitly acknowledges that reward models 'lack robustness when confronted with larger-scale real-world data and diffusion-generated data' (Sec. 5), the subjective evidence should be supplemented with a formal user study, ideally with multiple raters and a forced-choice protocol against the baseline models.","section":"Sec. 4.3, Qualitative Comparisons"}],"minor_comments":[{"comment":"The notation in Eq. (2) uses absolute-value bars, which should be clarified as an L1 norm; otherwise the equation is dimensionally ambiguous.","section":"Sec. 3.2, Eq. (2)"},{"comment":"The definition of DWT(·)_LL should be made explicit; the text refers to 'DWT(It)_LL' but Eq. (1) defines only the full DWT output.","section":"Sec. 3.2, Eq. (1)"},{"comment":"The row labels in Table 2 are inconsistent with the text: the paper's described default setting 'st1=20, st2=40' is listed only as 'Ours', while the first row 'st1∈[1,40], st2∈[41,50]' merges two different interval lengths. Please clarify which rows correspond to which sampling-step schedules.","section":"Sec. 4.4, Table 2"},{"comment":"The claim 'We are the first to introduce reward feedback learning into super-resolution fine-tuning' is strong and should be positioned more carefully against existing reward-finetuning works for diffusion models [4, 6, 35] and any prior use of perceptual rewards in restoration.","section":"Sec. 1, Contributions"}],"recommendation":"major_revision","confidential_remarks":"The paper presents a plausible and easy-to-reproduce recipe, and the code release is a plus. However, the central claim of subjective improvement is not currently supported: the main metric is a training reward, the paper's own reward-hacking baseline shows metric gains with degraded images, and there is no human study. A major revision that adds a human evaluation, multi-seed statistics, and a clearer separation between optimized and non-optimized metrics would make the paper publishable. I do not see a fundamental flaw in the method itself, so I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this paper: it proposes RFSR, a fine-tuning method for diffusion-based image super-resolution that applies a low-frequency DWT constraint early in denoising and reward feedback (CLIP-IQA + ImageReward) plus Gram-KL regularization late. It is plug-and-play across DiffBIR, PASD, and SeeSR, and it reports consistent gains on non-reference metrics. The core design is plausible and the motivation is concrete: the authors measure that ISR diffusion models stabilize low-frequency structure early and diverge from ground truth in high frequencies later, so they schedule supervision accordingly. The Gram-KL regularizer is a reasonable, targeted answer to reward hacking, and the ablations show it beats LoRA and KL baselines. Code is available. That is the good part.\n\nNow the soft spots. The headline claim is \"excellent subjective results,\" but the evidence does not yet support that. CLIP-IQA is both a training reward (Eq. 5) and a headline evaluation metric (Tables 1, 3, 4), so part of the reported gain is by construction. More tellingly, Table 4 shows the no-regularization baseline reaching CLIPIQA 0.8964 and Aesthetic 5.3612—higher than the proposed method's 0.7944 and 5.2683—yet the authors themselves describe that baseline as reward hacking with visibly degraded images. That means the metric can go up while subjective quality goes down. There is no human study, no error bars, no multiple seeds. The conclusion even concedes the reward model lacks robustness on larger-scale real-world and diffusion-generated data. So the central claim is not established; the reported improvements could be reward overoptimization.\n\nMinor issues: the \"first to introduce reward feedback learning into super-resolution\" contribution overstates novelty, since differentiable-reward fine-tuning is well established in T2I diffusion; the specific ISR application and timestep-aware combination are new, but the claim as written is too broad. Also, the training description is ambiguous—\"we enable gradient updates only in the final step\" needs a clearer explanation of how the intermediate-step losses actually contribute to the gradient.\n\nThese are real concerns but not fatal. The method is coherent, the frequency-domain analysis is a useful observation, and the plug-and-play gains on non-optimized metrics are encouraging. The fixes are straightforward: add a human evaluation, report multiple seeds with variance, clarify the gradient flow, and temper the novelty statement. This paper deserves a serious referee; I would send it to review with a request for major revision. I'd bring it to reading group to discuss evaluation methodology in reward fine-tuning, and I'd cite it if I worked in this area.","headline":"A sensible plug-and-play reward fine-tuning recipe for diffusion-based ISR with a coherent timestep-aware design, but the central subjective-quality claim rests on metrics that the method itself optimizes and that the paper's own ablation shows can diverge from visual quality.","tokens_in":12749,"tokens_out":1826,"would_cite":true,"duration_ms":20084,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Fine-tuning diffusion-based super-resolution models with timestep-aware reward feedback learning improves their perceptual and aesthetic output quality, as shown by higher no-reference quality scores on several established models.","keywords":["image super-resolution","diffusion models","reward feedback learning","timestep-aware training","CLIP-IQA","ImageReward","Gram-KL regularization","perceptual quality"],"falsifier":"A blind user study comparing original and RFSR-fine-tuned outputs of DiffBIR, PASD, and SeeSR on diverse real-world images, scored by human raters, would settle the claim; if human preference does not track the reported MANIQA, CLIPIQA, and Aesthetic gains, the core assumption fails. An additional holdout set of images from a distribution unlike the training data would test whether the reward models' judgments generalize.","tokens_in":11681,"feed_emoji":"🖼️","tokens_out":5743,"duration_ms":48035,"temperature":0.7,"pith_summary":"This paper claims that diffusion-based image super-resolution models, typically trained with denoising losses, can be further improved by fine-tuning them with reward feedback learning. The proposed method, RFSR, splits the denoising trajectory: early steps are constrained on low-frequency structure to keep the image faithful, while later steps are optimized with aesthetic and perceptual rewards to boost subjective quality. A Gram-KL regularizer is added to prevent the stylization artifacts that pure reward optimization causes. Experiments on DiffBIR, PASD, and SeeSR show consistent gains on no-reference perceptual and aesthetic metrics, suggesting the approach works as a plug-and-play improvement layer.","feed_headline":"Reward feedback improves super-resolution quality scores","feed_subtitle":"Timestep-aware fine-tuning lifts perceptual and aesthetic metrics on diffusion-based SR models.","key_machinery":"The method uses three components. A low-frequency structure constraint, computed with the discrete wavelet transform, matches the LL subband of early denoising outputs to the ground truth, keeping the overall layout stable. Reward feedback learning at late timesteps uses CLIP-IQA and ImageReward as differentiable reward models, pushing outputs toward higher perceptual and human-preference scores. A Gram-KL regularizer penalizes differences between VGG Gram matrices of the fine-tuned and frozen pretrained models, opposing the stylistic shifts characteristic of reward hacking. Gradients are updated only at the final denoising step to avoid instability, and the whole procedure fine-tunes an already-trained model in a plug-and-play fashion.","core_discovery":"The central claim is that reward feedback learning, selectively applied to the later denoising steps of a diffusion-based super-resolution model, can push output images toward higher perceptual and aesthetic quality while preserving the structural fidelity established early in the denoising process. The authors support this by showing, via discrete wavelet transform analysis, that low-frequency structure is settled early while high-frequency texture develops later and tends to diverge from the ground truth. They therefore apply a low-frequency constraint at large timesteps and a reward loss at small timesteps, which outperforms applying either constraint uniformly. For example, fine-tuning SeeSR with RFSR raises MANIQA from 0.5091 to 0.5954 and CLIPIQA from 0.6989 to 0.7944 on DIV2K-val.","pith_inferences":["Because CLIPIQA is used both as a training reward and as the headline evaluation metric, part of the measured gain is by construction; a blind human-preference study would be needed to confirm the improvements are genuinely perceptual.","The timestep-aware recipe could transfer to other conditional generation tasks such as inpainting or deblurring, where early structure and late texture also separate.","The divergence of high-frequency details from ground truth in late denoising steps suggests a fundamental tension between fidelity metrics and perceived quality; reward fine-tuning deliberately trades off the former for the latter.","The Gram-KL regularizer depends on VGG feature statistics; testing whether other feature extractors or style statistics (mean/covariance) behave similarly would clarify the mechanism."],"forward_implications":["Existing diffusion-based ISR models can be upgraded with RFSR without architectural changes or retraining from scratch.","The timestep-aware split — structure constraint early, reward late — offers a general recipe for fine-tuning other conditional diffusion restoration models.","The Gram-KL regularizer provides a lightweight counter to reward hacking that is orthogonal to LoRA or KL-based constraints.","The reported metric gains indicate that perceptual quality can be improved while keeping fidelity (LPIPS) close to the original model's level."],"supporting_citations":[{"why":"Introduces reward hacking in direct fine-tuning of diffusion models and proposes early-termination/LoRA remedies, motivating the Gram-KL regularizer.","marker":"[4]"},{"why":"Provides the ImageReward reward model and the technique of evaluating rewards at intermediate denoising steps.","marker":"[35]"},{"why":"Supplies the MANIQA, MUSIQ, and CLIPIQA quality scorers used as both reward losses and evaluation metrics.","marker":"[36]"},{"why":"Defines Gram matrices as style statistics, the basis of the Gram-KL regularization against stylization.","marker":"[7]"},{"why":"DiffBIR is one of the three diffusion-based ISR models that RFSR fine-tunes to demonstrate plug-and-play gains.","marker":"[16]"},{"why":"SeeSR is the model used for the main ablations and reports the largest relative gains after RFSR fine-tuning.","marker":"[34]"},{"why":"PASD, a semantic-cue diffusion ISR model, is fine-tuned with RFSR as another validation of the method.","marker":"[27]"}],"fun_headline_variants":["Reward feedback in late denoising boosts SR quality","Timestep-aware reward learning improves diffusion SR","RFSR: plug-and-play reward fine-tuning for ISR","Selective reward feedback enhances super-resolution diffusion","Later-stage reward improves perceptual SR metrics"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The reward models CLIP-IQA and ImageReward are reliable, differentiable proxies for human perceptual quality when used as training losses, so that optimizing them actually improves how people perceive the super-resolved images.","fun_headline_variants_meta":{"raw":{"variants":["Reward feedback in late denoising boosts SR quality","Timestep-aware reward learning improves diffusion SR","RFSR: plug-and-play reward fine-tuning for ISR","Selective reward feedback enhances super-resolution diffusion","Later-stage reward improves perceptual SR metrics"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000712,"raw_usage":{"total_tokens":3174,"prompt_tokens":886,"completion_tokens":2288,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":502,"completion_tokens_details":{"reasoning_tokens":2214}},"tokens_in":502,"tokens_out":2288,"duration_ms":14864,"temperature":1.0,"reasoning_tokens":2214,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T22:35:18.511641+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A blind user study comparing original and RFSR-fine-tuned outputs of DiffBIR, PASD, and SeeSR on diverse real-world images, scored by human raters, would settle the claim; if human preference does not track the reported MANIQA, CLIPIQA, and Aesthetic gains, the core assumption fails. An additional holdout set of images from a distribution unlike the training data would test whether the reward models' judgments generalize.","supporting_citations":[{"cited_title":"Imagere- ward: Learning and evaluating human preferences for text- to-image generation","cited_arxiv_id":null,"evidence_quote":"Provides the ImageReward reward model and the technique of evaluating rewards at intermediate denoising steps."},{"cited_title":"Maniqa: Multi-dimension attention network for no-reference image quality assessment","cited_arxiv_id":null,"evidence_quote":"Supplies the MANIQA, MUSIQ, and CLIPIQA quality scorers used as both reward losses and evaluation metrics."},{"cited_title":"Im- age style transfer using convolutional neural networks","cited_arxiv_id":null,"evidence_quote":"Defines Gram matrices as style statistics, the basis of the Gram-KL regularization against stylization."},{"cited_title":"Seesr: Towards semantics- aware real-world image super-resolution","cited_arxiv_id":null,"evidence_quote":"SeeSR is the model used for the main ablations and reports the largest relative gains after RFSR fine-tuning."},{"cited_title":"Pixel-aware stable diffusion for realistic image super- resolution and personalized stylization","cited_arxiv_id":null,"evidence_quote":"PASD, a semantic-cue diffusion ISR model, is fine-tuned with RFSR as another validation of the method."}],"review_version":1}