{"id":"a8626a83-6561-4342-999b-7f4ec72a382e","arxiv_id":"2412.18945","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":6,"one_line_summary":"A consistency-distillation method trains from a fixed partial-noise start and uses a replay bank plus DINO-v2 adversarial loss to improve few-step image and video stylization.","lead":"This paper introduces a way to speed up diffusion-based image and video style transfer by training a student model to follow a single teacher denoising trajectory from a fixed noise level, plus an adversarial loss for cleaner output. The authors report better style similarity and aesthetic scores than existing accelerators at 4 to 8 sampling steps.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline 'surpasses baselines' rests on a CFG=6-only comparison; the supplement admits most baselines' recommended CFG is 1–2 and that low CFG makes them less stylized, so the empirical margin may be a protocol artifact.","rationale":"The reader's eta-specific limitation is real but secondary: it is a usability constraint that the authors partially address by training separate models for eta=0.65, 0.75, 0.85, and 0.95, and it does not by itself undercut the headline comparison at the evaluated strength. The most load-bearing concern is the evaluation protocol. The paper's own supplementary text supplies the recommended CFG values for each baseline and acknowledges that low recommended CFG makes those baselines appear less stylized, yet the main quantitative table is produced at CFG=6 for all methods. Since the central claim is empirical superiority in style similarity and aesthetics, the absence of a per-method CFG tuning protocol means the observed ranking could be an artifact of choosing a CFG that favors STD. A concrete re-evaluation with recommended and best-validation CFGs, plus error bars, would settle whether the claim survives. I do not see an internal inconsistency in the method, and the trajectory-bank and adversarial-loss components are described precisely enough to reproduce; the concern is about the evidence supporting the 'surpasses' claim, not about the method's feasibility.","tokens_in":14692,"tokens_out":6581,"duration_ms":58008,"concrete_test":"Rerun the Table 1 protocol twice: (a) each method at its own recommended CFG from Supplementary 9.1 (LCM/TCD/SDXL-Lightning/MCM=1.0, PCM=1.6, TDD=2.0, Hyper-SD=6.0, STD=6.0); (b) each method at its best CFG chosen on a held-out subset from {1,1.6,2,4,6,8}. Report CSD, aesthetic, and warping error with at least 3 seeds for NFE=4,6,8. If STD's CSD lead over the best baseline is below about 0.02 or not significant in either protocol, the paper's headline claim is an artifact of CFG selection.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central quantitative claim (Abstract; Table 1) is that STD surpasses LCM, TCD, PCM, TDD, Hyper-SD, SDXL-Lightning, and MCM in CSD and aesthetics at NFE 4–8. In the main comparison, all methods are run at CFG=6 (Fig. 4 caption; Sec. 5.2), while Supplementary 9.1 states that the recommended CFGs are 1.0 for LCM/TCD/SDXL-Lightning/MCM, 1.6 for PCM, and 2.0 for TDD, and explicitly attributes the weaker stylization of baselines to this low recommended CFG. Thus the headline table compares methods outside their intended operating regimes and in a regime that the authors say is biased toward STD. The CFG sweep in Fig. 5 is only a line chart with no numeric values, no per-method optimum, and no error bars. Because CSD is the primary evidence for 'surpasses', the observed margin (e.g., 0.554 vs 0.522 for Hyper-SD at NFE=8) could be produced by CFG choice rather than by trajectory distillation; the aesthetic predictor also rewards the contrast/saturation that the asymmetric adversarial loss is designed to increase, so the two metrics are not independent controls. This is a protocol concern, not an internal contradiction, but it is load-bearing: without a per-method CFG tuning protocol, the central claim is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes single-trajectory distillation (STD) for accelerating diffusion-based image and video style transfer. The method starts from a fixed partial-noise strength eta, distills the teacher's reverse-time trajectory for that eta using a consistency-style loss, and introduces a trajectory bank to reuse teacher states during training. An asymmetric adversarial loss based on DINO-v2 features is added to improve saturation and reduce speckle noise. The authors report improved CSD and aesthetic scores over LCM, TCD, PCM, TDD, Hyper-SD, SDXL-Lightning, and MCM at NFE 4-8, and support the method with ablations and a theoretical motivation.","tokens_in":15066,"tokens_out":6151,"duration_ms":55498,"significance":"If the empirical claims are substantiated, the contribution is practically useful: it targets the common partial-noise editing setting rather than full text-to-image generation, and the trajectory bank is a sensible way to reduce the training cost of multi-step distillation. The idea of matching the student to the teacher's single trajectory for a fixed eta is simple and worth exploring. However, the current evaluation protocol and the theoretical justification have load-bearing weaknesses that prevent accepting the central 'surpasses all baselines' claim as stated.","major_comments":[{"comment":"The headline comparison fixes CFG=6 for all methods, while Supplementary 9.1 states that the recommended CFGs are 1.0 for LCM/TCD/SDXL-Lightning/MCM, 1.6 for PCM, and 2.0 for TDD, and explicitly explains that the baselines 'often appear less stylized' because of the low recommended CFG. This means the reported margins (e.g., CSD 0.554 vs 0.522 for Hyper-SD at NFE=8) may be artifacts of the CFG protocol rather than evidence that STD is superior. Please report results at each method's recommended CFG and at each method's best CFG, and provide per-method tuning details. Figure 5, which is offered as a CFG sweep, is a line chart without numeric values, per-method optima, or error bars, so it does not resolve this concern.","section":"Section 5.2 / Table 1 / Supplementary 9.1"},{"comment":"The theorem claims that PF-ODE trajectories from any two forward points are inconsistent for an imperfect teacher, but the proof only bounds the distance between x_s and the one-step DDIM denoised sample by C_{t,s} * delta_phi. An upper bound of this form does not establish that the trajectories are necessarily non-equivalent; the actual error could be zero even when delta_phi > 0. In addition, C_{t,s} as defined can be negative for some (t,s), and the proof treats vector differences as scalars. Please either replace the theorem with a rigorous lower-bound or non-equivalence argument, or explicitly present it as a heuristic motivation rather than a formal theorem.","section":"Section 4.1 / Eq. (9)-(10) / Supplementary 7"},{"comment":"The central empirical claim is not supported with appropriate statistics. Table 1 reports single-run metric values without variance, confidence intervals, or significance tests, and Figure 5 is a line chart with no numeric values or error bars. Given the small differences between STD and the best baselines (e.g., CSD 0.554 vs 0.522; aesthetic 5.190 vs 5.163 at NFE=8), repeated-run statistics are needed to substantiate the word 'surpasses'. Furthermore, the asymmetric adversarial loss is explicitly designed to increase saturation and contrast (Section 5.3), and the aesthetic predictor may reward exactly those low-level properties, so the aesthetic and CSD metrics are not independent controls. A human evaluation or an alternative aesthetic metric would help.","section":"Section 5.2 / Table 1 / Figure 5"}],"minor_comments":[{"comment":"There is a duplicated sentence: 'Distillation can be specifically tailored to the complete trajectory for a particular eta, called single-trajectory distillation. This approach reduces error and improves alignment...' appears twice in consecutive paragraphs.","section":"Section 4.1"},{"comment":"In the conclusion, 'trajetories' should be 'trajectories'.","section":"Section 6"},{"comment":"Equation (20) defines epsilon ~ N(0, sigma^2 I), but the DDIM derivation that follows sets sigma_t = 0. Please clarify the notation and the role of sigma in the forward process.","section":"Supplementary 7, Eq. (20)"},{"comment":"Line 5 samples (x0, xt_{n+1}, c, t_{n+1}) from the trajectory bank, but earlier text defines bank entries as (x0, x_hat_t, c, t). The notation should be made consistent.","section":"Algorithm 1"},{"comment":"The text says 'We evaluated each method with CFG values of 2, 4, 6, and 8' and refers to Figure 5, but the figure does not report the numeric scores. Please provide the underlying numbers in a table or as a supplementary table.","section":"Section 5.2"}],"recommendation":"major_revision","confidential_remarks":"The CFG-protocol concern raised by the stress-test is valid and is explicitly acknowledged in the paper's own supplement, so it must be addressed before publication. The authors should also ensure that the project page and code, promised in the abstract, are actually available and linked in the final version."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take. The core idea is a sensible specialization of consistency distillation for partial-noise editing: fix the starting noise strength eta, distill a single trajectory from x_eta to x0, and use a trajectory bank to avoid repeated multi-step teacher rollouts during training. That combination, plus the asymmetric adversarial loss comparing the student's output at step s with a real image noised to r<s, is a genuine, if narrow, extension of TDD/PCM/MCM. The paper is clearly written, the method is well specified, and the ablations support the internal design choices.\n\nThe problem is the headline comparison. Table 1 evaluates every method at CFG=6, and the supplement concedes that most baselines recommend CFG 1-2 and that their lower CFG is exactly why they look less stylized. That is a protocol that biases the comparison toward STD. A line chart at CFG 2, 4, 6, 8 with no numbers and no per-method optimum doesn't fix it. So the claim that STD 'surpasses' LCM, TCD, PCM, TDD, Hyper-SD, SDXL-Lightning, and MCM is not established by the evidence in the paper. This is load-bearing, not a minor quibble, because CSD and aesthetics are the only quantitative evidence.\n\nSecond, the theory is oversold. The theorem in Section 4.1 bounds the difference between a forward sample and a one-step DDIM denoising from a different starting point; that is a bound on a single step, not on the whole trajectory. The abstract's 'reduced error of the whole trajectory' is not what is proved.\n\nThird, no error bars or significance tests anywhere, and each eta requires a separate model. Those are minor-to-moderate; the first two are the real issues.\n\nCredit where due: the trajectory bank is a practical training acceleration, the choice of DINO-v2 features for the discriminator is well motivated for video memory, and the authors are honest enough to print the recommended-CFG admission in the supplement. The internal logic holds; the external comparison doesn't.\n\nWho is this for? Someone building few-step stylization pipelines will get useful engineering cues. The paper deserves a serious referee, but it needs major revisions: per-method CFG tuning with reported numbers, error bars, a corrected theorem statement, and ideally code.","headline":"A sensible distillation recipe for stylization, but the 'surpasses baselines' claim rests on a CFG protocol that the authors' own supplement admits favors them.","tokens_in":15576,"tokens_out":2657,"would_cite":false,"duration_ms":22299,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that few-step style transfer is made accurate by distilling a single fixed denoising trajectory rather than aligning many forward trajectories at their starts.","keywords":["single trajectory distillation","consistency model","style transfer","diffusion model acceleration","trajectory bank","adversarial loss","video stylization","partial noise editing"],"falsifier":"Measure the student's predicted trajectory against the teacher's at intermediate timesteps for a fixed $\\eta=0.75$: if STD's whole-trajectory error is not measurably smaller than LCM's or TCD's at the same number of steps, the central claim fails. Alternatively, evaluate a model trained at $\\eta=0.75$ at $\\eta=0.8$ without retraining: a sharp drop in style-similarity or aesthetic scores would confirm the single-trajectory specialization and expose the fixed-strength assumption.","tokens_in":14475,"feed_emoji":"🎨","tokens_out":7526,"duration_ms":60347,"temperature":0.7,"pith_summary":"Image and video stylization usually adds a fixed amount of noise to the input and then denoises, which is slow. Prior consistency-model accelerators align only the first step of the student with an imperfect teacher, so errors propagate along the whole trajectory. The paper proposes single-trajectory distillation (STD): fix the noise strength used at inference and make the student self-consistent along the teacher's entire denoising trajectory from that exact point. A trajectory bank stores intermediate states to avoid repeated ODE solves during training, and an asymmetric adversarial loss reduces speckle while boosting saturation. On SDXL-based image and video stylization, STD reports higher style-similarity and aesthetic scores than LCM, TCD, PCM, TDD, Hyper-SD, SDXL-Lightning, and MCM at four to eight function evaluations.","feed_headline":"One fixed denoising trajectory beats many for few-step style transfer","feed_subtitle":"Distilling the teacher's whole path from one noise level lifts style and aesthetic scores at 4–8 steps.","key_machinery":"Single-trajectory distillation (STD) is a consistency-distillation objective defined on the one denoising trajectory that starts from the fixed partial-noise state $x_{\\tau_\\eta}=\\alpha_{\\tau_\\eta}x_0+\\sigma_{\\tau_\\eta}\\epsilon$, instead of on the bundle of forward-diffusion trajectories used by prior consistency models. The loss is $L_{\\text{STD}}=\\|f_\\theta(\\hat{x}^{\\phi,\\eta}_{t_{n+1}},t_{n+1},s)-f_{\\theta^-}(\\hat{x}^{\\phi,\\eta}_{t_n},t_n,s)\\|_2^2$, with $\\hat{x}^{\\phi,\\eta}_{t}$ obtained from the teacher's ODE solver applied to $x_{\\tau_\\eta}$. Two supporting mechanisms carry it: a trajectory bank that stores intermediate states of that single trajectory so training never re-computes them from scratch, and an asymmetric adversarial loss that compares the student's prediction at timestep $s$ to real images noised to $r<s$ via DINO-v2 features.","core_discovery":"The central claim is that in an imperfect denoising model, the PF-ODE trajectories starting from different points on the forward diffusion path are inconsistent, so consistency distillation that starts from forward-diffusion samples fits the teacher badly everywhere except the first step. Since stylization inference always starts from one specific noise level $\\eta$, the paper defines a single trajectory $x_{\\tau_\\eta} \\to \\hat{x}^\\phi_0$ and distills self-consistency along it: the student learns $f_\\theta(x_t,t,s)=f_{\\theta^-}(\\hat{x}^\\phi_t,t,s)$ for $t\\in[0,\\tau_\\eta]$, where $\\hat{x}^\\phi_t$ is produced by denoising $x_{\\tau_\\eta}$ with the teacher, not by re-adding noise to $x_0$. A theorem gives the bound $\\|x_s-\\hat{x}^\\phi_s\\|<C_{t,s}\\cdot\\delta_\\phi$ for a DDIM solver, showing why initial-step alignment fails as $t$ and $s$ separate. With the trajectory bank and the asymmetric adversarial loss, STD achieves the reported gains at 4–8 steps.","pith_inferences":["One could test whether a single model conditioned on $\\eta$ (for example, through a small embedding) regains the per-strength gains without training four separate checkpoints; the paper's bound suggests nearby $\\eta$ trajectories are close, so interpolation may succeed.","The trajectory bank acts as a replay buffer with non-stationary states; systematically varying its size, update rate, and sampling probability $\\rho$ could reveal trade-offs and might explain residual speckle artifacts.","The bound $C_{t,s}\\delta_\\phi$ quantifies how trajectory error grows with the spread of starting timesteps, implying that the advantage of STD over initial-step methods should increase as the teacher model becomes less accurate—a prediction that could be tested with deliberately weak teachers.","Inference in the paper uses a TCD scheduler while training uses a DDIM solver; training with the exact sampling-time solver could close the remaining distribution gap and is a direct, testable variant."],"forward_implications":["At NFE 4–8, stylized images and videos from an SDXL-class model reach higher style-similarity and aesthetic scores than the listed acceleration baselines, lowering the computational barrier for practical stylization.","Because the trajectory bank makes single-trajectory training cheap, the method scales to video stylization, where multi-step ODE solves during training would be prohibitively expensive.","The asymmetric adversarial loss, which constrains predictions at timestep $s$ against real images at $r<s$, is matched to the teacher-trajectory learning target; aligning at timestep 0 instead degrades style and aesthetics, as shown in the paper's ablation table.","The theorem implies that any partial-noise editing task with a fixed noise strength, not just stylization, can use the same single-trajectory approach; the paper names image inpainting as future work."],"supporting_citations":[{"why":"Supplies the consistency-model self-consistency objective that STD modifies and extends to a single fixed trajectory.","marker":"[30]"},{"why":"Defines the latent consistency model distillation loss that STD compares against and refines for partial-noise editing.","marker":"[19]"},{"why":"Introduces SDEdit, the partial-noise editing paradigm that motivates starting every trajectory from a fixed noise level.","marker":"[21]"},{"why":"Provides the DDIM solver used both to define trajectories in training and to derive the bound in the theorem.","marker":"[28]"},{"why":"Presents trajectory consistency distillation, the main trajectory-alignment baseline, and also supplies the scheduler used at STD inference.","marker":"[40]"},{"why":"Offers the phased consistency model baseline and the adversarial-consistency idea that STD contrasts with its asymmetric version.","marker":"[32]"},{"why":"Supplies the video motion consistency model baseline whose training settings and adversarial setup STD adapts.","marker":"[39]"},{"why":"Delivers the DINO-v2 features that the asymmetric adversarial discriminator uses for semantic-level constraints.","marker":"[23]"}],"fun_headline_variants":["Whole-trajectory distillation yields 4-step style transfer","Distill the full denoise path, not just the start, for style","Single trajectory distillation accelerates style transfer to 4-8 steps","One clean path: single trajectory distillation for style","Distill one ODE path, cut style transfer steps to 4"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method assumes that every application will start denoising from exactly the same fixed noise strength the model was trained on, and it trains a separate model for each strength.","fun_headline_variants_meta":{"raw":{"variants":["Whole-trajectory distillation yields 4-step style transfer","Distill the full denoise path, not just the start, for style","Single trajectory distillation accelerates style transfer to 4-8 steps","One clean path: single trajectory distillation for style","Distill one ODE path, cut style transfer steps to 4"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000347,"raw_usage":{"total_tokens":1915,"prompt_tokens":976,"completion_tokens":939,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":852}},"tokens_in":592,"tokens_out":939,"duration_ms":8066,"temperature":1.0,"reasoning_tokens":852,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T04:18:24.925375+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure the student's predicted trajectory against the teacher's at intermediate timesteps for a fixed $\\eta=0.75$: if STD's whole-trajectory error is not measurably smaller than LCM's or TCD's at the same number of steps, the central claim fails. Alternatively, evaluate a model trained at $\\eta=0.75$ at $\\eta=0.8$ without retraining: a sharp drop in style-similarity or aesthetic scores would confirm the single-trajectory specialization and expose the fixed-strength assumption.","supporting_citations":[],"review_version":1}