Pith. sign in

REVIEW 2 cited by

Boosting Diffusion Models with Moving Average Sampling in Frequency Domain

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.17870 v1 pith:GU6LMAXR submitted 2024-03-26 cs.CV cs.MM

classification cs.CVcs.MM
keywords averagemodelsmovingdiffusionsamplesfrequencymasfsampling
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models have recently brought a powerful revolution in image generation. Despite showing impressive generative capabilities, most of these models rely on the current sample to denoise the next one, possibly resulting in denoising instability. In this paper, we reinterpret the iterative denoising process as model optimization and leverage a moving average mechanism to ensemble all the prior samples. Instead of simply applying moving average to the denoised samples at different timesteps, we first map the denoised samples to data space and then perform moving average to avoid distribution shift across timesteps. In view that diffusion models evolve the recovery from low-frequency components to high-frequency details, we further decompose the samples into different frequency components and execute moving average separately on each component. We name the complete approach "Moving Average Sampling in Frequency domain (MASF)". MASF could be seamlessly integrated into mainstream pre-trained diffusion models and sampling schedules. Extensive experiments on both unconditional and conditional diffusion models demonstrate that our MASF leads to superior performances compared to the baselines, with almost negligible additional complexity cost.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CoCA redistributes a single final image reward across denoising steps using cosine similarity between intermediate and final latents, improving RL fine-tuning sample efficiency on four human preference rewards.

  2. ANT: Adaptive Neural Temporal-Aware Text-to-Motion Model

    cs.CV 2025-06 conditional novelty 5.0 of 10

    ANT makes text embeddings change across denoising steps and schedules classifier-free guidance to decay, improving text-motion alignment in diffusion text-to-motion models.

Pith tools