Pith. sign in

REVIEW 2 cited by

FreSca: Scaling in Frequency Space Enhances Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2504.02154 v3 pith:5TXRWEV6 submitted 2025-04-02 cs.CV

classification cs.CV
keywords controldiffusionfrescalatentmodelsspaceacrossdifference
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Latent diffusion models (LDMs) have achieved remarkable success in a variety of image tasks, yet achieving fine-grained, disentangled control over global structures versus fine details remains challenging. This paper explores frequency-based control within latent diffusion models. We first systematically analyze frequency characteristics across pixel space, VAE latent space, and internal LDM representations. This reveals that the "noise difference" term, derived from classifier-free guidance at each step t, is a uniquely effective and semantically rich target for manipulation. Building on this insight, we introduce FreSca, a novel and plug-and-play framework that decomposes noise difference into low- and high-frequency components and applies independent scaling factors to them via spatial or energy-based cutoffs. Essentially, FreSca operates without any model retraining or architectural change, offering model- and task-agnostic control. We demonstrate its versatility and effectiveness in improving generation quality and structural emphasis on multiple architectures (e.g., SD3, SDXL) and across applications including image generation, editing, depth estimation, and video synthesis, thereby unlocking a new dimension of expressive control within LDMs.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Articulate3D: Zero-Shot Text-Driven 3D Object Posing

    cs.CV 2025-08 conditional novelty 6.0 of 10

    A zero-shot pipeline that reposes 3D meshes by generating text-conditioned target images with rewired multi-view attention and aligning mesh keypoints to them.

  2. Step-level Reward for Free in RL-based T2I Diffusion Model Fine-tuning

    cs.CV 2025-05 conditional novelty 6.0 of 10

    CoCA redistributes a single final image reward across denoising steps using cosine similarity between intermediate and final latents, improving RL fine-tuning sample efficiency on four human preference rewards.

Pith tools