Pith. sign in

REVIEW 3 cited by

Bigger is not Always Better: Scaling Properties of Latent Diffusion Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2404.01367 v2 pith:2CUJEHAY submitted 2024-04-01 cs.CV cs.LG

classification cs.CVcs.LG
keywords modelsdiffusionsamplingefficiencyfindingsinferencescalinglatent
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We study the scaling properties of latent diffusion models (LDMs) with an emphasis on their sampling efficiency. While improved network architecture and inference algorithms have shown to effectively boost sampling efficiency of diffusion models, the role of model size -- a critical determinant of sampling efficiency -- has not been thoroughly examined. Through empirical analysis of established text-to-image diffusion models, we conduct an in-depth investigation into how model size influences sampling efficiency across varying sampling steps. Our findings unveil a surprising trend: when operating under a given inference budget, smaller models frequently outperform their larger equivalents in generating high-quality results. Moreover, we extend our study to demonstrate the generalizability of the these findings by applying various diffusion samplers, exploring diverse downstream tasks, evaluating post-distilled models, as well as comparing performance relative to training compute. These findings open up new pathways for the development of LDM scaling strategies which can be employed to enhance generative capabilities within limited inference budgets.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Towards Precise Scaling Laws for Video Diffusion Transformers

    cs.CV 2024-11 conditional novelty 6.0 of 10

    Video diffusion transformers follow scaling laws, and tuning learning rate and batch size per model and data size makes those laws precise enough to predict loss and optimal model size at larger scales.

  2. Zoomed In, Diffused Out: Towards Local Degradation-Aware Multi-Diffusion for Extreme Image Super-Resolution

    cs.CV 2024-11 conditional novelty 5.0 of 10

    A training-free recipe uses MultiDiffusion with per-tile degradation-aware text prompts to make frozen text-to-image diffusion models super-resolve images up to 8K.

  3. Foundations of GenIR

    cs.IR 2025-01 unverdicted novelty 1.0 of 10

    A survey chapter proposing that generative AI reshapes information access through two paradigms, information generation and information synthesis.

Pith tools