Pith. sign in

REVIEW 3 cited by

Visual Generation Without Guidance

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.15420 v2 pith:6DYJZXU3 submitted 2025-01-26 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords modelsvisualsamplingacrossconditionaldirectlyguidanceguidance-free
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Classifier-Free Guidance (CFG) has been a default technique in various visual generative models, yet it requires inference from both conditional and unconditional models during sampling. We propose to build visual models that are free from guided sampling. The resulting algorithm, Guidance-Free Training (GFT), matches the performance of CFG while reducing sampling to a single model, halving the computational cost. Unlike previous distillation-based approaches that rely on pretrained CFG networks, GFT enables training directly from scratch. GFT is simple to implement. It retains the same maximum likelihood objective as CFG and differs mainly in the parameterization of conditional models. Implementing GFT requires only minimal modifications to existing codebases, as most design choices and hyperparameters are directly inherited from CFG. Our extensive experiments across five distinct visual models demonstrate the effectiveness and versatility of GFT. Across domains of diffusion, autoregressive, and masked-prediction modeling, GFT consistently achieves comparable or even lower FID scores, with similar diversity-fidelity trade-offs compared with CFG baselines, all while being guidance-free. Code will be available at https://github.com/thu-ml/GFT.

Discussion (0). Sign in to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Perceptual Flow Matching for Few-Step Generative Modeling

    cs.CV 2026-07 conditional novelty 6.5 of 10

    Supervising flow matching on decoded samples in pretrained perceptual feature space yields high-quality 4–8 step generators without teacher models or distillation.

  2. UniGen-AR: Unifying Visual Generation with Auto-Regressive Modeling

    cs.CV 2026-07 conditional novelty 6.0 of 10

    An MLLM-conditioned next-scale VAR decoder handles 15+ unified visual generation tasks with competitive quality and substantially lower latency than diffusion baselines.

  3. FlowerDance: MeanFlow for Efficient and Refined 3D Dance Generation

    cs.CV 2025-11 conditional novelty 5.0 of 10

    FlowerDance pairs MeanFlow few-step flow matching with a bidirectional Mamba backbone and physical-consistency losses, reporting state-of-the-art dance quality at 2008 FPS on FineDance and AIST++.

Pith tools