Pith. sign in

REVIEW 11 cited by

Inductive Moment Matching

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2503.07565 v7 pith:GTGE4LLL submitted 2025-03-10 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords modelsmatchingdiffusionfew-stepinductiveinferencemodelmoment
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models and Flow Matching generate high-quality samples but are slow at inference, and distilling them into few-step models often leads to instability and extensive tuning. To resolve these trade-offs, we propose Inductive Moment Matching (IMM), a new class of generative models for one- or few-step sampling with a single-stage training procedure. Unlike distillation, IMM does not require pre-training initialization and optimization of two networks; and unlike Consistency Models, IMM guarantees distribution-level convergence and remains stable under various hyperparameters and standard model architectures. IMM surpasses diffusion models on ImageNet-256x256 with 1.99 FID using only 8 inference steps and achieves state-of-the-art 2-step FID of 1.98 on CIFAR-10 for a model trained from scratch.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 11 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. DriftXpress: Faster Drifting Models via Projected RKHS Fields

    cs.LG 2026-05 unverdicted novelty 7.0 of 10

    DriftXpress approximates the attraction field of drifting models with a Nyström landmark projection, reducing training time by 2.6–6.7× at comparable FID.

  2. Understanding LoRA as Knowledge Memory: An Empirical Analysis

    cs.LG 2026-03 conditional novelty 7.0 of 10

    LoRA modules function as composable knowledge memories for LLMs with measurable storage capacity, internalization efficiency, and advantages in multi-module long-context reasoning.

  3. Amortized Moment Matching for Visual Generation

    cs.LG 2026-07 accept novelty 6.0 of 10

    Amortized Fréchet Distance uses neural nets to match conditional means and covariances, yielding stronger one-step visual generators than explicit FD-loss or multi-step teachers.

  4. Flow Map Learning via Nongradient Vector Flow

    cs.LG 2026-07 conditional novelty 6.0 of 10

    SGFlow learns the integral map of a probability-flow ODE via a stop-gradient loss whose only stationary point is the true flow map, and it reaches the best-in-comparison FID at 10 steps on CIFAR-10.

  5. Parallel Decoding Distillation for Fast Image and Video Generation

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A trajectory-based distillation method trains a student to predict multiple mean velocities per network evaluation, enabling 4-8 step generation with competitive quality and improved diversity.

  6. Autoregressive One-Step Generative Modeling for Dynamical System Forecasting

    cs.LG 2026-05 unverdicted novelty 6.0 of 10

    MeLISA extends pixel-space MeanFlow to one-step window-conditioned autoregressive forecasting, improving long-horizon turbulence statistics over neural-operator baselines.

  7. MENO: MeanFlow-Enhanced Neural Operators for Dynamical Systems

    cs.LG 2026-04 unverdicted novelty 6.0 of 10

    MENO restores multi-scale structure in neural-operator PDE surrogates via one-step improved MeanFlow, claiming up to 2× better power-spectrum accuracy and up to 14× faster inference than DDIM enhancement.

  8. Dual-End Consistency Model

    cs.CV 2026-02 unverdicted novelty 6.0 of 10

    DE-CM trains a flow-map consistency model on three sub-trajectories (coupling, instantaneous, noise-to-noisy) and reports 1.70 FID one-step on ImageNet 256.

  9. Scalable GANs with Transformers

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A transformer-only GAN trained in VAE latent space with multi-level noise supervision and width-scaled learning rates achieves FID 2.96 on ImageNet-256 in 40 epochs.

  10. Transition Models: Rethinking the Generative Learning Objective

    cs.LG 2025-09 conditional novelty 6.0 of 10

    TiM trains a single diffusion-type model on arbitrary time-interval transitions, achieving strong one-step and multi-step text-to-image generation with 865M parameters.

  11. Diffuse and Disperse: Image Generation with Representation Regularization

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Adding a dispersion regularizer to intermediate features of diffusion and flow models consistently improves FID on ImageNet and one-step generation, with no additional parameters, pretraining, or external data.

Pith tools