Pith. sign in

REVIEW 8 cited by

Improved Precision and Recall Metric for Assessing Generative Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1904.06991 v3 pith:2NLZQLG5 submitted 2019-04-15 stat.ML cs.LGcs.NE

classification stat.MLcs.LGcs.NE
keywords metricestimategenerativeidentifyimprovedmethodsmodelquality
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

The ability to automatically estimate the quality and coverage of the samples produced by a generative model is a vital requirement for driving algorithm research. We present an evaluation metric that can separately and reliably measure both of these aspects in image generation tasks by forming explicit, non-parametric representations of the manifolds of real and generated data. We demonstrate the effectiveness of our metric in StyleGAN and BigGAN by providing several illustrative examples where existing metrics yield uninformative or contradictory results. Furthermore, we analyze multiple design variants of StyleGAN to better understand the relationships between the model architecture, training methods, and the properties of the resulting sample distribution. In the process, we identify new variants that improve the state-of-the-art. We also perform the first principled analysis of truncation methods and identify an improved method. Finally, we extend our metric to estimate the perceptual quality of individual samples, and use this to study latent space interpolations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents

    cs.LG 2026-07 conditional novelty 6.0 of 10

    FairDiffuseVQVAE reaches state-of-the-art fairness on the standard tabular benchmark (DPR 0.702, EOR 0.686) by uniform protected-attribute sampling at inference, paying ~15 AUC points of utility.

  2. GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification

    cs.CV 2026-07 conditional novelty 6.0 of 10

    GenSyn10 provides 60k CIFAR-10-aligned images from FLUX.2, HunyuanImage-3.0, and Qwen-Image-2512, showing detectors lose 4–18 points of accuracy on an unseen generator.

  3. DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A new evaluation framework, DIMCIM, measures default-mode diversity and prompted generalization in text-to-image models, finding a scale trade-off and a 0.85 correlation between default diversity and training data diversity.

  4. Phantom Evidence: How and Why Generative AI Manufactures False Positives in Science

    q-bio.NC 2026-07 conditional novelty 5.0 of 10

    Evidence from a convincing-looking AI output is bounded by how often such outputs occur when the target is absent; generative AI inflates that denominator, and the paper calls the resulting overestimate 'phantom evidence'.

  5. HistoFID- Calibrating Frechet-distance evaluation across pathology foundation models

    eess.IV 2026-07 conditional novelty 5.0 of 10

    Raw Fréchet distances in pathology vary ~30-fold across encoders; dividing by each encoder's own within-cohort floor cuts cross-encoder variation by ~89% within and ~58% across cohorts.

  6. Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    SKD-CAG erases adversarial text triggers from diffusion models by distilling the model's own clean outputs through cross-attention guidance, claiming 100% and 93% removal for pixel and style backdoors.

  7. Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models

    cs.CV 2025-06 conditional novelty 5.0 of 10

    A probabilistic overlap measure for object positions yields a human-aligned spatial relationship metric and a training-free generation guidance method for text-to-image models.

  8. Paired and Unpaired Image to Image Translation using Generative Adversarial Networks

    cs.CV 2025-05 conditional novelty 2.0 of 10

    Using Pix2Pix and CycleGAN on four standard datasets, the paper finds that L1 loss and paired supervision beat alternatives in FID, precision, and recall, a conclusion already found in the cited literature.

Pith tools