REVIEW 8 cited by
Improved Precision and Recall Metric for Assessing Generative Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The ability to automatically estimate the quality and coverage of the samples produced by a generative model is a vital requirement for driving algorithm research. We present an evaluation metric that can separately and reliably measure both of these aspects in image generation tasks by forming explicit, non-parametric representations of the manifolds of real and generated data. We demonstrate the effectiveness of our metric in StyleGAN and BigGAN by providing several illustrative examples where existing metrics yield uninformative or contradictory results. Furthermore, we analyze multiple design variants of StyleGAN to better understand the relationships between the model architecture, training methods, and the properties of the resulting sample distribution. In the process, we identify new variants that improve the state-of-the-art. We also perform the first principled analysis of truncation methods and identify an improved method. Finally, we extend our metric to estimate the perceptual quality of individual samples, and use this to study latent space interpolations.
Forward citations
Cited by 8 Pith papers
-
FairDiffuseVQVAE: Sampling-Time Fairness in Tabular Diffusion via Conditional Refinement of Vector-Quantized Latents
FairDiffuseVQVAE reaches state-of-the-art fairness on the standard tabular benchmark (DPR 0.702, EOR 0.686) by uniform protected-attribute sampling at inference, paying ~15 AUC points of utility.
-
GenSyn10: A Multi-Generative AI Dataset For Benchmarking Image Classification
GenSyn10 provides 60k CIFAR-10-aligned images from FLUX.2, HunyuanImage-3.0, and Qwen-Image-2512, showing detectors lose 4–18 points of accuracy on an unseen generator.
-
DIMCIM: A Quantitative Evaluation Framework for Default-mode Diversity and Generalization in Text-to-Image Generative Models
A new evaluation framework, DIMCIM, measures default-mode diversity and prompted generalization in text-to-image models, finding a scale trade-off and a 0.85 correlation between default diversity and training data diversity.
-
Phantom Evidence: How and Why Generative AI Manufactures False Positives in Science
Evidence from a convincing-looking AI output is bounded by how often such outputs occur when the target is absent; generative AI inflates that denominator, and the paper calls the resulting overestimate 'phantom evidence'.
-
HistoFID- Calibrating Frechet-distance evaluation across pathology foundation models
Raw Fréchet distances in pathology vary ~30-fold across encoders; dividing by each encoder's own within-cohort floor cuts cross-encoder variation by ~89% within and ~58% across cohorts.
-
Sealing The Backdoor: Unlearning Adversarial Text Triggers In Diffusion Models Using Knowledge Distillation
SKD-CAG erases adversarial text triggers from diffusion models by distilling the model's own clean outputs through cross-attention guidance, claiming 100% and 93% removal for pixel and style backdoors.
-
Why Settle for Mid: A Probabilistic Viewpoint to Spatial Relationship Alignment in Text-to-image Models
A probabilistic overlap measure for object positions yields a human-aligned spatial relationship metric and a training-free generation guidance method for text-to-image models.
-
Paired and Unpaired Image to Image Translation using Generative Adversarial Networks
Using Pix2Pix and CycleGAN on four standard datasets, the paper finds that L1 loss and paired supervision beat alternatives in FID, precision, and recall, a conclusion already found in the cited literature.
Discussion (0). Continue with ORCID to comment.