Pith. sign in

REVIEW 1 cited by

Are Diffusion Models Vision-And-Language Reasoners?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2305.16397 v3 pith:MLFV2KWQ submitted 2023-05-25 cs.CV cs.AIcs.CL

classification cs.CVcs.AIcs.CL
keywords diffusionmodelsevaluationstablebenchmarkgenerativetasksvision-and-language
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Text-conditioned image generation models have recently shown immense qualitative success using denoising diffusion processes. However, unlike discriminative vision-and-language models, it is a non-trivial task to subject these diffusion-based generative models to automatic fine-grained quantitative evaluation of high-level phenomena such as compositionality. Towards this goal, we perform two innovations. First, we transform diffusion-based models (in our case, Stable Diffusion) for any image-text matching (ITM) task using a novel method called DiffusionITM. Second, we introduce the Generative-Discriminative Evaluation Benchmark (GDBench) benchmark with 7 complex vision-and-language tasks, bias evaluation and detailed analysis. We find that Stable Diffusion + DiffusionITM is competitive on many tasks and outperforms CLIP on compositional tasks like like CLEVR and Winoground. We further boost its compositional performance with a transfer setup by fine-tuning on MS-COCO while retaining generative capabilities. We also measure the stereotypical bias in diffusion models, and find that Stable Diffusion 2.1 is, for the most part, less biased than Stable Diffusion 1.5. Overall, our results point in an exciting direction bringing discriminative and generative model evaluation closer. We will release code and benchmark setup soon.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Conditional Diffusion Models are Medical Image Classifiers that Provide Explainability and Uncertainty for Free

    cs.CV 2025-02 conditional novelty 4.0 of 10

    Conditional diffusion models can classify medical images by comparing reconstruction errors, and the per-noise-level majority vote also yields explanation and uncertainty byproducts.

Pith tools