REVIEW 2 cited by
Your Diffusion Model is Secretly a Zero-Shot Classifier
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhibit impressive compositional generalization abilities. Almost all use cases thus far have solely focused on sampling; however, diffusion models can also provide conditional density estimates, which are useful for tasks beyond image generation. In this paper, we show that the density estimates from large-scale text-to-image diffusion models like Stable Diffusion can be leveraged to perform zero-shot classification without any additional training. Our generative approach to classification, which we call Diffusion Classifier, attains strong results on a variety of benchmarks and outperforms alternative methods of extracting knowledge from diffusion models. Although a gap remains between generative and discriminative approaches on zero-shot recognition tasks, our diffusion-based approach has significantly stronger multimodal compositional reasoning ability than competing discriminative approaches. Finally, we use Diffusion Classifier to extract standard classifiers from class-conditional diffusion models trained on ImageNet. Our models achieve strong classification performance using only weak augmentations and exhibit qualitatively better "effective robustness" to distribution shift. Overall, our results are a step toward using generative over discriminative models for downstream tasks. Results and visualizations at https://diffusion-classifier.github.io/
Forward citations
Cited by 2 Pith papers
-
ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features
ConceptAttention shows that linear projections in the output space of DiT attention layers yield sharper concept-localizing saliency maps than cross-attention maps, reaching state-of-the-art zero-shot segmentation.
-
Conditional Diffusion Models are Medical Image Classifiers that Provide Explainability and Uncertainty for Free
Conditional diffusion models can classify medical images by comparing reconstruction errors, and the per-noise-level majority vote also yields explanation and uncertainty byproducts.
Discussion (0). Continue with ORCID to comment.