Pith. sign in

REVIEW 4 cited by

Unleashing Text-to-Image Diffusion Models for Visual Perception

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2303.02153 v1 pith:S6XBZFD6 submitted 2023-03-03 cs.CV

classification cs.CV
keywords diffusionmodelspre-trainedvisualperceptiontext-to-imagefeaturessegmentation
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion models (DMs) have become the new trend of generative models and have demonstrated a powerful ability of conditional synthesis. Among those, text-to-image diffusion models pre-trained on large-scale image-text pairs are highly controllable by customizable prompts. Unlike the unconditional generative models that focus on low-level attributes and details, text-to-image diffusion models contain more high-level knowledge thanks to the vision-language pre-training. In this paper, we propose VPD (Visual Perception with a pre-trained Diffusion model), a new framework that exploits the semantic information of a pre-trained text-to-image diffusion model in visual perception tasks. Instead of using the pre-trained denoising autoencoder in a diffusion-based pipeline, we simply use it as a backbone and aim to study how to take full advantage of the learned knowledge. Specifically, we prompt the denoising decoder with proper textual inputs and refine the text features with an adapter, leading to a better alignment to the pre-trained stage and making the visual contents interact with the text prompts. We also propose to utilize the cross-attention maps between the visual features and the text features to provide explicit guidance. Compared with other pre-training methods, we show that vision-language pre-trained diffusion models can be faster adapted to downstream visual perception tasks using the proposed VPD. Extensive experiments on semantic segmentation, referring image segmentation and depth estimation demonstrates the effectiveness of our method. Notably, VPD attains 0.254 RMSE on NYUv2 depth estimation and 73.3% oIoU on RefCOCO-val referring image segmentation, establishing new records on these two benchmarks. Code is available at https://github.com/wl-zhao/VPD

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Video Depth without Video Models

    cs.CV 2024-11 conditional novelty 6.0 of 10

    A single-image latent diffusion model extended with cross-frame attention and global scale-shift alignment produces state-of-the-art video depth without a video diffusion model.

  2. Marigold-DC: Zero-Shot Monocular Depth Completion with Guided Diffusion

    cs.CV 2024-12 conditional novelty 5.0 of 10

    Sparse depth points injected as test-time guidance into a pretrained monocular depth diffusion model achieve strong zero-shot depth completion across indoor and outdoor scenes.

  3. An Efficient Framework for Enhancing Discriminative Models via Diffusion Techniques

    cs.CV 2024-12 conditional novelty 5.0 of 10

    DBMEF improves discriminative image classifiers by reclassifying low-confidence predictions with a diffusion model, without any training.

  4. Vision Generalist Model: A Survey

    cs.CV 2025-06 conditional novelty 3.0 of 10

    A structured review of vision generalist models, classifying them into encoding-based and sequence-to-sequence frameworks and summarizing datasets, benchmarks, techniques, and open problems.

Pith tools