Pith. sign in

REVIEW 8 cited by

What Matters When Repurposing Diffusion Models for General Dense Perception Tasks?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2403.06090 v4 pith:2L4IJG76 submitted 2024-03-10 cs.CV

classification cs.CV
keywords perceptiontasksdiffusionfine-tuningdensevisualdatamodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Extensive pre-training with large data is indispensable for downstream geometry and semantic visual perception tasks. Thanks to large-scale text-to-image (T2I) pretraining, recent works show promising results by simply fine-tuning T2I diffusion models for dense perception tasks. However, several crucial design decisions in this process still lack comprehensive justification, encompassing the necessity of the multi-step stochastic diffusion mechanism, training strategy, inference ensemble strategy, and fine-tuning data quality. In this work, we conduct a thorough investigation into critical factors that affect transfer efficiency and performance when using diffusion priors. Our key findings are: 1) High-quality fine-tuning data is paramount for both semantic and geometry perception tasks. 2) The stochastic nature of diffusion models has a slightly negative impact on deterministic visual perception tasks. 3) Apart from fine-tuning the diffusion model with only latent space supervision, task-specific image-level supervision is beneficial to enhance fine-grained details. These observations culminate in the development of GenPercept, an effective deterministic one-step fine-tuning paradigm tailed for dense visual perception tasks. Different from the previous multi-step methods, our paradigm has a much faster inference speed, and can be seamlessly integrated with customized perception decoders and loss functions for image-level supervision, which is critical to improving the fine-grained details of predictions. Comprehensive experiments on diverse dense visual perceptual tasks, including monocular depth estimation, surface normal estimation, image segmentation, and matting, are performed to demonstrate the remarkable adaptability and effectiveness of our proposed method.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 8 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Unified Video Dense Prediction from Disjoint Data

    cs.CV 2026-07 conditional novelty 7.0 of 10

    A single video backbone predicts eight dense scene tasks from separate single-task datasets via latent distillation from diffusion-based specialists, with no co-annotated data or pseudo-labels.

  2. Video Generation Models are General-Purpose Vision Learners

    cs.CV 2026-07 conditional novelty 6.0 of 10

    A video-diffusion backbone fine-tuned as a single-step multi-task perceiver matches or beats specialists on depth, normals, pose and segmentation, with high data efficiency and sim-to-real transfer.

  3. LuxDiT: Lighting Estimation with Video Diffusion Transformer

    cs.GR 2025-09 conditional novelty 6.0 of 10

    A video diffusion transformer fine-tuned on synthetic and real data predicts HDR environment maps from images/videos, cutting peak light-direction error by roughly 45% on sunny outdoor scenes versus DiffusionLight.

  4. SDMatte: Grafting Diffusion Models for Interactive Matting

    cs.CV 2025-08 conditional novelty 6.0 of 10

    SDMatte adapts Stable Diffusion to interactive matting via visual-prompt cross-attention, opacity/coordinate embeddings, and masked self-attention, reporting SOTA results on multiple benchmarks.

  5. BenchDepth: Are We on the Right Way to Evaluate Depth Foundation Models?

    cs.CV 2025-07 conditional novelty 6.0 of 10

    BenchDepth evaluates eight depth foundation models by their performance on five downstream tasks, finding Depth Anything V2's relative version to be the most practically useful.

  6. UniGeo: Taming Video Diffusion for Unified Consistent Geometry Estimation

    cs.CV 2025-05 conditional novelty 6.0 of 10

    Fine-tuning a pretrained video diffusion transformer to predict geometry in one shared global frame produces consistent, camera-free surface normals and coordinates across entire video clips.

  7. FUMO: Prior-Modulated Diffusion for Single Image Reflection Removal

    cs.CV 2026-03 conditional novelty 5.0 of 10

    A coarse-to-fine diffusion SIRR method gates ControlNet residuals with a VLM intensity prior times a multi-scale high-frequency prior, then refines geometry and detail in image space.

  8. DidSee: Diffusion-Based Depth Completion for Material-Agnostic Robotic Perception and Manipulation

    cs.CV 2025-06 conditional novelty 5.0 of 10

    DidSee is a diffusion-based depth completion model that combines a zero terminal-SNR noise scheduler, single-step training, and a semantic segmentation enhancer to achieve state-of-the-art results on non-Lambertian objects.

Pith tools