REVIEW 6 cited by
Marigold: Affordable Adaptation of Diffusion-Based Image Generators for Image Analysis
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The success of deep learning in computer vision over the past decade has hinged on large labeled datasets and strong pretrained models. In data-scarce settings, the quality of these pretrained models becomes crucial for effective transfer learning. Image classification and self-supervised learning have traditionally been the primary methods for pretraining CNNs and transformer-based architectures. Recently, the rise of text-to-image generative models, particularly those using denoising diffusion in a latent space, has introduced a new class of foundational models trained on massive, captioned image datasets. These models' ability to generate realistic images of unseen content suggests they possess a deep understanding of the visual world. In this work, we present Marigold, a family of conditional generative models and a fine-tuning protocol that extracts the knowledge from pretrained latent diffusion models like Stable Diffusion and adapts them for dense image analysis tasks, including monocular depth estimation, surface normals prediction, and intrinsic decomposition. Marigold requires minimal modification of the pre-trained latent diffusion model's architecture, trains with small synthetic datasets on a single GPU over a few days, and demonstrates state-of-the-art zero-shot generalization. Project page: https://marigoldcomputervision.github.io
Forward citations
Cited by 6 Pith papers
-
GeoStereo: A Unified Stereo Geometry Estimation Framework for Disparity and Surface Normal
A unified stereo framework couples feed-forward disparity matching with a diffusion-based normal estimator through disparity-to-normal initialization and warped right-view conditioning, claiming zero-shot SOTA on seve...
-
Towards Consistent Video Geometry Estimation
ViGeo is a feed-forward transformer for video geometry that introduces dynamic chunking attention and a completion-based data refinement framework to achieve SOTA on depth, normals, and point map estimation.
-
GSFix3D: Diffusion-Guided Repair of Novel Views in Gaussian Splatting
A per-scene fine-tuned latent diffusion model with dual mesh-3DGS conditioning and random mask augmentation improves novel-view repair in Gaussian Splatting, outperforming DIFIX baselines on ScanNet++ and Replica.
-
Single-Step Latent Diffusion for Underwater Image Restoration
SLURPP combines pretrained latent diffusion priors with a physics-based scene-medium decomposition to restore underwater images in one inference step, beating prior diffusion methods in speed and quality.
-
PoseMaster: A Unified 3D Native Framework for Stylized Pose Generation
PoseMaster produces a 3D character mesh from one image and a target 3D skeleton, preserving identity and pose in a single unified model, and it outperforms two-stage 2D-to-3D baselines on the VRoid pose canonicalizati...
-
LTM: Large-scale Terrain Model for Wildfire-prone Landscapes
A ray-tracing pipeline aligns ground-level image pixels to outdated DEM rasters for real-time 3D terrain reconstruction in wildfire zones, validated primarily through a custom simulator.
Discussion (0). Sign in to comment.