REVIEW 14 cited by
SegDiff: Image Segmentation with Diffusion Probabilistic Models
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Diffusion Probabilistic Methods are employed for state-of-the-art image generation. In this work, we present a method for extending such models for performing image segmentation. The method learns end-to-end, without relying on a pre-trained backbone. The information in the input image and in the current estimation of the segmentation map is merged by summing the output of two encoders. Additional encoding layers and a decoder are then used to iteratively refine the segmentation map, using a diffusion model. Since the diffusion model is probabilistic, it is applied multiple times, and the results are merged into a final segmentation map. The new method produces state-of-the-art results on the Cityscapes validation set, the Vaihingen building segmentation benchmark, and the MoNuSeg dataset.
Forward citations
Cited by 14 Pith papers
-
Multi-Channel Uncertainty-Weighted Score Matching for Conditional Diffusion in Medical UDA
A UDA method uses Bezier style transfer plus an uncertainty-weighted conditional diffusion model to generate labeled target-style images, improving cross-modality medical segmentation.
-
DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval
A diffusion-based disentanglement method that separates object content from domain style achieves state-of-the-art unsupervised cross-domain image retrieval on three benchmarks.
-
Q-Sched: Pushing the Boundaries of Few-Step Diffusion Models with Quantization-Aware Scheduling
Q-Sched's quantization-aware scheduler with a reference-free JAQ loss lets 2-8 step quantized diffusion models reach lower FID than full-precision baselines.
-
SDMatte: Grafting Diffusion Models for Interactive Matting
SDMatte adapts Stable Diffusion to interactive matting via visual-prompt cross-attention, opacity/coordinate embeddings, and masked self-attention, reporting SOTA results on multiple benchmarks.
-
Flow Stochastic Segmentation Networks
Flow-SSNs model high-rank pixel covariances for ambiguous medical image segmentation by mapping a learned diagonal-Gaussian prior through a lightweight flow, outperforming prior SOTA with fewer parameters.
-
Diffusion-FS: Multimodal Free-Space Prediction via Diffusion for Autonomous Driving
A self-supervised diffusion model predicts multimodal drivable corridors as contour points in monocular images, and beats two segmentation baselines on CARLA and nuScenes.
-
UniSegDiff: Boosting Unified Lesion Segmentation via a Staged Diffusion Model
UniSegDiff uses staged training and inference with alternating mask/noise prediction targets plus STAPLE fusion of multiple samples to reach state-of-the-art lesion segmentation across six datasets and modalities.
-
FreeDNA: Endowing Domain Adaptation of Diffusion-Based Dense Prediction with Training-Free Domain Noise Alignment
Aligning the L2-norm statistics of noise predictions during diffusion sampling improves domain adaptation for dense prediction, with a source-free version guided by high-confidence regions.
-
Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation
Massive activations in DiTs are timestep-driven detail channels; suppressing them guides finer sampling and AdaLN-modulating them yields more discriminative dense features.
-
IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control
The claimed IDC-Net framework is absent; the body text is an unrelated instance-segmentation paper.
-
From Variability To Accuracy: Conditional Bernoulli Diffusion Models with Consensus-Driven Correction for Thin Structure Segmentation
A consensus-driven correction on top of a conditional Bernoulli diffusion model improves recall in thin orbital bone segmentation, though not all metrics beat baselines.
-
LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion
LangScene-X generates RGB, normal, and semantic videos from sparse views to reconstruct 3D language-embedded Gaussian fields that support open-ended text queries.
-
Volumetric Directional Diffusion: Anchoring Uncertainty Quantification in Anatomical Consensus for Ambiguous Medical Image Segmentation
Anchoring a 3D diffusion model to a deterministic consensus prior and sampling only boundary residuals improves uncertainty alignment while keeping anatomical structure intact.
-
SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3
A frozen DINOv3 backbone plus a simple MLP head reportedly beats specialized segmentation models on six benchmarks, but the evidence lacks statistical rigor.
Discussion (0). Sign in to comment.