Pith. sign in

REVIEW 14 cited by

SegDiff: Image Segmentation with Diffusion Probabilistic Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2112.00390 v3 pith:DVN2QY6B submitted 2021-12-01 cs.CV cs.AIcs.LG

classification cs.CVcs.AIcs.LG
keywords segmentationdiffusionimagemethodprobabilisticmergedmodelmodels
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Diffusion Probabilistic Methods are employed for state-of-the-art image generation. In this work, we present a method for extending such models for performing image segmentation. The method learns end-to-end, without relying on a pre-trained backbone. The information in the input image and in the current estimation of the segmentation map is merged by summing the output of two encoders. Additional encoding layers and a decoder are then used to iteratively refine the segmentation map, using a diffusion model. Since the diffusion model is probabilistic, it is applied multiple times, and the results are merged into a final segmentation map. The new method produces state-of-the-art results on the Cityscapes validation set, the Vaihingen building segmentation benchmark, and the MoNuSeg dataset.

Discussion (0). Sign in to comment.

Forward citations

Cited by 14 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Multi-Channel Uncertainty-Weighted Score Matching for Conditional Diffusion in Medical UDA

    cs.CV 2025-09 conditional novelty 7.0 of 10

    A UDA method uses Bezier style transfer plus an uncertainty-weighted conditional diffusion model to generate labeled target-style images, improving cross-modality medical segmentation.

  2. DUDE: Diffusion-Based Unsupervised Cross-Domain Image Retrieval

    cs.CV 2025-09 conditional novelty 6.0 of 10

    A diffusion-based disentanglement method that separates object content from domain style achieves state-of-the-art unsupervised cross-domain image retrieval on three benchmarks.

  3. Q-Sched: Pushing the Boundaries of Few-Step Diffusion Models with Quantization-Aware Scheduling

    cs.CV 2025-09 conditional novelty 6.0 of 10

    Q-Sched's quantization-aware scheduler with a reference-free JAQ loss lets 2-8 step quantized diffusion models reach lower FID than full-precision baselines.

  4. SDMatte: Grafting Diffusion Models for Interactive Matting

    cs.CV 2025-08 conditional novelty 6.0 of 10

    SDMatte adapts Stable Diffusion to interactive matting via visual-prompt cross-attention, opacity/coordinate embeddings, and masked self-attention, reporting SOTA results on multiple benchmarks.

  5. Flow Stochastic Segmentation Networks

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Flow-SSNs model high-rank pixel covariances for ambiguous medical image segmentation by mapping a learned diagonal-Gaussian prior through a lightweight flow, outperforming prior SOTA with fewer parameters.

  6. Diffusion-FS: Multimodal Free-Space Prediction via Diffusion for Autonomous Driving

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A self-supervised diffusion model predicts multimodal drivable corridors as contour points in monocular images, and beats two segmentation baselines on CARLA and nuScenes.

  7. UniSegDiff: Boosting Unified Lesion Segmentation via a Staged Diffusion Model

    eess.IV 2025-07 conditional novelty 6.0 of 10

    UniSegDiff uses staged training and inference with alternating mask/noise prediction targets plus STAPLE fusion of multiple samples to reach state-of-the-art lesion segmentation across six datasets and modalities.

  8. FreeDNA: Endowing Domain Adaptation of Diffusion-Based Dense Prediction with Training-Free Domain Noise Alignment

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Aligning the L2-norm statistics of noise predictions during diffusion sampling improves domain adaptation for dense prediction, with a source-free version guided by high-confidence regions.

  9. Awakening Diffusion Transformers: Eliciting Stronger Generation and Understanding via Massive Activation Modulation

    cs.CV 2026-07 conditional novelty 5.0 of 10

    Massive activations in DiTs are timestep-driven detail channels; suppressing them guides finer sampling and AdaLN-modulating them yields more discriminative dense features.

  10. IDCNet: Guided Video Diffusion for Metric-Consistent RGBD Scene Generation with Precise Camera Control

    cs.CV 2025-08 unverdicted novelty 5.0 of 10

    The claimed IDC-Net framework is absent; the body text is an unrelated instance-segmentation paper.

  11. From Variability To Accuracy: Conditional Bernoulli Diffusion Models with Consensus-Driven Correction for Thin Structure Segmentation

    eess.IV 2025-07 conditional novelty 5.0 of 10

    A consensus-driven correction on top of a conditional Bernoulli diffusion model improves recall in thin orbital bone segmentation, though not all metrics beat baselines.

  12. LangScene-X: Reconstruct Generalizable 3D Language-Embedded Scenes with TriMap Video Diffusion

    cs.CV 2025-07 conditional novelty 5.0 of 10

    LangScene-X generates RGB, normal, and semantic videos from sparse views to reconstruct 3D language-embedded Gaussian fields that support open-ended text queries.

  13. Volumetric Directional Diffusion: Anchoring Uncertainty Quantification in Anatomical Consensus for Ambiguous Medical Image Segmentation

    cs.CV 2026-03 conditional novelty 4.0 of 10

    Anchoring a 3D diffusion model to a deterministic consensus prior and sampling only boundary residuals improves uncertainty alignment while keeping anatomical structure intact.

  14. SegDINO: An Efficient Design for Medical and Natural Image Segmentation with DINO-V3

    cs.CV 2025-08 reject novelty 4.0 of 10

    A frozen DINOv3 backbone plus a simple MLP head reportedly beats specialized segmentation models on six benchmarks, but the evidence lacks statistical rigor.

Pith tools