Pith. sign in

REVIEW 5 cited by

DFormer: Diffusion-guided Transformer for Universal Image Segmentation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2306.03437 v2 pith:57L6ZMRF submitted 2023-06-06 cs.CV

classification cs.CV
keywords segmentationdformermasksimagediffusion-baseduniversaldenoisingfeatures
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

This paper introduces an approach, named DFormer, for universal image segmentation. The proposed DFormer views universal image segmentation task as a denoising process using a diffusion model. DFormer first adds various levels of Gaussian noise to ground-truth masks, and then learns a model to predict denoising masks from corrupted masks. Specifically, we take deep pixel-level features along with the noisy masks as inputs to generate mask features and attention masks, employing diffusion-based decoder to perform mask prediction gradually. At inference, our DFormer directly predicts the masks and corresponding categories from a set of randomly-generated masks. Extensive experiments reveal the merits of our proposed contributions on different image segmentation tasks: panoptic segmentation, instance segmentation, and semantic segmentation. Our DFormer outperforms the recent diffusion-based panoptic segmentation method Pix2Seq-D with a gain of 3.6% on MS COCO val2017 set. Further, DFormer achieves promising semantic segmentation performance outperforming the recent diffusion-based method by 2.2% on ADE20K val set. Our source code and models will be publicly on https://github.com/cp3wan/DFormer

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. D3RM: A Discrete Denoising Diffusion Refinement Model for Piano Transcription

    cs.SD 2025-01 conditional novelty 6.0 of 10

    D3RM uses discrete diffusion with neighborhood attention and an asymmetric train/inference masking schedule to improve piano transcription F1 over a feed-forward baseline and DiffRoll on MAESTRO.

  2. PreSem-Surf: RGB-D Surface Reconstruction with Progressive Semantic Modeling and SG-MLP Pre-Rendering Mechanism

    cs.GR 2025-08 unverdicted novelty 5.0 of 10

    PreSem-Surf combines RGB, depth, and semantic labels with an MLP pre-rendering mechanism to reconstruct scene surfaces, reporting best C-L1, F-score, and IoU on seven synthetic scenes.

  3. D-Cube: Exploiting Hyper-Features of Diffusion Model for Robust Medical Classification

    cs.CV 2024-11 conditional novelty 5.0 of 10

    D-Cube combines selected diffusion-model feature maps with ResNet sub-features and custom losses to improve medical image classification.

  4. Unleashing Diffusion and State Space Models for Medical Image Segmentation

    cs.CV 2025-06 conditional novelty 4.0 of 10

    DSM integrates k-means attention, Mamba state-space layers, diffusion-guided boundary refinement, and CLIP text prompts to segment seen organs and unseen tumors in CT images.

  5. Computationally Efficient Diffusion Models in Medical Imaging: A Comprehensive Review

    eess.IV 2025-05 reject novelty 3.0 of 10

    A survey of DDPM, LDM, and WDM diffusion models for medical imaging, organized around training and inference efficiency.

Pith tools