Pith. sign in

REVIEW 1 cited by

DualDiff: Dual-branch Diffusion Model for Autonomous Driving with Semantic Fusion

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2505.01857 v1 pith:XRYDTC3B submitted 2025-05-03 cs.CV

classification cs.CV
keywords scenedrivingdualdiffinformationbackgroundcontroldiffusiondual-branch
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Accurate and high-fidelity driving scene reconstruction relies on fully leveraging scene information as conditioning. However, existing approaches, which primarily use 3D bounding boxes and binary maps for foreground and background control, fall short in capturing the complexity of the scene and integrating multi-modal information. In this paper, we propose DualDiff, a dual-branch conditional diffusion model designed to enhance multi-view driving scene generation. We introduce Occupancy Ray Sampling (ORS), a semantic-rich 3D representation, alongside numerical driving scene representation, for comprehensive foreground and background control. To improve cross-modal information integration, we propose a Semantic Fusion Attention (SFA) mechanism that aligns and fuses features across modalities. Furthermore, we design a foreground-aware masked (FGM) loss to enhance the generation of tiny objects. DualDiff achieves state-of-the-art performance in FID score, as well as consistently better results in downstream BEV segmentation and 3D object detection tasks.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MultiEditor: Controllable Multimodal Object Editing for Driving Scenarios Using 3D Gaussian Splatting Priors

    cs.AI 2025-07 conditional novelty 6.0 of 10

    A dual-branch diffusion framework jointly edits images and LiDAR point clouds in driving scenes using 3D Gaussian Splatting object priors, improving fidelity and boosting detection of rare vehicle classes.

Pith tools