Pith. sign in

REVIEW 10 cited by

UniScene: Unified Occupancy-centric Driving Scene Generation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.05435 v2 pith:6HWX4PY5 submitted 2024-12-06 cs.CV

classification cs.CV
keywords generationdatasceneuniscenedrivinggeneratingoccupancylidar
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Generating high-fidelity, controllable, and annotated training data is critical for autonomous driving. Existing methods typically generate a single data form directly from a coarse scene layout, which not only fails to output rich data forms required for diverse downstream tasks but also struggles to model the direct layout-to-data distribution. In this paper, we introduce UniScene, the first unified framework for generating three key data forms - semantic occupancy, video, and LiDAR - in driving scenes. UniScene employs a progressive generation process that decomposes the complex task of scene generation into two hierarchical steps: (a) first generating semantic occupancy from a customized scene layout as a meta scene representation rich in both semantic and geometric information, and then (b) conditioned on occupancy, generating video and LiDAR data, respectively, with two novel transfer strategies of Gaussian-based Joint Rendering and Prior-guided Sparse Modeling. This occupancy-centric approach reduces the generation burden, especially for intricate scenes, while providing detailed intermediate representations for the subsequent generation stages. Extensive experiments demonstrate that UniScene outperforms previous SOTAs in the occupancy, video, and LiDAR generation, which also indeed benefits downstream driving tasks. Project page: https://arlo0o.github.io/uniscene/

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 10 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. MapDiffusion: Generative Diffusion for Vectorized Online HD Map Construction and Uncertainty Estimation in Autonomous Driving

    cs.CV 2025-07 conditional novelty 6.0 of 10

    MapDiffusion uses a diffusion decoder conditioned on a latent BEV grid to sample multiple vectorized HD maps, reporting a 5.3% relative mAP gain over StreamMapNet on nuScenes and uncertainty that increases in occluded areas.

  2. GS-Occ3D: Scaling Vision-only Occupancy Reconstruction with Gaussian Splatting

    cs.CV 2025-07 conditional novelty 6.0 of 10

    A camera-only Gaussian-surfel pipeline reconstructs full Waymo scenes, converts them to binary occupancy labels, and trains CVT-Occ to generalize on Occ3D-Waymo and Occ3D-nuScenes at a level close to or above LiDAR-la...

  3. COME: Adding Scene-Centric Forecasting Control to Occupancy World Model

    cs.CV 2025-06 conditional novelty 6.0 of 10

    COME adds a scene-centric forecasting branch as a ControlNet-style condition to a diffusion occupancy world model, improving static-scene consistency and beating prior methods on Occ3D-nuScenes while hiding a stronger...

  4. Genesis: Multimodal Driving Scene Generation with Spatio-Temporal and Cross-Modal Consistency

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A joint video and LiDAR generation framework for driving scenes, conditioned on shared scene layouts, VLM captions, and BEV features, achieves SOTA generation and downstream perception gains on nuScenes.

  5. Controllable Radar Simulation with Waveform Parameter Embedding

    eess.SP 2025-06 reject novelty 6.0 of 10

    A hybrid radar-cube simulator conditions a 3D U-Net on four fitted waveform parameters and is reported to match or beat real radar on 2D detection and segmentation, though its headline experiments do not control for t...

  6. Impromptu VLA: Open Weights and Open Data for Driving Vision-Language-Action Models

    cs.CV 2025-05 conditional novelty 6.0 of 10

    A new 80K-clip dataset of unstructured driving scenarios with Q&A annotations improves VLA performance on NeuroNCAP and nuScenes benchmarks.

  7. Scaling Up Occupancy-centric Driving Scene Generation: Dataset and Method

    cs.CV 2025-10 conditional novelty 5.0 of 10

    UniScenev2 scales occupancy-centric driving-scene generation to NuPlan scale, releasing a 3.6M-frame semantic-occupancy dataset and jointly generating occupancy, video, and LiDAR that beats published baselines on its ...

  8. LongDWM: Cross-Granularity Distillation for Building a Long-Term Driving World Model

    cs.CV 2025-06 reject novelty 5.0 of 10

    A hierarchical coarse-to-fine diffusion transformer with cross-granularity distillation improves long-term driving video prediction, but the reported gains may be inflated by future-derived text prompts and a selected...

  9. Challenger: Affordable Adversarial Driving Video Generation

    cs.CV 2025-05 conditional novelty 5.0 of 10

    A framework for automatic generation of photorealistic adversarial driving videos, shown to sharply increase collision rates of end-to-end autonomous driving models.

  10. DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion

    cs.CV 2025-10 conditional novelty 4.0 of 10

    DriveGen3D makes long driving-video synthesis and 3D scene reconstruction practical by caching only the conditional diffusion branch, quantizing cross-view attention, and fusing temporal context into a feed-forward Ga...

Pith tools