Pith. sign in

REVIEW 7 cited by

InfiniCube: Unbounded and Controllable Dynamic 3D Driving Scene Generation with World-Guided Video Models

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2412.03934 v2 pith:TFJV462T submitted 2024-12-05 cs.CV cs.AIcs.GR

classification cs.CVcs.AIcs.GR
keywords dynamicgenerationcontrollabledrivingmodelsceneunboundedvideo
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We present InfiniCube, a scalable method for generating unbounded dynamic 3D driving scenes with high fidelity and controllability. Previous methods for scene generation either suffer from limited scales or lack geometric and appearance consistency along generated sequences. In contrast, we leverage the recent advancements in scalable 3D representation and video models to achieve large dynamic scene generation that allows flexible controls through HD maps, vehicle bounding boxes, and text descriptions. First, we construct a map-conditioned sparse-voxel-based 3D generative model to unleash its power for unbounded voxel world generation. Then, we re-purpose a video model and ground it on the voxel world through a set of carefully designed pixel-aligned guidance buffers, synthesizing a consistent appearance. Finally, we propose a fast feed-forward approach that employs both voxel and pixel branches to lift the dynamic videos to dynamic 3D Gaussians with controllable objects. Our method can generate controllable and realistic 3D driving scenes, and extensive experiments validate the effectiveness and superiority of our model.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 7 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Streaming Multi-Agent Autoregressive Diffusion Model with World State Registers

    cs.CV 2026-07 conditional novelty 6.0 of 10

    Adding persistently updated, supervised world-state register tokens to streaming multi-agent diffusion improves cross-agent consistency and visual quality in two-agent Minecraft generation.

  2. Sat2City: 3D City Generation from A Single Satellite Image with Cascaded Latent Diffusion

    cs.CV 2025-07 conditional novelty 6.0 of 10

    Sat2City generates explicit 3D city geometry and appearance from a height-map condition using cascaded latent diffusion on sparse voxel grids, beating prior methods on a new synthetic city dataset.

  3. Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    Post-trained Cosmos world models generate controllable multi-view driving videos and LiDAR; augmenting real AV training data with these synthetic clips improves downstream perception and policy metrics, especially in ...

  4. Dreamland: Controllable World Creation with Simulator and Generative Models

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A three-stage hybrid pipeline uses an intermediate layered world representation to refine simulator-rendered driving scenes into realistic, controllable images and videos.

  5. RGE-GS: Reward-Guided Expansive Driving Scene Reconstruction via Diffusion Priors

    cs.CV 2025-06 conditional novelty 5.0 of 10

    RGE-GS uses a learned pixel-wise reward filter and convergence-aware Gaussian training to improve cross-lane novel view synthesis for driving scenes.

  6. Observable Performance Does Not Fully Reflect Adaptive System Organization: A Multi-Level Analysis of Gait Dynamics Under Occlusal Constraint

    cs.LG 2026-05 unverdicted novelty 4.0 of 10

    In one Parkinson's patient, six occlusal probes produce overlapping gait scores and UMAP embeddings, so observable performance does not uniquely identify adaptive system state under VDO constraint.

  7. DriveGen3D: Boosting Feed-Forward Driving Scene Generation with Efficient Video Diffusion

    cs.CV 2025-10 conditional novelty 4.0 of 10

    DriveGen3D makes long driving-video synthesis and 3D scene reconstruction practical by caching only the conditional diffusion branch, quantizing cross-view attention, and fusing temporal context into a feed-forward Ga...

Pith tools