Pith. sign in

REVIEW 3 major objections 6 minor 136 references

LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding

T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read LSD-3D generates explicit, real-time 3D driving scenes with geometry grounded by a proxy mesh and refined by 2D diffusion priors.

desk verdict A serious systems paper with a plausible pipeline and a large unmeasured gap: the 'accurate geometry' claim is never directly tested, and the FVD table alone shouldn't convince you of geometric fidelity. read the letter →

arxiv 2508.19204 v1 pith:IFU7VZIO submitted 2025-08-26 cs.CV cs.AIcs.GR

classification cs.CVcs.AIcs.GR
keywords 3DscenegenerationdrivingsimulationGaussiansplattingscoredistillationdiffusionmodelsgeometrygroundingnovelviewsynthesisautonomous
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

LSD-3D sets out to close the gap between two ways of making driving data: reconstructing real scenes, which are physically grounded but static, and generating videos with diffusion models, which are controllable but not causally consistent. Its route is to first generate a coarse 3D mesh of a street scene, optionally conditioned on a map layout or point cloud, and then refine that proxy into a textured set of Gaussian splats by distilling a fine-tuned 2D latent diffusion model. The intended outcome is an explicit 3D scene that can be rendered in real time from unlimited new trajectories while staying geometrically plausible. That matters because such scenes could supply closed-loop simulators and training data with object permanence and controllability that pure video generation does not offer. The paper reports that the approach matches video models on prompt adherence while beating them on novel-view consistency.

What carries the argument

Geometry-Grounded Distillation Sampling (GGDS) is the mechanism that carries the argument. It is an image-space distillation step: render the Gaussian scene from a viewpoint, encode and noise the image, run a small fixed number of DDIM denoising steps, and backpropagate the pixel and LPIPS loss between the rendered image and the denoised one into the splat parameters. Two features are load-bearing: DDIM inversion replaces random noise sampling so that the objectives stay consistent across optimization steps, and the denoiser is conditioned on disparity maps rendered from the proxy mesh, with normal and disparity regularizers in Eq. (4), so that the diffusion prior refines texture without abandoning the generated layout. This is what lets the method avoid the collapse the ablation reports for vanilla score distillation and random noise sampling on non-overlapping viewpoints.

What would settle it

Render generated scenes under both in-distribution prompts (city street, night) and out-of-distribution prompts (heavy snow, desert), and compare the rendered depth and normal maps against the proxy mesh depth and normals on held-out viewpoints. If geometric error grows sharply for out-of-distribution prompts while texture quality stays high, the geometry-grounding claim fails.

Watch

Extended reading notes

Core claim

The central claim is that score distillation can be made to work for large-scale outdoor driving scenes, not just objects, provided the 3D optimization is anchored to a coarse geometric proxy. The method generates a voxel occupancy grid from a hierarchical latent voxel diffusion model, turns it into a surface mesh, seeds millions of planar Gaussian splats on that mesh, and optimizes them with GGDS. GGDS renders each viewpoint, encodes the image, denoises with a fixed number of steps, and uses DDIM inversion instead of random noise so that consecutive optimization steps agree; the denoised result is compared with the rendered image in pixel and perceptual space. To keep the splats from drifting off the proxy, the diffusion process is conditioned on disparity maps rendered from the mesh and the splats are regularized against the mesh normals and disparity. The paper's claim is that this yields an explicit, causally generated 3D scene with high-fidelity texture, and its Waymo experiments show lower FID and FVD on novel views than video-generation baselines fitted to Gaussian splats.

Load-bearing premise

The load-bearing assumption is that conditioning a fine-tuned 2D diffusion model on disparity maps rendered from the proxy mesh, together with normal and disparity regularizers, is enough to make the distillation converge to a scene that respects the proxy geometry rather than overriding it.

Editorial extensions

If this is right

  • Generated scenes are explicit Gaussian splats, so they can be rendered at over 60 fps at 960p, enabling real-time closed-loop simulation along arbitrary trajectories.
  • Because the geometry is explicit and causal, a generated environment remains 3D-consistent when the viewpoint leaves the recorded path, unlike video diffusion baselines whose quality degrades off-trajectory.
  • Text prompts and map layouts steer scene content, so the same pipeline can produce many versions of one road network under different weather, season, and lighting.
  • The scene representation is composable: dynamic actor assets can be placed, relit by the environment map, and rendered through the ego sensor stack, which is what a closed-loop driving simulator needs.
  • Prompt adherence stays at the level of video-based generation while the representation adds causality and explicit geometry.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • A testable extension the paper leaves implicit: because geometry and appearance are supplied by separate modules, one could swap the voxel-diffusion proxy for a hand-authored layout or an HD map and expect the same texture-refinement step to work, provided the disparity conditioning stays well aligned.
  • The main risk the authors identify is out-of-distribution prompts: if the fine-tuned diffusion prior does not respect disparity conditioning for rare conditions such as heavy snow, geometry drift would appear exactly where synthetic training data is most needed.
  • Since the 2D prior is fine-tuned on driving data, scene diversity is bounded by that prior; a natural follow-up would be to blend multiple priors or add a second prior trained on rare weather and terrain.
  • The same geometry-grounded distillation recipe may transfer to indoor or off-road environments if a proxy mesh can be generated, because the geometry losses do not use driving-specific semantics.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. The paper presents LSD-3D, a pipeline for generating large-scale 3D driving scenes by combining a generated proxy mesh with image-space distillation from a fine-tuned latent diffusion model. The scene is represented as 2D Gaussian splats initialized from the proxy mesh, an environment map, and a text/map-conditioned background; the GGDS procedure refines the splats using DDIM inversion, disparity conditioning, and geometry-grounding regularizers. The central claim is that the method produces explicit, causal, real-time renderable 3D scenes with accurate geometry and object permanence, supporting arbitrary novel trajectories. The paper validates the design with ablations and compares against WonderWorld, Vista, MagicDrive3D, and GEN3C using FID, FDDINOv2, FVD, and CLIP scores.

Significance. If the geometry claim holds, this is a meaningful step beyond both reconstruction-only simulators and trajectory-bound video diffusion models: the method offers explicit 3D scenes, real-time rendering, map-conditioned control, and composability with dynamic actors. The paper deserves credit for a clear system design, for ablations showing that vanilla SDS fails and that DDIM-inversion consistency is needed, and for reporting a large improvement in novel-view FVD over video-based baselines. However, the evidence for the central claim of accurate geometry is currently indirect: all quantitative metrics operate on rendered images or videos, and the proxy mesh generator is never evaluated on its own. Because the headline contributions are geometry grounding, causality, and 3D consistency, the evaluation needs direct geometric measurements before the claim can be accepted.

major comments (3)
  1. [§4.2, Table 2] The paper's central claim—accurate geometry with object permanence and causal novel view synthesis—is not directly measured. All metrics in Table 2 are image-space or video-space appearance metrics (FID, FDDINOv2, FVD, CLIP), and the ablation in Fig. 3 is quantified only with FID. These metrics can improve even when the underlying 3D geometry is wrong, because a splat cloud can render plausible disparity and normal maps while encoding incorrect absolute geometry (e.g., a flat billboard facing the camera). The geometry losses in Eq. 4 are necessary but not sufficient for the claimed accuracy. I request a direct geometry evaluation: (i) compare the generated voxel occupancy or proxy mesh against held-out Waymo LiDAR scans using IoU, Chamfer distance, or F1; (ii) render depth and normal maps from the final Gaussian scene along novel trajectories and compare them against LiDAR ground truth or against the proxy mesh, quantifying the deviation induced by distillation; and (iii) report a metric that directly tests object permanence, such as the consistency of detected object positions across viewpoints. Without such measurements, the 'accurate geometry' claim remains unsupported.
  2. [§3.2] The proxy mesh generator is a load-bearing component but is never validated on its own. The hierarchical voxel diffusion model is trained from scratch on aggregated point clouds and maps, and NKSR produces the coarse mesh that conditions all subsequent distillation; yet the paper reports no evaluation of the generated occupancy or mesh quality, and no ablation isolates the effect of proxy quality on the final scene. If the proxy contains systematic errors—collapsed facades, missing road surfaces, misplaced buildings—every downstream scene inherits those errors. I request a quantitative evaluation of the proxy generator (e.g., occupancy IoU and surface Chamfer distance against held-out LiDAR scans) and an experiment in which the proxy is deliberately degraded or replaced to show that GGDS cannot repair a fundamentally wrong proxy.
  3. [§4.2, footnote 1] The comparison to MagicDrive3D relies on the authors' own reimplementation because the official code and models are unavailable. The footnote states this, but the main text presents the reimplemented baseline on equal footing with the other methods. This is a comparability risk: the reimplementation uses a different backbone (MagicDriveDiT) and the 2DGS optimization, and small implementation differences could explain the observed FID/FVD gaps. I ask the authors to release the reimplementation code and full hyperparameters for exact reproduction, or to validate it against any official results that become available. In addition, Table 2 reports no error bars or significance tests for any method; with only 40 generation scenes and stochastic pipelines, the reported differences may be within run-to-run noise. Please report means and standard errors over at least three seeds per method.
minor comments (6)
  1. [§3.3, Eq. (2)] The notation for DDIM inversion is confusing: the equation writes 'zt,i = DDIM−1(zt−1,i, αt, αt−1)' but the expression resembles the forward DDIM update from zt−1 to zt. Please clarify the indexing and define the exact mapping used in the algorithm, and distinguish it from the standard DDIM inversion notation.
  2. [§3.3, Eq. (1)] The loss in Eq. (1) is written with an incompletely defined norm for the first term (no subscript), and the text says 'the noisy latent zt is the denoised for N steps'. Please define all norms and correct the sentence for clarity.
  3. [§3.1] The reference to 'Huang et al. [2024]' for 2D oriented planar splats should be expanded with the full author list and venue, or else use the already-cited 2DGS reference [36] consistently.
  4. [§4.2, Evaluation Metrics] The FVD reference is described as 'a subset of the respective training dataset [82, 6]', but the subset size and selection procedure are unspecified. Since all methods are compared against the same reference, please state the exact reference distribution and confirm that all methods use the same subset.
  5. [Table 1] Table 1 uses checkmarks and parenthesized checkmarks without a legend. Please add a footnote defining (✓) and the difference between ✓ and (✓).
  6. [§3.3, Eq. (4)] The relative weights of Lnorm, Ldisp, the TV loss, and the 2DGS regularization are deferred to the supplement. Since these weights control the strength of geometry grounding, please provide at least the final values or schedule in the main text or an appendix table.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the generation pipeline is self-contained and the central claims rest on external baselines and independently trained components.

full rationale

I walked the claimed derivation chain: a proxy mesh is produced by a from-scratch hierarchical voxel diffusion model conditioned on map layouts and trained on Waymo point clouds, followed by NKSR surface reconstruction; Gaussians are initialized on that mesh; GGDS then distills a fine-tuned 2D latent diffusion model under disparity conditioning from the rendered proxy depth, with geometry regularization in Eq. 4 pulling rendered normals and disparity back to the proxy. No step defines a predicted quantity in terms of a fitted quantity, and no fitted parameter is renamed as a prediction. The geometry losses enforce consistency with the proxy rather than deriving the proxy from the final output, so the final scene geometry is not equivalent to its input by construction. Quantitative evaluation is against external baselines using shared fine-tuned T2I and 2DGS pipelines, with FID/FVD/FDDINOv2 references drawn from Waymo; this is standard generative evaluation, not circular. The only self-citation is [62] in a list of neural reconstruction works, and it is not load-bearing for any of the paper's contributions. The paper's 'accurate geometry' claim is under-validated because no direct LiDAR/mesh IoU or Chamfer evaluation is reported, but under-validation is a correctness risk, not circularity.

Assumptions & free parameters 6 free parameters · 5 assumptions · 0 invented entities

The central claim rests on two learned priors, the map-conditioned voxel diffusion model on Waymo and a fine-tuned SDXL, plus several hand-set hyperparameters. Geometry grounding is protected by disparity conditioning and losses whose sufficiency is only shown by ablations on 40 scenes. No new physical or conceptual entities are introduced; proxy meshes, environment maps, and 2D Gaussian splats are established representations.

free parameters (6)
  • Denoising steps N = 5
    Set to 5 independently of noise level; the paper states this improves quality at low t but offers no systematic sweep.
  • Minimum noise level tmin annealing schedule = not specified
    Used to anneal generation strength coarse-to-fine; exact schedule is deferred to the supplementary.
  • SGLD perturbation scale lambda_noise = not specified
    Added noise in Eq. 3 to stabilize Langevin dynamics; the value is not reported.
  • Weights for Lnorm, Ldisp, TV, and 2DGS regularization = not specified
    Loss composition in Sec. 3.3 is deferred to the supplementary; these weights determine the geometry versus texture trade-off.
  • Chunk size and overlap for map-conditioned outpainting = 100m x 100m chunks with overlapping zones
    Defines the scale of generated geometry and the potential for accumulated drift across chunks.
  • Gaussian count and pruning bounds = 1.8 to 2.2 million initialized, maximum 4 million
    Scene capacity and rendering quality are sensitive to these caps; they are set by compute budget rather than by analysis.
assumptions (5)
  • domain assumption The map-conditioned hierarchical voxel diffusion model p(V|M), trained from scratch on Waymo point clouds, generates a proxy geometry mesh that is a sufficient scaffold for the scene.
    Sec. 3.2. There is no direct evaluation of mesh accuracy, so all geometry claims inherit this assumption.
  • domain assumption The fine-tuned 2D latent diffusion model used in GGDS provides a strong enough image prior for outdoor driving scenes, including prompts for weather, season, and time of day.
    Sec. 3.3. The method relies on this prior for all appearance and for many structural details.
  • ad hoc to paper Disparity conditioning and the geometry losses in Eq. 4 are sufficient to prevent the optimized Gaussians from drifting away from the proxy mesh.
    Sec. 3.3. This is the mechanism that grounds geometry; the paper validates it only with ablations on a small set of scenes.
  • domain assumption DDIM inversion with fixed N steps provides a consistent optimization target across different viewpoints and noise levels.
    Sec. 3.3, Eq. 2. Optimization stability is central to GGDS, but there is no theoretical guarantee; it is an empirical design choice.
  • domain assumption The Waymo Open Dataset is representative enough to train the geometry and appearance priors for diverse novel scenes.
    Sec. 4.2. If the training distribution is too narrow, the diversity claims, including desert, snow, and night scenes, outpace what the priors can support.

how reviews work

0 comments
Cite this review

Pith. "Pith review of LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding." pith.science (2026). https://pith.science/paper/IFU7VZIO

@misc{pith2026250819204,
  author       = {Pith},
  title        = {Pith review of: LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/IFU7VZIO}},
  note         = {Machine review of arXiv:2508.19204}
}
read the original abstract

Large-scale scene data is essential for training and testing in robot learning. Neural reconstruction methods have promised the capability of reconstructing large physically-grounded outdoor scenes from captured sensor data. However, these methods have baked-in static environments and only allow for limited scene control -- they are functionally constrained in scene and trajectory diversity by the captures from which they are reconstructed. In contrast, generating driving data with recent image or video diffusion models offers control, however, at the cost of geometry grounding and causality. In this work, we aim to bridge this gap and present a method that directly generates large-scale 3D driving scenes with accurate geometry, allowing for causal novel view synthesis with object permanence and explicit 3D geometry estimation. The proposed method combines the generation of a proxy geometry and environment representation with score distillation from learned 2D image priors. We find that this approach allows for high controllability, enabling the prompt-guided geometry and high-fidelity texture and structure that can be conditioned on map layouts -- producing realistic and geometrically consistent 3D generations of complex driving scenes.

Figures

Figures reproduced from arXiv: 2508.19204 by the authors.

Figure 1
Figure 1. Geometry-Grounded Large-Scale 3D Scene Generation. We generate a large-scale scene as a combination of a coarse geometric layout, an environment map, and a set of Gaussians for texture details, discussed in Sec. 3.1. The geometric layout is either generated, conditioned on a map, or predicted from point-cloud data and guides the overall scene structure. We can further control the setting with a scene prompt, describ… view at source ↗
Figure 2
Figure 2. Geometry-Grounded LSD -3D Generations. We visualize 3D scenes generated via our method, alongside the corresponding map of surface normals and selection of novel viewpoints at street level for each of them. In the first two columns, we provide samples of scenes with diversity in time-of-day, season, location, and scene type. In the third column, we provide examples of generated scenes with a point cloud condition an… view at source ↗
Figure 3
Figure 3. Ablation Experiments. We qualitatively validate the core components of our optimization method. With vanilla SDS, scenes completely fail to converge, necessitating our Gaussian Optimization approach. The proposed texture regularization and initialization approach ensures that scenes converge to a reasonable color distribution, while scenes without them fail. The bottom table reports FID scores without the same compo… view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Qualitative Comparisons to Video and Scene Generation Methods. Our approach generates an accurate and 3D￾consistent scene representation, enabling high-quality novel view synthesis and the generation of unlimited off-trajectory view￾points. In contrast, existing baseli…
Figure 5
Figure 5. Figure 5: Composability with Dynamic Actors. We sim￾ulate driving trajectories in a residential street scenes for a Waymo-representative [82] sensor stack. From bottome to top, we show a third-person view of the ego capture vehicle followed by a set of rendered front cameras. Bo…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

136 extracted references · 34 canonical work pages

  1. [1]

    J.; and Guerrero, P

    Anciukevicius, T.; Xu, Z.; Fisher, M.; Henderson, P.; Bilen, H.; Mitra, N. J.; and Guerrero, P. 2022. RenderDiffu- sion: Image Diffusion for 3D Reconstruction, Inpainting and Generation. arXiv

  2. [2]

    J.; Tagliasac- chi, A.; and Lindell, D

    Bahmani, S.; Skorokhodov, I.; Rong, V .; Wetzstein, G.; Guibas, L.; Wonka, P.; Tulyakov, S.; Park, J. J.; Tagliasac- chi, A.; and Lindell, D. B. 2023. 4d-fy: Text-to-4d genera- tion using hybrid score distillation sampling. arXiv preprint arXiv:2311.17984

  3. [3]

    Blattmann, A.; Dockhorn, T.; Kulal, S.; Mendelevitch, D.; Kilian, M.; Lorenz, D.; Levi, Y .; English, Z.; V oleti, V .; Letts, A.; et al. 2023. Stable video diffusion: Scaling la- tent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127

  4. [4]

    Borkman, S.; Crespi, A.; Dhakad, S.; Ganguly, S.; Hogins, J.; Jhang, Y .-C.; Kamalzadeh, M.; Li, B.; Leal, S.; Parisi, P.; et al. 2021. Unity perception: Generate synthetic data for computer vision. arXiv preprint arXiv:2107.04259

  5. [5]

    Brock, A.; Donahue, J.; and Simonyan, K. 2018. Large scale GAN training for high fidelity natural image synthe- sis. arXiv preprint arXiv:1809.11096

  6. [6]

    H.; V ora, S.; Liong, V

    Caesar, H.; Bankiti, V .; Lang, A. H.; V ora, S.; Liong, V . E.; Xu, Q.; Krishnan, A.; Pan, Y .; Baldan, G.; and Beijbom, O

  7. [7]

    Chen, A.; Zheng, W.; Wang, Y .; Zhang, X.; Zhan, K.; Jia, P.; Keutzer, K.; and Zhang, S. 2025. GeoDrive: 3D Geometry- Informed Driving World Model with Precise Action Con- trol. arXiv:2505.22421

  8. [8]

    M.; Ivanovic, B.; Litany, O.; Gojcic, Z.; Fidler, S.; Pavone, M.; Song, L.; and Wang, Y

    Chen, Z.; Yang, J.; Huang, J.; de Lutio, R.; Esturo, J. M.; Ivanovic, B.; Litany, O.; Gojcic, Z.; Fidler, S.; Pavone, M.; Song, L.; and Wang, Y . 2025. OmniRe: Omni Urban Scene Reconstruction. In The Thirteenth International Conference on Learning Representations

Show all 136 references
  1. [9]

    F.; Dideriksen, T.; Arora, H.; Guillaumin, M.; and Malik, J

    Collins, J.; Goel, S.; Deng, K.; Luthra, A.; Xu, L.; Gun- dogdu, E.; Zhang, X.; Yago Vicente, T. F.; Dideriksen, T.; Arora, H.; Guillaumin, M.; and Malik, J. 2022. ABO: Dataset and Benchmarks for Real-World 3D Object Under- standing. CVPR

  2. [10]

    Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; and Schiele, B

  3. [11]

    Dauner, D.; Hallgarten, M.; Li, T.; Weng, X.; Huang, Z.; Yang, Z.; Li, H.; Gilitschenski, I.; Ivanovic, B.; Pavone, M.; Geiger, A.; and Chitta, K. 2024. NA VSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Bench- marking. In Advances in Neural Information Proces...

  4. [12]

    Y .; VanderBilt, E.; Kembhavi, A.; V ondrick, C.; Gkioxari, G.; Ehsani, K.; Schmidt, L.; and Farhadi, A

    Deitke, M.; Liu, R.; Wallingford, M.; Ngo, H.; Michel, O.; Kusupati, A.; Fan, A.; Laforte, C.; V oleti, V .; Gadre, S. Y .; VanderBilt, E.; Kembhavi, A.; V ondrick, C.; Gkioxari, G.; Ehsani, K.; Schmidt, L.; and Farhadi, A. 2023. Objaverse- XL: A Universe of 10M+ 3D Objects. a...

  5. [13]

    Deng, B.; Tucker, R.; Li, Z.; Guibas, L.; Snavely, N.; and Wetzstein, G. 2024. Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffu- sion. In SIGGRAPH 2024 Conference Papers

  6. [14]

    Dhariwal, P.; and Nichol, A. 2021. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34: 8780–8794

  7. [15]

    Dockhorn, T.; Vahdat, A.; and Kreis, K. 2021. Score-based generative modeling with critically-damped langevin diffu- sion. arXiv preprint arXiv:2112.07068

  8. [16]

    Dosovitskiy, A.; Ros, G.; Codevilla, F.; Lopez, A.; and Koltun, V . 2017. CARLA: An open urban driving simulator. In Conference on robot learning, 1–16. PMLR

  9. [17]

    Esser, P.; Rombach, R.; and Ommer, B. 2021. Taming trans- formers for high-resolution image synthesis. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 12873–12883

  10. [18]

    R.; Zhou, Y .; Yang, Z.; Chouard, A.; Sun, P.; Ngiam, J.; Vasudevan, V .; Mc- Cauley, A.; Shlens, J.; and Anguelov, D

    Ettinger, S.; Cheng, S.; Caine, B.; Liu, C.; Zhao, H.; Prad- han, S.; Chai, Y .; Sapp, B.; Qi, C. R.; Zhou, Y .; Yang, Z.; Chouard, A.; Sun, P.; Ngiam, J.; Vasudevan, V .; Mc- Cauley, A.; Shlens, J.; and Anguelov, D. 2021. Large Scale Interactive Motion Forecasting for Autonom...

  11. [19]

    Feng, L.; Li, Q.; Peng, Z.; Tan, S.; and Zhou, B. 2023. Traf- ficgen: Learning to generate diverse and realistic traffic sce- narios. In 2023 IEEE International Conference on Robotics and Automation (ICRA), 3567–3575. IEEE

  12. [20]

    R.; Yang, Y .-H.; Keetha, N

    Fischer, T.; Bul `o, S. R.; Yang, Y .-H.; Keetha, N. V .; Porzi, L.; M ¨uller, N.; Schwarz, K.; Luiten, J.; Pollefeys, M.; and Kontschieder, P. 2025. FlowR: Flowing from Sparse to Dense 3D Reconstructions. arXiv preprint arXiv:2504.01647

  13. [21]

    Gao, R.; Chen, K.; Li, Z.; Hong, L.; Li, Z.; and Xu, Q

  14. [22]

    Gao, R.; Chen, K.; Xiao, B.; Hong, L.; Li, Z.; and Xu, Q. 2024. MagicDriveDiT: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control. arXiv:2411.13807

  15. [23]

    Gao, R.; Chen, K.; Xie, E.; Hong, L.; Li, Z.; Yeung, D.- Y .; and Xu, Q. 2023. Magicdrive: Street view genera- tion with diverse 3d geometry control. arXiv preprint arXiv:2310.02601

  16. [24]

    P.; Barron, J

    Gao*, R.; Holynski*, A.; Henzler, P.; Brussee, A.; Martin- Brualla, R.; Srinivasan, P. P.; Barron, J. T.; and Poole*, B

  17. [25]

    Gao, S.; Yang, J.; Chen, L.; Chitta, K.; Qiu, Y .; Geiger, A.; Zhang, J.; and Li, H. 2024. Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllabil- ity. In Advances in Neural Information Processing Systems (NeurIPS)

  18. [26]

    Geiger, A.; Lenz, P.; and Urtasun, R. 2012. Are we ready for autonomous driving? the kitti vision benchmark suite. In 2012 IEEE conference on computer vision and pattern recognition, 3354–3361. IEEE

  19. [27]

    Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y

  20. [28]

    Advances in Neural Information Processing Systems

    CAT3D: Create Anything in 3D with Multi-View Dif- fusion Models. Advances in Neural Information Processing Systems

  21. [29]

    D.; Agarwal, R.; Roelofs, R.; Lu, Y .; Montali, N.; Mougin, P.; Yang, Z.; White, B.; Faust, A.; McAllister, R.; Anguelov, D.; and Sapp, B

    Gulino, C.; Fu, J.; Luo, W.; Tucker, G.; Bronstein, E.; Lu, Y .; Harb, J.; Pan, X.; Wang, Y .; Chen, X.; Co-Reyes, J. D.; Agarwal, R.; Roelofs, R.; Lu, Y .; Montali, N.; Mougin, P.; Yang, Z.; White, B.; Faust, A.; McAllister, R.; Anguelov, D.; and Sapp, B. 2023. Waymax: An Acc...

  22. [30]

    Harvey, W.; Naderiparizi, S.; Masrani, V .; Weilbach, C.; and Wood, F. 2022. Flexible diffusion modeling of long videos. Advances in Neural Information Processing Systems , 35: 27953–27965

  23. [31]

    Hess, G.; Lindstr ¨om, C.; Fatemi, M.; Petersson, C.; and Svensson, L. 2025. Splatad: Real-time lidar and camera ren- dering with 3d gaussian splatting for autonomous driving. In Proceedings of the Computer Vision and Pattern Recog- nition Conference, 11982–11992

  24. [32]

    P.; Poole, B.; Norouzi, M.; Fleet, D

    Ho, J.; Chan, W.; Saharia, C.; Whang, J.; Gao, R.; Gritsenko, A.; Kingma, D. P.; Poole, B.; Norouzi, M.; Fleet, D. J.; et al

  25. [33]

    Greene, N. 1986. Environment mapping and other appli- cations of world projections. IEEE computer graphics and Applications, 6(11): 21–29

  26. [34]

    Ho, J.; Salimans, T.; Gritsenko, A.; Chan, W.; Norouzi, M.; and Fleet, D. J. 2022. Video diffusion models. Advances in Neural Information Processing Systems, 35: 8633–8646

  27. [35]

    Hong, W.; Ding, M.; Zheng, W.; Liu, X.; and Tang, J. 2022. CogVideo: Large-scale Pretraining for Text-to-Video Gen- eration via Transformers. arXiv preprint arXiv:2205.15868

  28. [36]

    Huang, B.; Yu, Z.; Chen, A.; Geiger, A.; and Gao, S. 2024. 2D Gaussian Splatting for Geometrically Accurate Radiance Fields. In SIGGRAPH 2024 Conference Papers. Association for Computing Machinery

  29. [37]

    Huang, J.; Gojcic, Z.; Atzmon, M.; Litany, O.; Fidler, S.; and Williams, F. 2023. Neural Kernel Surface Reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4369–4379

  30. [38]

    Hwang, S.; Kim, M.-J.; Kang, T.; Kang, J.; and Choo, J

  31. [39]

    Ho, J.; Jain, A.; and Abbeel, P. 2020. Denoising diffusion probabilistic models. Advances in neural information pro- cessing systems, 33: 6840–6851

  32. [40]

    Karras, T.; Laine, S.; Aittala, M.; Hellsten, J.; Lehtinen, J.; and Aila, T. 2020. Analyzing and improving the image qual- ity of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 8110–8119

  33. [41]

    Kazemkhani, S.; Pandya, A.; Cornelisse, D.; Shacklett, B.; and Vinitsky, E. 2025. GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS. In Proceedings of the International Conference on Learning Representations (ICLR)

  34. [42]

    Kerbl, B.; Kopanas, G.; Leimk ¨uhler, T.; and Drettakis, G

  35. [43]

    Kheradmand, S.; Rebain, D.; Sharma, G.; Sun, W.; Tseng, J.; Isack, H.; Kar, A.; Tagliasacchi, A.; and Yi, K. M

  36. [44]

    W.; Brown, B.; Yin, K.; Kreis, K.; Schwarz, K.; Li, D.; Rombach, R.; Torralba, A.; and Fidler, S

    Kim, S. W.; Brown, B.; Yin, K.; Kreis, K.; Schwarz, K.; Li, D.; Rombach, R.; Torralba, A.; and Fidler, S. 2023. NeuralField-LDM: Scene Generation With Hierarchical La- tent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (...

  37. [45]

    In European Conference on Computer Vision, 1–18

    Vegs: View extrapolation of urban scenes in 3d gaus- sian splatting using learned priors. In European Conference on Computer Vision, 1–18. Springer

  38. [46]

    Jun, H.; and Nichol, A. 2023. Shap-E: Generating Condi- tional 3D Implicit Functions. arXiv:2305.02463

  39. [47]

    Li, H.; Shi, H.; Zhang, W.; Wu, W.; Liao, Y .; Wang, L.; Lee, L.-h.; and Zhou, P. 2024. DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sam- pling. arXiv preprint arXiv:2404.03575

  40. [48]

    Li, Q.; Peng, Z.; Feng, L.; Zhang, Q.; Xue, Z.; and Zhou, B

  41. [49]

    Li, Y .; Zou, Z.-X.; Liu, Z.; Wang, D.; Liang, Y .; Yu, Z.; Liu, X.; Guo, Y .-C.; Liang, D.; Ouyang, W.; et al. 2025. Tri- poSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models. arXiv preprint arXiv:2502.06608

  42. [50]

    Liang, Y .; Yang, X.; Lin, J.; Li, H.; Xu, X.; and Chen, Y . 2024. Luciddreamer: Towards high-fidelity text-to-3d generation via interval score matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6517–6526

  43. [51]

    Liao, Y .; Xie, J.; and Geiger, A. 2022. Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(3): 3292–3310

  44. [52]

    arXiv:2404.09591

    3D Gaussian Splatting as Markov Chain Monte Carlo. arXiv:2404.09591

  45. [53]

    Liu, X.; Zhou, C.; and Huang, S. 2024. 3dgs-enhancer: Enhancing unbounded 3d gaussian splatting with view- consistent 2d diffusion priors. Advances in Neural Infor- mation Processing Systems, 37: 133305–133327

  46. [54]

    P.; and Welling, M

    Kingma, D. P.; and Welling, M. 2013. Auto-encoding vari- ational bayes. arXiv preprint arXiv:1312.6114

  47. [55]

    Lee, J.; Lee, S.; Jo, C.; Im, W.; Seon, J.; and Yoon, S.-E

  48. [56]

    In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

    SemCity: Semantic Scene Generation with Triplane Diffusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

  49. [57]

    Luo, S.; Tan, Y .; Huang, L.; Li, J.; and Zhao, H. 2023. Latent consistency models: Synthesizing high-resolution images with few-step inference. arXiv preprint arXiv:2310.04378

  50. [58]

    Miao, X.; Duan, H.; Ojha, V .; Song, J.; Shah, T.; Long, Y .; and Ranjan, R. 2024. Dreamer XL: Towards High- Resolution Text-to-3D Generation via Trajectory Score Matching. arXiv preprint arXiv:2405.11252

  51. [59]

    IEEE Transactions on Pattern Analysis and Machine Intelligence

    Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning. IEEE Transactions on Pattern Analysis and Machine Intelligence

  52. [60]

    NVIDIA; Agarwal, N.; Ali, A.; Bala, M.; Balaji, Y .; Barker, E.; Cai, T.; Chattopadhyay, P.; Chen, Y .; Cui, Y .; Ding, Y .; Dworakowski, D.; Fan, J.; Fenzi, M.; Ferroni, F.; Fidler, S.; Fox, D.; Ge, S.; Ge, Y .; Gu, J.; Gururani, S.; He, E.; Huang, J.; Huffman, J.; Jannaty, P...

  53. [61]

    Oquab, M.; Darcet, T.; Moutakanni, T.; V o, H.; Szafraniec, M.; Khalidov, V .; Fernandez, P.; Haziza, D.; Massa, F.; El- Nouby, A.; et al. 2023. Dinov2: Learning robust visual fea- tures without supervision.arXiv preprint arXiv:2304.07193

  54. [62]

    Ost, J.; Mannan, F.; Thuerey, N.; Knodt, J.; and Heide, F

  55. [63]

    H.; Lee, H.-Y .; Menapace, W.; Chai, M.; Siarohin, A.; Yang, M.-H.; and Tulyakov, S

    Lin, C. H.; Lee, H.-Y .; Menapace, W.; Chai, M.; Siarohin, A.; Yang, M.-H.; and Tulyakov, S. 2023. Infinicity: Infinite- scale city synthesis. In Proceedings of the IEEE/CVF inter- national conference on computer vision, 22808–22818

  56. [64]

    T.; and Mildenhall, B

    Poole, B.; Jain, A.; Barron, J. T.; and Mildenhall, B. 2022. DreamFusion: Text-to-3D using 2D Diffusion. arXiv

  57. [65]

    Ljungbergh, W.; Tonderski, A.; Johnander, J.; Caesar, H.; ˚Astr¨om, K.; Felsberg, M.; and Petersson, C. 2024. Neu- roNCAP: Photorealistic Closed-loop Safety Testing for Au- tonomous Driving. arXiv preprint arXiv:2404.07762

  58. [66]

    Lu, J.; Huang, Z.; Zhang, J.; Yang, Z.; and Zhang, L. 2024. WoV oGen: World V olume-aware Diffusion for Controllable Multi-camera Driving Scene Generation. In European Con- ference on Computer Vision (ECCV)

  59. [67]

    Lu, Y .; Ren, X.; Yang, J.; Shen, T.; Wu, Z.; Gao, J.; Wang, Y .; Chen, S.; Chen, M.; Fidler, S.; and Huang, J. 2024. In- finiCube: Unbounded and Controllable Dynamic 3D Driv- ing Scene Generation with World-Guided Video Models. arXiv:2412.03934

  60. [68]

    Z.; Chen, R.; Kim, S

    Ren, X.; Lu, Y .; Cao, T.; Gao, R.; Huang, S.; Sabour, A.; Shen, T.; Pfaff, T.; Wu, J. Z.; Chen, R.; Kim, S. W.; Gao, J.; Leal-Taixe, L.; Chen, M.; Fidler, S.; and Ling, H

  61. [69]

    Z.; Ling, H.; Chen, M.; Fidler, F., Sanja annd Williams; and Huang, J

    Ren, X.; Lu, Y .; Liang, H.; Wu, J. Z.; Ling, H.; Chen, M.; Fidler, F., Sanja annd Williams; and Huang, J. 2024. SCube: Instant Large-Scale Scene Reconstruction using V oxSplats. In The Thirty-eighth Annual Conference on Neural Informa- tion Processing Systems

  62. [70]

    P.; Tancik, M.; Barron, J

    Mildenhall, B.; Srinivasan, P. P.; Tancik, M.; Barron, J. T.; Ramamoorthi, R.; and Ng, R. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Com- munications of the ACM, 65(1): 99–106

  63. [71]

    Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Om- mer, B. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition , 10684– 10695

  64. [72]

    H.; Tabatabaee, H.; Lu, Q.; Lemke, S.; Moˇzeiko, M.; Boise, E.; Uhm, G.; Gerow, M.; Mehta, S.; et al

    Rong, G.; Shin, B. H.; Tabatabaee, H.; Lu, Q.; Lemke, S.; Moˇzeiko, M.; Boise, E.; Uhm, G.; Gerow, M.; Mehta, S.; et al. 2020. Lgsvl simulator: A high fidelity simulator for autonomous driving. In 2020 IEEE 23rd International con- ference on intelligent transportation systems ...

  65. [73]

    R.; Lagun, D.; Fei-Fei, L.; Sun, D.; et al

    Sargent, K.; Li, Z.; Shah, T.; Herrmann, C.; Yu, H.-X.; Zhang, Y .; Chan, E. R.; Lagun, D.; Fei-Fei, L.; Sun, D.; et al

  66. [74]

    Sauer, A.; Schwarz, K.; and Geiger, A. 2022. Stylegan-xl: Scaling stylegan to large diverse datasets. In ACM SIG- GRAPH 2022 conference proceedings, 1–10

  67. [75]

    Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; M¨uller, J.; Penna, J.; and Rombach, R. 2023. Sdxl: Im- proving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952

  68. [76]

    Shriram, J.; Trevithick, A.; Liu, L.; and Ramamoorthi, R. 2024. Realmdreamer: Text-driven 3d scene genera- tion with inpainting and depth diffusion. arXiv preprint arXiv:2404.07199

  69. [77]

    W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I

    Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021. Learning Transferable Visual Models From Natural Language Supervision. In In- ternational Conference on Machine Learning

  70. [78]

    Razavi, A.; Van den Oord, A.; and Vinyals, O. 2019. Gener- ating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems, 32

  71. [79]

    Ren, X.; Huang, J.; Zeng, X.; Museth, K.; Fidler, S.; and Williams, F. 2024. XCube: Large-Scale 3D Generative Modeling using Sparse V oxel Hierarchies. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

  72. [80]

    Song, Y .; Sun, Z.; and Yin, X. 2024. SDXS: Real-Time One- Step Latent Diffusion Models with Image Conditions.arXiv preprint arXiv:2403.16627

  73. [81]

    L.; Taylor, E.; and Loaiza-Ganem, G

    Stein, G.; Cresswell, J.; Hosseinzadeh, R.; Sui, Y .; Ross, B.; Villecroze, V .; Liu, Z.; Caterini, A. L.; Taylor, E.; and Loaiza-Ganem, G. 2023. Exposing flaws of generative model evaluation metrics and their unfair treatment of dif- fusion models. Advances in Neural Informat...

  74. [82]

    Sun, P.; Kretzschmar, H.; Dotiwalla, X.; Chouard, A.; Pat- naik, V .; Tsui, P.; Guo, J.; Zhou, Y .; Chai, Y .; Caine, B.; Vasudevan, V .; Han, W.; Ngiam, J.; Zhao, H.; Timofeev, A.; Ettinger, S.; Krivokon, M.; Gao, A.; Joshi, A.; Zhang, Y .; Shlens, J.; Chen, Z.; and Anguelov,...

  75. [83]

    Ren, X.; Shen, T.; Huang, J.; Ling, H.; Lu, Y .; Nimier-David, M.; M ¨uller, T.; Keller, A.; Fidler, S.; and Gao, J. 2025. Gen3c: 3d-informed world-consistent video generation with precise camera control. In Proceedings of the Computer Vi- sion and Pattern Recognition Conferen...

  76. [84]

    Team, A.; Zhu, H.; Wang, Y .; Zhou, J.; Chang, W.; Zhou, Y .; Li, Z.; Chen, J.; Shen, C.; Pang, J.; et al. 2025. Aether: Geometric-aware unified world modeling. arXiv preprint arXiv:2503.18945

  77. [85]

    Team, G. 2024. Mochi 1. https://github.com/genmoai/ models

  78. [86]

    Team, T. H. 2025. Hunyuan3D 2.0: Scaling Diffusion Mod- els for High Resolution Textured 3D Assets Generation. arXiv:2501.12202

  79. [87]

    arXiv preprint arXiv:2310.17994

    Zeronvs: Zero-shot 360-degree view synthesis from a single real image. arXiv preprint arXiv:2310.17994

  80. [88]

    Tonderski, A.; Lindstr ¨om, C.; Hess, G.; Ljungbergh, W.; Svensson, L.; and Petersson, C. 2023. NeuRAD: Neu- ral rendering for autonomous driving. arXiv preprint arXiv:2311.15260

  81. [89]

    Shah, S.; Dey, D.; Lovett, C.; and Kapoor, A. 2018. Airsim: High-fidelity visual and physical simulation for autonomous vehicles. In Field and Service Robotics: Results of the 11th International Conference, 621–635. Springer

  82. [90]

    Vahdat, A.; Kreis, K.; and Kautz, J. 2021. Score-based gen- erative modeling in latent space. Advances in neural infor- mation processing systems, 34: 11287–11302

  83. [91]

    R.; Chan, E

    Shue, J. R.; Chan, E. R.; Po, R.; Ankner, Z.; Wu, J.; and Wetzstein, G. 2023. 3d neural field generation using triplane diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20875–20886

  84. [92]

    Singer, U.; Polyak, A.; Hayes, T.; Yin, X.; An, J.; Zhang, S.; Hu, Q.; Yang, H.; Ashual, O.; Gafni, O.; et al. 2022. Make- a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792

  85. [93]

    Song, J.; Meng, C.; and Ermon, S. 2020. Denoising Diffu- sion Implicit Models. arXiv:2010.02502

  86. [94]

    Williams, F.; Gojcic, Z.; Khamis, S.; Zorin, D.; Bruna, J.; Fidler, S.; and Litany, O. 2021. Neural Fields as Learnable Kernels for 3D Reconstruction. arXiv:2111.13674

  87. [95]

    Z.; Zhang, Y .; Turki, H.; Ren, X.; Gao, J.; Shou, M

    Wu, J. Z.; Zhang, Y .; Turki, H.; Ren, X.; Gao, J.; Shou, M. Z.; Fidler, S.; Gojcic, Z.; and Ling, H. 2025. Di- fix3d+: Improving 3d reconstructions with single-step dif- fusion models. In Proceedings of the Computer Vision and Pattern Recognition Conference, 26024–26035

  88. [96]

    T.; and Holynski, A

    Wu, R.; Gao, R.; Poole, B.; Trevithick, A.; Zheng, C.; Barron, J. T.; and Holynski, A. 2024. CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models. arXiv:2411.18613

  89. [97]

    Talwar, D.; Guruswamy, S.; Ravipati, N.; and Eirinaki, M

  90. [98]

    In 2020 IEEE International Conference On Artificial Intelligence Testing (AITest) , 73–

    Evaluating validity of synthetic data in perception tasks for autonomous vehicles. In 2020 IEEE International Conference On Artificial Intelligence Testing (AITest) , 73–

  91. [99]

    Xiang, J.; Lv, Z.; Xu, S.; Deng, Y .; Wang, R.; Zhang, B.; Chen, D.; Tong, X.; and Yang, J. 2024. Structured 3d la- tents for scalable and versatile 3d generation.arXiv preprint arXiv:2412.01506

  92. [100]

    Xie, H.; Chen, Z.; Hong, F.; and Liu, Z. 2024. Citydreamer: Compositional generative model of unbounded 3d cities. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, 9666–9675

  93. [101]

    Xie, K.; Lorraine, J.; Cao, T.; Gao, J.; Lucas, J.; Torralba, A.; Fidler, S.; and Zeng, X. 2024. LATTE3D: Large-scale Amortized Text-To-Enhanced3D Synthesis. arXiv preprint arXiv:2403.15385

  94. [102]

    Thies, J.; Zollh ¨ofer, M.; and Nießner, M. 2019. Deferred neural rendering: Image synthesis using neural textures. Acm Transactions on Graphics (TOG), 38(4): 1–12

  95. [103]

    Yang, J.; Gao, S.; Qiu, Y .; Chen, L.; Li, T.; Dai, B.; Chitta, K.; Wu, P.; Zeng, J.; Luo, P.; Zhang, J.; Geiger, A.; Qiao, Y .; and Li, H. 2024. Generalized Predictive Model for Au- tonomous Driving. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern R...

  96. [104]

    Vahdat, A.; and Kautz, J. 2020. NV AE: A deep hierarchi- cal variational autoencoder. Advances in neural information processing systems, 33: 19667–19679

  97. [105]

    Yang, S.; Hou, L.; Huang, H.; Ma, C.; Wan, P.; Zhang, D.; Chen, X.; and Liao, J. 2024. Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion. arXiv preprint arXiv:2402.03162

  98. [106]

    Wan, T.; Wang, A.; Ai, B.; Wen, B.; Mao, C.; Xie, C.-W.; Chen, D.; Yu, F.; Zhao, H.; Yang, J.; et al. 2025. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314

  99. [107]

    Wang, P.; Xu, D.; Fan, Z.; Wang, D.; Mohan, S.; Iandola, F.; Ranjan, R.; Li, Y .; Liu, Q.; Wang, Z.; et al. 2023. Taming Mode Collapse in Score Distillation for Text-to-3D Genera- tion. arXiv preprint arXiv:2401.00909

  100. [108]

    Wang, X.; Zhu, Z.; Huang, G.; Chen, X.; and Lu, J. 2023. Drivedreamer: Towards real-world-driven world models for autonomous driving. arXiv preprint arXiv:2309.09777

  101. [109]

    Yi, T.; Fang, J.; Wang, J.; Wu, G.; Xie, L.; Zhang, X.; Liu, W.; Tian, Q.; and Wang, X. 2024. GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models. In CVPR

  102. [110]

    T.; and Wu, J

    Yu, H.-X.; Duan, H.; Herrmann, C.; Freeman, W. T.; and Wu, J. 2024. WonderWorld: Interactive 3D Scene Genera- tion from a Single Image. arXiv preprint arXiv:2406.09394

  103. [111]

    T.; Cole, F.; Sun, D.; Snavely, N.; Wu, J.; et al

    Yu, H.-X.; Duan, H.; Hur, J.; Sargent, K.; Rubinstein, M.; Freeman, W. T.; Cole, F.; Sun, D.; Snavely, N.; Wu, J.; et al

  104. [112]

    P.; Verbin, D.; Barron, J

    Wu, R.; Mildenhall, B.; Henzler, P.; Park, K.; Gao, R.; Wat- son, D.; Srinivasan, P. P.; Verbin, D.; Barron, J. T.; Poole, B.; et al. 2024. Reconfusion: 3d reconstruction with diffu- sion priors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognit...

  105. [113]

    Wu, Z.; Liu, T.; Luo, L.; Zhong, Z.; Chen, J.; Xiao, H.; Hou, C.; Lou, H.; Chen, Y .; Yang, R.; et al. 2023. Mars: An instance-aware, modular and realistic simulator for au- tonomous driving. In CAAI International Conference on Ar- tificial Intelligence, 3–15. Springer

  106. [114]

    A.; Shechtman, E.; and Wang, O

    Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In CVPR

  107. [115]

    Zhang, S.; Zhang, Y .; Zheng, Q.; Ma, R.; Hua, W.; Bao, H.; Xu, W.; and Zou, C. 2024. 3D-SceneDreamer: Text- Driven 3D-Consistent Scene Generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10170–10180

  108. [116]

    Zhang, Z.; Long, F.; Pan, Y .; Qiu, Z.; Yao, T.; Cao, Y .; and Mei, T. 2024. TRIP: Temporal Residual Learning with Im- age Noise Prior for Image-to-Video Diffusion Models.arXiv preprint arXiv:2403.17005

  109. [117]

    Xu, Y .; Chai, M.; Shi, Z.; Peng, S.; Skorokhodov, I.; Siaro- hin, A.; Yang, C.; Shen, Y .; Lee, H.-Y .; Zhou, B.; et al

  110. [118]

    In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, 4402–4412

    Discoscene: Spatially disentangled generative radi- ance fields for controllable 3d-aware scene synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, 4402–4412

  111. [119]

    Zyrianov, V .; Che, H.; Liu, Z.; and Wang, S. 2024. Li- darDM: Generative LiDAR Simulation in a Generated World. arXiv preprint arXiv:2404.02903

  112. [120]

    Yang, J.; Huang, J.; Chen, Y .; Wang, Y .; Li, B.; You, Y .; Igl, M.; Sharma, A.; Karkus, P.; Xu, D.; Ivanovic, B.; Wang, Y .; and Pavone, M. 2025. STORM: Spatio-Temporal Re- construction Model for Large-scale Outdoor Scenes. arXiv preprint arXiv:2501.00602

  113. [122]

    Yang, Y .; Yang, Y .; Guo, H.; Xiong, R.; Wang, Y .; and Liao, Y . 2023. Urbangiraffe: Representing urban scenes as com- positional generative neural feature fields. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, 9199–9210

  114. [123]

    J.; and Urtasun, R

    Yang, Z.; Chen, Y .; Wang, J.; Manivasagam, S.; Ma, W.- C.; Yang, A. J.; and Urtasun, R. 2023. UniSim: A Neural Closed-Loop Sensor Simulator. In CVPR

  115. [124]

    Yang, Z.; Teng, J.; Zheng, W.; Ding, M.; Huang, S.; Xu, J.; Yang, Y .; Hong, W.; Zhang, X.; Feng, G.; et al. 2024. CogVideoX: Text-to-Video Diffusion Models with An Ex- pert Transformer. arXiv preprint arXiv:2408.06072

  116. [128]

    In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6658–6667

    Wonderjourney: Going from anywhere to everywhere. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6658–6667

  117. [129]

    R.; Liu, G.; and Zhou, B

    Zhang, J.; Zhang, Q.; Zhang, L.; Kompella, R. R.; Liu, G.; and Zhou, B. 2024. Urban Scene Diffusion through Seman- tic Occupancy Map. arXiv preprint arXiv:2403.11697

  118. [130]

    Zhang, L.; Rao, A.; and Agrawala, M. 2023. Adding condi- tional control to text-to-image diffusion models. InProceed- ings of the IEEE/CVF international conference on computer vision, 3836–3847

  119. [134]

    Zhengwentai, S. 2023. clip-score: CLIP Score for PyTorch. https://github.com/taited/clip-score. Version 0.2.1

  120. [135]

    Zhou, L.; Du, Y .; and Wu, J. 2021. 3D Shape Generation and Completion Through Point-V oxel Diffusion. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), 5826–5835

  121. [2014]

    Advances in neural in- formation processing systems, 27

    Generative adversarial nets. Advances in neural in- formation processing systems, 27

  122. [2016]

    In Proceedings of the IEEE conference on com- puter vision and pattern recognition, 3213–3223

    The cityscapes dataset for semantic urban scene un- derstanding. In Proceedings of the IEEE conference on com- puter vision and pattern recognition, 3213–3223

  123. [2020]

    nuScenes: A multimodal dataset for autonomous driv- ing. In CVPR

  124. [2021]

    In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2856–2865

    Neural scene graphs for dynamic scenes. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2856–2865

  125. [2022]

    arXiv preprint arXiv:2210.02303

    Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303

  126. [2023]

    ACM Transactions on Graphics, 42(4)

    3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on Graphics, 42(4)

  127. [2024]

    arXiv preprint arXiv:2405.14475

    MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes. arXiv preprint arXiv:2405.14475

  128. [2025]

    Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.