REVIEW 3 major objections 6 minor 136 references
LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding
T0 review · 3 major / 6 minor · reviewed 2026-08-15 · deepseek-v4-flash
Pith's one-line read LSD-3D generates explicit, real-time 3D driving scenes with geometry grounded by a proxy mesh and refined by 2D diffusion priors.
desk verdict A serious systems paper with a plausible pipeline and a large unmeasured gap: the 'accurate geometry' claim is never directly tested, and the FVD table alone shouldn't convince you of geometric fidelity. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Geometry-Grounded Distillation Sampling (GGDS) is the mechanism that carries the argument. It is an image-space distillation step: render the Gaussian scene from a viewpoint, encode and noise the image, run a small fixed number of DDIM denoising steps, and backpropagate the pixel and LPIPS loss between the rendered image and the denoised one into the splat parameters. Two features are load-bearing: DDIM inversion replaces random noise sampling so that the objectives stay consistent across optimization steps, and the denoiser is conditioned on disparity maps rendered from the proxy mesh, with normal and disparity regularizers in Eq. (4), so that the diffusion prior refines texture without abandoning the generated layout. This is what lets the method avoid the collapse the ablation reports for vanilla score distillation and random noise sampling on non-overlapping viewpoints.
What would settle it
Render generated scenes under both in-distribution prompts (city street, night) and out-of-distribution prompts (heavy snow, desert), and compare the rendered depth and normal maps against the proxy mesh depth and normals on held-out viewpoints. If geometric error grows sharply for out-of-distribution prompts while texture quality stays high, the geometry-grounding claim fails.
Extended reading notes
Core claim
The central claim is that score distillation can be made to work for large-scale outdoor driving scenes, not just objects, provided the 3D optimization is anchored to a coarse geometric proxy. The method generates a voxel occupancy grid from a hierarchical latent voxel diffusion model, turns it into a surface mesh, seeds millions of planar Gaussian splats on that mesh, and optimizes them with GGDS. GGDS renders each viewpoint, encodes the image, denoises with a fixed number of steps, and uses DDIM inversion instead of random noise so that consecutive optimization steps agree; the denoised result is compared with the rendered image in pixel and perceptual space. To keep the splats from drifting off the proxy, the diffusion process is conditioned on disparity maps rendered from the mesh and the splats are regularized against the mesh normals and disparity. The paper's claim is that this yields an explicit, causally generated 3D scene with high-fidelity texture, and its Waymo experiments show lower FID and FVD on novel views than video-generation baselines fitted to Gaussian splats.
Load-bearing premise
The load-bearing assumption is that conditioning a fine-tuned 2D diffusion model on disparity maps rendered from the proxy mesh, together with normal and disparity regularizers, is enough to make the distillation converge to a scene that respects the proxy geometry rather than overriding it.
Editorial extensions
If this is right
- Generated scenes are explicit Gaussian splats, so they can be rendered at over 60 fps at 960p, enabling real-time closed-loop simulation along arbitrary trajectories.
- Because the geometry is explicit and causal, a generated environment remains 3D-consistent when the viewpoint leaves the recorded path, unlike video diffusion baselines whose quality degrades off-trajectory.
- Text prompts and map layouts steer scene content, so the same pipeline can produce many versions of one road network under different weather, season, and lighting.
- The scene representation is composable: dynamic actor assets can be placed, relit by the environment map, and rendered through the ego sensor stack, which is what a closed-loop driving simulator needs.
- Prompt adherence stays at the level of video-based generation while the representation adds causality and explicit geometry.
Reading between the lines
- A testable extension the paper leaves implicit: because geometry and appearance are supplied by separate modules, one could swap the voxel-diffusion proxy for a hand-authored layout or an HD map and expect the same texture-refinement step to work, provided the disparity conditioning stays well aligned.
- The main risk the authors identify is out-of-distribution prompts: if the fine-tuned diffusion prior does not respect disparity conditioning for rare conditions such as heavy snow, geometry drift would appear exactly where synthetic training data is most needed.
- Since the 2D prior is fine-tuned on driving data, scene diversity is bounded by that prior; a natural follow-up would be to blend multiple priors or add a second prior trained on rare weather and terrain.
- The same geometry-grounded distillation recipe may transfer to indoor or off-road environments if a proxy mesh can be generated, because the geometry losses do not use driving-specific semantics.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper presents LSD-3D, a pipeline for generating large-scale 3D driving scenes by combining a generated proxy mesh with image-space distillation from a fine-tuned latent diffusion model. The scene is represented as 2D Gaussian splats initialized from the proxy mesh, an environment map, and a text/map-conditioned background; the GGDS procedure refines the splats using DDIM inversion, disparity conditioning, and geometry-grounding regularizers. The central claim is that the method produces explicit, causal, real-time renderable 3D scenes with accurate geometry and object permanence, supporting arbitrary novel trajectories. The paper validates the design with ablations and compares against WonderWorld, Vista, MagicDrive3D, and GEN3C using FID, FDDINOv2, FVD, and CLIP scores.
Significance. If the geometry claim holds, this is a meaningful step beyond both reconstruction-only simulators and trajectory-bound video diffusion models: the method offers explicit 3D scenes, real-time rendering, map-conditioned control, and composability with dynamic actors. The paper deserves credit for a clear system design, for ablations showing that vanilla SDS fails and that DDIM-inversion consistency is needed, and for reporting a large improvement in novel-view FVD over video-based baselines. However, the evidence for the central claim of accurate geometry is currently indirect: all quantitative metrics operate on rendered images or videos, and the proxy mesh generator is never evaluated on its own. Because the headline contributions are geometry grounding, causality, and 3D consistency, the evaluation needs direct geometric measurements before the claim can be accepted.
major comments (3)
- [§4.2, Table 2] The paper's central claim—accurate geometry with object permanence and causal novel view synthesis—is not directly measured. All metrics in Table 2 are image-space or video-space appearance metrics (FID, FDDINOv2, FVD, CLIP), and the ablation in Fig. 3 is quantified only with FID. These metrics can improve even when the underlying 3D geometry is wrong, because a splat cloud can render plausible disparity and normal maps while encoding incorrect absolute geometry (e.g., a flat billboard facing the camera). The geometry losses in Eq. 4 are necessary but not sufficient for the claimed accuracy. I request a direct geometry evaluation: (i) compare the generated voxel occupancy or proxy mesh against held-out Waymo LiDAR scans using IoU, Chamfer distance, or F1; (ii) render depth and normal maps from the final Gaussian scene along novel trajectories and compare them against LiDAR ground truth or against the proxy mesh, quantifying the deviation induced by distillation; and (iii) report a metric that directly tests object permanence, such as the consistency of detected object positions across viewpoints. Without such measurements, the 'accurate geometry' claim remains unsupported.
- [§3.2] The proxy mesh generator is a load-bearing component but is never validated on its own. The hierarchical voxel diffusion model is trained from scratch on aggregated point clouds and maps, and NKSR produces the coarse mesh that conditions all subsequent distillation; yet the paper reports no evaluation of the generated occupancy or mesh quality, and no ablation isolates the effect of proxy quality on the final scene. If the proxy contains systematic errors—collapsed facades, missing road surfaces, misplaced buildings—every downstream scene inherits those errors. I request a quantitative evaluation of the proxy generator (e.g., occupancy IoU and surface Chamfer distance against held-out LiDAR scans) and an experiment in which the proxy is deliberately degraded or replaced to show that GGDS cannot repair a fundamentally wrong proxy.
- [§4.2, footnote 1] The comparison to MagicDrive3D relies on the authors' own reimplementation because the official code and models are unavailable. The footnote states this, but the main text presents the reimplemented baseline on equal footing with the other methods. This is a comparability risk: the reimplementation uses a different backbone (MagicDriveDiT) and the 2DGS optimization, and small implementation differences could explain the observed FID/FVD gaps. I ask the authors to release the reimplementation code and full hyperparameters for exact reproduction, or to validate it against any official results that become available. In addition, Table 2 reports no error bars or significance tests for any method; with only 40 generation scenes and stochastic pipelines, the reported differences may be within run-to-run noise. Please report means and standard errors over at least three seeds per method.
minor comments (6)
- [§3.3, Eq. (2)] The notation for DDIM inversion is confusing: the equation writes 'zt,i = DDIM−1(zt−1,i, αt, αt−1)' but the expression resembles the forward DDIM update from zt−1 to zt. Please clarify the indexing and define the exact mapping used in the algorithm, and distinguish it from the standard DDIM inversion notation.
- [§3.3, Eq. (1)] The loss in Eq. (1) is written with an incompletely defined norm for the first term (no subscript), and the text says 'the noisy latent zt is the denoised for N steps'. Please define all norms and correct the sentence for clarity.
- [§3.1] The reference to 'Huang et al. [2024]' for 2D oriented planar splats should be expanded with the full author list and venue, or else use the already-cited 2DGS reference [36] consistently.
- [§4.2, Evaluation Metrics] The FVD reference is described as 'a subset of the respective training dataset [82, 6]', but the subset size and selection procedure are unspecified. Since all methods are compared against the same reference, please state the exact reference distribution and confirm that all methods use the same subset.
- [Table 1] Table 1 uses checkmarks and parenthesized checkmarks without a legend. Please add a footnote defining (✓) and the difference between ✓ and (✓).
- [§3.3, Eq. (4)] The relative weights of Lnorm, Ldisp, the TV loss, and the 2DGS regularization are deferred to the supplement. Since these weights control the strength of geometry grounding, please provide at least the final values or schedule in the main text or an appendix table.
Circularity Check
No significant circularity: the generation pipeline is self-contained and the central claims rest on external baselines and independently trained components.
full rationale
I walked the claimed derivation chain: a proxy mesh is produced by a from-scratch hierarchical voxel diffusion model conditioned on map layouts and trained on Waymo point clouds, followed by NKSR surface reconstruction; Gaussians are initialized on that mesh; GGDS then distills a fine-tuned 2D latent diffusion model under disparity conditioning from the rendered proxy depth, with geometry regularization in Eq. 4 pulling rendered normals and disparity back to the proxy. No step defines a predicted quantity in terms of a fitted quantity, and no fitted parameter is renamed as a prediction. The geometry losses enforce consistency with the proxy rather than deriving the proxy from the final output, so the final scene geometry is not equivalent to its input by construction. Quantitative evaluation is against external baselines using shared fine-tuned T2I and 2DGS pipelines, with FID/FVD/FDDINOv2 references drawn from Waymo; this is standard generative evaluation, not circular. The only self-citation is [62] in a list of neural reconstruction works, and it is not load-bearing for any of the paper's contributions. The paper's 'accurate geometry' claim is under-validated because no direct LiDAR/mesh IoU or Chamfer evaluation is reported, but under-validation is a correctness risk, not circularity.
Assumptions & free parameters
free parameters (6)
- Denoising steps N =
5
- Minimum noise level tmin annealing schedule =
not specified
- SGLD perturbation scale lambda_noise =
not specified
- Weights for Lnorm, Ldisp, TV, and 2DGS regularization =
not specified
- Chunk size and overlap for map-conditioned outpainting =
100m x 100m chunks with overlapping zones
- Gaussian count and pruning bounds =
1.8 to 2.2 million initialized, maximum 4 million
assumptions (5)
- domain assumption The map-conditioned hierarchical voxel diffusion model p(V|M), trained from scratch on Waymo point clouds, generates a proxy geometry mesh that is a sufficient scaffold for the scene.
- domain assumption The fine-tuned 2D latent diffusion model used in GGDS provides a strong enough image prior for outdoor driving scenes, including prompts for weather, season, and time of day.
- ad hoc to paper Disparity conditioning and the geometry losses in Eq. 4 are sufficient to prevent the optimized Gaussians from drifting away from the proxy mesh.
- domain assumption DDIM inversion with fixed N steps provides a consistent optimization target across different viewpoints and noise levels.
- domain assumption The Waymo Open Dataset is representative enough to train the geometry and appearance priors for diverse novel scenes.
Cite this review
Pith. "Pith review of LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding." pith.science (2026). https://pith.science/paper/IFU7VZIO
@misc{pith2026250819204,
author = {Pith},
title = {Pith review of: LSD-3D: Large-Scale 3D Driving Scene Generation with Geometry Grounding},
year = {2026},
howpublished = {\url{https://pith.science/paper/IFU7VZIO}},
note = {Machine review of arXiv:2508.19204}
}
read the original abstract
Large-scale scene data is essential for training and testing in robot learning. Neural reconstruction methods have promised the capability of reconstructing large physically-grounded outdoor scenes from captured sensor data. However, these methods have baked-in static environments and only allow for limited scene control -- they are functionally constrained in scene and trajectory diversity by the captures from which they are reconstructed. In contrast, generating driving data with recent image or video diffusion models offers control, however, at the cost of geometry grounding and causality. In this work, we aim to bridge this gap and present a method that directly generates large-scale 3D driving scenes with accurate geometry, allowing for causal novel view synthesis with object permanence and explicit 3D geometry estimation. The proposed method combines the generation of a proxy geometry and environment representation with score distillation from learned 2D image priors. We find that this approach allows for high controllability, enabling the prompt-guided geometry and high-fidelity texture and structure that can be conditioned on map layouts -- producing realistic and geometrically consistent 3D generations of complex driving scenes.
Figures
Figures from the paper (2 more)
Reference graph
Works this paper leans on
-
[1]
J.; and Guerrero, P
Anciukevicius, T.; Xu, Z.; Fisher, M.; Henderson, P.; Bilen, H.; Mitra, N. J.; and Guerrero, P. 2022. RenderDiffu- sion: Image Diffusion for 3D Reconstruction, Inpainting and Generation. arXiv
2022
-
[2]
J.; Tagliasac- chi, A.; and Lindell, D
Bahmani, S.; Skorokhodov, I.; Rong, V .; Wetzstein, G.; Guibas, L.; Wonka, P.; Tulyakov, S.; Park, J. J.; Tagliasac- chi, A.; and Lindell, D. B. 2023. 4d-fy: Text-to-4d genera- tion using hybrid score distillation sampling. arXiv preprint arXiv:2311.17984
arXiv 2023
-
[3]
Blattmann, A.; Dockhorn, T.; Kulal, S.; Mendelevitch, D.; Kilian, M.; Lorenz, D.; Levi, Y .; English, Z.; V oleti, V .; Letts, A.; et al. 2023. Stable video diffusion: Scaling la- tent video diffusion models to large datasets. arXiv preprint arXiv:2311.15127
arXiv 2023
-
[4]
Borkman, S.; Crespi, A.; Dhakad, S.; Ganguly, S.; Hogins, J.; Jhang, Y .-C.; Kamalzadeh, M.; Li, B.; Leal, S.; Parisi, P.; et al. 2021. Unity perception: Generate synthetic data for computer vision. arXiv preprint arXiv:2107.04259
arXiv 2021
-
[5]
Brock, A.; Donahue, J.; and Simonyan, K. 2018. Large scale GAN training for high fidelity natural image synthe- sis. arXiv preprint arXiv:1809.11096
arXiv 2018
-
[6]
H.; V ora, S.; Liong, V
Caesar, H.; Bankiti, V .; Lang, A. H.; V ora, S.; Liong, V . E.; Xu, Q.; Krishnan, A.; Pan, Y .; Baldan, G.; and Beijbom, O
-
[7]
Chen, A.; Zheng, W.; Wang, Y .; Zhang, X.; Zhan, K.; Jia, P.; Keutzer, K.; and Zhang, S. 2025. GeoDrive: 3D Geometry- Informed Driving World Model with Precise Action Con- trol. arXiv:2505.22421
arXiv 2025
-
[8]
M.; Ivanovic, B.; Litany, O.; Gojcic, Z.; Fidler, S.; Pavone, M.; Song, L.; and Wang, Y
Chen, Z.; Yang, J.; Huang, J.; de Lutio, R.; Esturo, J. M.; Ivanovic, B.; Litany, O.; Gojcic, Z.; Fidler, S.; Pavone, M.; Song, L.; and Wang, Y . 2025. OmniRe: Omni Urban Scene Reconstruction. In The Thirteenth International Conference on Learning Representations
2025
Show all 136 references
-
[9]
F.; Dideriksen, T.; Arora, H.; Guillaumin, M.; and Malik, J
Collins, J.; Goel, S.; Deng, K.; Luthra, A.; Xu, L.; Gun- dogdu, E.; Zhang, X.; Yago Vicente, T. F.; Dideriksen, T.; Arora, H.; Guillaumin, M.; and Malik, J. 2022. ABO: Dataset and Benchmarks for Real-World 3D Object Under- standing. CVPR
2022
-
[10]
Cordts, M.; Omran, M.; Ramos, S.; Rehfeld, T.; Enzweiler, M.; Benenson, R.; Franke, U.; Roth, S.; and Schiele, B
-
[11]
Dauner, D.; Hallgarten, M.; Li, T.; Weng, X.; Huang, Z.; Yang, Z.; Li, H.; Gilitschenski, I.; Ivanovic, B.; Pavone, M.; Geiger, A.; and Chitta, K. 2024. NA VSIM: Data-Driven Non-Reactive Autonomous Vehicle Simulation and Bench- marking. In Advances in Neural Information Proces...
2024
-
[12]
Y .; VanderBilt, E.; Kembhavi, A.; V ondrick, C.; Gkioxari, G.; Ehsani, K.; Schmidt, L.; and Farhadi, A
Deitke, M.; Liu, R.; Wallingford, M.; Ngo, H.; Michel, O.; Kusupati, A.; Fan, A.; Laforte, C.; V oleti, V .; Gadre, S. Y .; VanderBilt, E.; Kembhavi, A.; V ondrick, C.; Gkioxari, G.; Ehsani, K.; Schmidt, L.; and Farhadi, A. 2023. Objaverse- XL: A Universe of 10M+ 3D Objects. a...
2023 arXiv
-
[13]
Deng, B.; Tucker, R.; Li, Z.; Guibas, L.; Snavely, N.; and Wetzstein, G. 2024. Streetscapes: Large-scale Consistent Street View Generation Using Autoregressive Video Diffu- sion. In SIGGRAPH 2024 Conference Papers
2024
-
[14]
Dhariwal, P.; and Nichol, A. 2021. Diffusion models beat gans on image synthesis. Advances in neural information processing systems, 34: 8780–8794
2021
-
[15]
Dockhorn, T.; Vahdat, A.; and Kreis, K. 2021. Score-based generative modeling with critically-damped langevin diffu- sion. arXiv preprint arXiv:2112.07068
2021 arXiv
-
[16]
Dosovitskiy, A.; Ros, G.; Codevilla, F.; Lopez, A.; and Koltun, V . 2017. CARLA: An open urban driving simulator. In Conference on robot learning, 1–16. PMLR
2017
-
[17]
Esser, P.; Rombach, R.; and Ommer, B. 2021. Taming trans- formers for high-resolution image synthesis. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, 12873–12883
2021
-
[18]
R.; Zhou, Y .; Yang, Z.; Chouard, A.; Sun, P.; Ngiam, J.; Vasudevan, V .; Mc- Cauley, A.; Shlens, J.; and Anguelov, D
Ettinger, S.; Cheng, S.; Caine, B.; Liu, C.; Zhao, H.; Prad- han, S.; Chai, Y .; Sapp, B.; Qi, C. R.; Zhou, Y .; Yang, Z.; Chouard, A.; Sun, P.; Ngiam, J.; Vasudevan, V .; Mc- Cauley, A.; Shlens, J.; and Anguelov, D. 2021. Large Scale Interactive Motion Forecasting for Autonom...
2021
-
[19]
Feng, L.; Li, Q.; Peng, Z.; Tan, S.; and Zhou, B. 2023. Traf- ficgen: Learning to generate diverse and realistic traffic sce- narios. In 2023 IEEE International Conference on Robotics and Automation (ICRA), 3567–3575. IEEE
2023
-
[20]
R.; Yang, Y .-H.; Keetha, N
Fischer, T.; Bul `o, S. R.; Yang, Y .-H.; Keetha, N. V .; Porzi, L.; M ¨uller, N.; Schwarz, K.; Luiten, J.; Pollefeys, M.; and Kontschieder, P. 2025. FlowR: Flowing from Sparse to Dense 3D Reconstructions. arXiv preprint arXiv:2504.01647
2025 arXiv
-
[21]
Gao, R.; Chen, K.; Li, Z.; Hong, L.; Li, Z.; and Xu, Q
-
[22]
Gao, R.; Chen, K.; Xiao, B.; Hong, L.; Li, Z.; and Xu, Q. 2024. MagicDriveDiT: High-Resolution Long Video Generation for Autonomous Driving with Adaptive Control. arXiv:2411.13807
2024 arXiv
-
[23]
Gao, R.; Chen, K.; Xie, E.; Hong, L.; Li, Z.; Yeung, D.- Y .; and Xu, Q. 2023. Magicdrive: Street view genera- tion with diverse 3d geometry control. arXiv preprint arXiv:2310.02601
2023 arXiv
-
[24]
P.; Barron, J
Gao*, R.; Holynski*, A.; Henzler, P.; Brussee, A.; Martin- Brualla, R.; Srinivasan, P. P.; Barron, J. T.; and Poole*, B
-
[25]
Gao, S.; Yang, J.; Chen, L.; Chitta, K.; Qiu, Y .; Geiger, A.; Zhang, J.; and Li, H. 2024. Vista: A Generalizable Driving World Model with High Fidelity and Versatile Controllabil- ity. In Advances in Neural Information Processing Systems (NeurIPS)
2024
-
[26]
Geiger, A.; Lenz, P.; and Urtasun, R. 2012. Are we ready for autonomous driving? the kitti vision benchmark suite. In 2012 IEEE conference on computer vision and pattern recognition, 3354–3361. IEEE
2012
-
[27]
Goodfellow, I.; Pouget-Abadie, J.; Mirza, M.; Xu, B.; Warde-Farley, D.; Ozair, S.; Courville, A.; and Bengio, Y
-
[28]
Advances in Neural Information Processing Systems
CAT3D: Create Anything in 3D with Multi-View Dif- fusion Models. Advances in Neural Information Processing Systems
-
[29]
D.; Agarwal, R.; Roelofs, R.; Lu, Y .; Montali, N.; Mougin, P.; Yang, Z.; White, B.; Faust, A.; McAllister, R.; Anguelov, D.; and Sapp, B
Gulino, C.; Fu, J.; Luo, W.; Tucker, G.; Bronstein, E.; Lu, Y .; Harb, J.; Pan, X.; Wang, Y .; Chen, X.; Co-Reyes, J. D.; Agarwal, R.; Roelofs, R.; Lu, Y .; Montali, N.; Mougin, P.; Yang, Z.; White, B.; Faust, A.; McAllister, R.; Anguelov, D.; and Sapp, B. 2023. Waymax: An Acc...
2023
-
[30]
Harvey, W.; Naderiparizi, S.; Masrani, V .; Weilbach, C.; and Wood, F. 2022. Flexible diffusion modeling of long videos. Advances in Neural Information Processing Systems , 35: 27953–27965
2022
-
[31]
Hess, G.; Lindstr ¨om, C.; Fatemi, M.; Petersson, C.; and Svensson, L. 2025. Splatad: Real-time lidar and camera ren- dering with 3d gaussian splatting for autonomous driving. In Proceedings of the Computer Vision and Pattern Recog- nition Conference, 11982–11992
2025
-
[32]
P.; Poole, B.; Norouzi, M.; Fleet, D
Ho, J.; Chan, W.; Saharia, C.; Whang, J.; Gao, R.; Gritsenko, A.; Kingma, D. P.; Poole, B.; Norouzi, M.; Fleet, D. J.; et al
-
[33]
Greene, N. 1986. Environment mapping and other appli- cations of world projections. IEEE computer graphics and Applications, 6(11): 21–29
1986
-
[34]
Ho, J.; Salimans, T.; Gritsenko, A.; Chan, W.; Norouzi, M.; and Fleet, D. J. 2022. Video diffusion models. Advances in Neural Information Processing Systems, 35: 8633–8646
2022
-
[35]
Hong, W.; Ding, M.; Zheng, W.; Liu, X.; and Tang, J. 2022. CogVideo: Large-scale Pretraining for Text-to-Video Gen- eration via Transformers. arXiv preprint arXiv:2205.15868
2022 arXiv
-
[36]
Huang, B.; Yu, Z.; Chen, A.; Geiger, A.; and Gao, S. 2024. 2D Gaussian Splatting for Geometrically Accurate Radiance Fields. In SIGGRAPH 2024 Conference Papers. Association for Computing Machinery
2024
-
[37]
Huang, J.; Gojcic, Z.; Atzmon, M.; Litany, O.; Fidler, S.; and Williams, F. 2023. Neural Kernel Surface Reconstruction. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 4369–4379
2023
-
[38]
Hwang, S.; Kim, M.-J.; Kang, T.; Kang, J.; and Choo, J
-
[39]
Ho, J.; Jain, A.; and Abbeel, P. 2020. Denoising diffusion probabilistic models. Advances in neural information pro- cessing systems, 33: 6840–6851
2020
-
[40]
Karras, T.; Laine, S.; Aittala, M.; Hellsten, J.; Lehtinen, J.; and Aila, T. 2020. Analyzing and improving the image qual- ity of stylegan. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 8110–8119
2020
-
[41]
Kazemkhani, S.; Pandya, A.; Cornelisse, D.; Shacklett, B.; and Vinitsky, E. 2025. GPUDrive: Data-driven, multi-agent driving simulation at 1 million FPS. In Proceedings of the International Conference on Learning Representations (ICLR)
2025
-
[42]
Kerbl, B.; Kopanas, G.; Leimk ¨uhler, T.; and Drettakis, G
-
[43]
Kheradmand, S.; Rebain, D.; Sharma, G.; Sun, W.; Tseng, J.; Isack, H.; Kar, A.; Tagliasacchi, A.; and Yi, K. M
-
[44]
W.; Brown, B.; Yin, K.; Kreis, K.; Schwarz, K.; Li, D.; Rombach, R.; Torralba, A.; and Fidler, S
Kim, S. W.; Brown, B.; Yin, K.; Kreis, K.; Schwarz, K.; Li, D.; Rombach, R.; Torralba, A.; and Fidler, S. 2023. NeuralField-LDM: Scene Generation With Hierarchical La- tent Diffusion Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (...
2023
-
[45]
In European Conference on Computer Vision, 1–18
Vegs: View extrapolation of urban scenes in 3d gaus- sian splatting using learned priors. In European Conference on Computer Vision, 1–18. Springer
-
[46]
Jun, H.; and Nichol, A. 2023. Shap-E: Generating Condi- tional 3D Implicit Functions. arXiv:2305.02463
2023 arXiv
-
[47]
Li, H.; Shi, H.; Zhang, W.; Wu, W.; Liao, Y .; Wang, L.; Lee, L.-h.; and Zhou, P. 2024. DreamScene: 3D Gaussian-based Text-to-3D Scene Generation via Formation Pattern Sam- pling. arXiv preprint arXiv:2404.03575
2024 arXiv
-
[48]
Li, Q.; Peng, Z.; Feng, L.; Zhang, Q.; Xue, Z.; and Zhou, B
-
[49]
Li, Y .; Zou, Z.-X.; Liu, Z.; Wang, D.; Liang, Y .; Yu, Z.; Liu, X.; Guo, Y .-C.; Liang, D.; Ouyang, W.; et al. 2025. Tri- poSG: High-Fidelity 3D Shape Synthesis using Large-Scale Rectified Flow Models. arXiv preprint arXiv:2502.06608
2025 arXiv
-
[50]
Liang, Y .; Yang, X.; Lin, J.; Li, H.; Xu, X.; and Chen, Y . 2024. Luciddreamer: Towards high-fidelity text-to-3d generation via interval score matching. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6517–6526
2024
-
[51]
Liao, Y .; Xie, J.; and Geiger, A. 2022. Kitti-360: A novel dataset and benchmarks for urban scene understanding in 2d and 3d. IEEE Transactions on Pattern Analysis and Machine Intelligence, 45(3): 3292–3310
2022
- [52]
-
[53]
Liu, X.; Zhou, C.; and Huang, S. 2024. 3dgs-enhancer: Enhancing unbounded 3d gaussian splatting with view- consistent 2d diffusion priors. Advances in Neural Infor- mation Processing Systems, 37: 133305–133327
2024
-
[54]
P.; and Welling, M
Kingma, D. P.; and Welling, M. 2013. Auto-encoding vari- ational bayes. arXiv preprint arXiv:1312.6114
2013 arXiv
-
[55]
Lee, J.; Lee, S.; Jo, C.; Im, W.; Seon, J.; and Yoon, S.-E
-
[56]
In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
SemCity: Semantic Scene Generation with Triplane Diffusion. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition
-
[57]
Luo, S.; Tan, Y .; Huang, L.; Li, J.; and Zhao, H. 2023. Latent consistency models: Synthesizing high-resolution images with few-step inference. arXiv preprint arXiv:2310.04378
2023 arXiv
-
[58]
Miao, X.; Duan, H.; Ojha, V .; Song, J.; Shah, T.; Long, Y .; and Ranjan, R. 2024. Dreamer XL: Towards High- Resolution Text-to-3D Generation via Trajectory Score Matching. arXiv preprint arXiv:2405.11252
2024 arXiv
-
[59]
IEEE Transactions on Pattern Analysis and Machine Intelligence
Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning. IEEE Transactions on Pattern Analysis and Machine Intelligence
-
[60]
NVIDIA; Agarwal, N.; Ali, A.; Bala, M.; Balaji, Y .; Barker, E.; Cai, T.; Chattopadhyay, P.; Chen, Y .; Cui, Y .; Ding, Y .; Dworakowski, D.; Fan, J.; Fenzi, M.; Ferroni, F.; Fidler, S.; Fox, D.; Ge, S.; Ge, Y .; Gu, J.; Gururani, S.; He, E.; Huang, J.; Huffman, J.; Jannaty, P...
2025
-
[61]
Oquab, M.; Darcet, T.; Moutakanni, T.; V o, H.; Szafraniec, M.; Khalidov, V .; Fernandez, P.; Haziza, D.; Massa, F.; El- Nouby, A.; et al. 2023. Dinov2: Learning robust visual fea- tures without supervision.arXiv preprint arXiv:2304.07193
2023 arXiv
-
[62]
Ost, J.; Mannan, F.; Thuerey, N.; Knodt, J.; and Heide, F
-
[63]
H.; Lee, H.-Y .; Menapace, W.; Chai, M.; Siarohin, A.; Yang, M.-H.; and Tulyakov, S
Lin, C. H.; Lee, H.-Y .; Menapace, W.; Chai, M.; Siarohin, A.; Yang, M.-H.; and Tulyakov, S. 2023. Infinicity: Infinite- scale city synthesis. In Proceedings of the IEEE/CVF inter- national conference on computer vision, 22808–22818
2023
-
[64]
T.; and Mildenhall, B
Poole, B.; Jain, A.; Barron, J. T.; and Mildenhall, B. 2022. DreamFusion: Text-to-3D using 2D Diffusion. arXiv
2022
-
[65]
Ljungbergh, W.; Tonderski, A.; Johnander, J.; Caesar, H.; ˚Astr¨om, K.; Felsberg, M.; and Petersson, C. 2024. Neu- roNCAP: Photorealistic Closed-loop Safety Testing for Au- tonomous Driving. arXiv preprint arXiv:2404.07762
2024 arXiv
-
[66]
Lu, J.; Huang, Z.; Zhang, J.; Yang, Z.; and Zhang, L. 2024. WoV oGen: World V olume-aware Diffusion for Controllable Multi-camera Driving Scene Generation. In European Con- ference on Computer Vision (ECCV)
2024
-
[67]
Lu, Y .; Ren, X.; Yang, J.; Shen, T.; Wu, Z.; Gao, J.; Wang, Y .; Chen, S.; Chen, M.; Fidler, S.; and Huang, J. 2024. In- finiCube: Unbounded and Controllable Dynamic 3D Driv- ing Scene Generation with World-Guided Video Models. arXiv:2412.03934
2024 arXiv
-
[68]
Z.; Chen, R.; Kim, S
Ren, X.; Lu, Y .; Cao, T.; Gao, R.; Huang, S.; Sabour, A.; Shen, T.; Pfaff, T.; Wu, J. Z.; Chen, R.; Kim, S. W.; Gao, J.; Leal-Taixe, L.; Chen, M.; Fidler, S.; and Ling, H
-
[69]
Z.; Ling, H.; Chen, M.; Fidler, F., Sanja annd Williams; and Huang, J
Ren, X.; Lu, Y .; Liang, H.; Wu, J. Z.; Ling, H.; Chen, M.; Fidler, F., Sanja annd Williams; and Huang, J. 2024. SCube: Instant Large-Scale Scene Reconstruction using V oxSplats. In The Thirty-eighth Annual Conference on Neural Informa- tion Processing Systems
2024
-
[70]
P.; Tancik, M.; Barron, J
Mildenhall, B.; Srinivasan, P. P.; Tancik, M.; Barron, J. T.; Ramamoorthi, R.; and Ng, R. 2021. Nerf: Representing scenes as neural radiance fields for view synthesis. Com- munications of the ACM, 65(1): 99–106
2021
-
[71]
Rombach, R.; Blattmann, A.; Lorenz, D.; Esser, P.; and Om- mer, B. 2022. High-resolution image synthesis with latent diffusion models. In Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition , 10684– 10695
2022
-
[72]
H.; Tabatabaee, H.; Lu, Q.; Lemke, S.; Moˇzeiko, M.; Boise, E.; Uhm, G.; Gerow, M.; Mehta, S.; et al
Rong, G.; Shin, B. H.; Tabatabaee, H.; Lu, Q.; Lemke, S.; Moˇzeiko, M.; Boise, E.; Uhm, G.; Gerow, M.; Mehta, S.; et al. 2020. Lgsvl simulator: A high fidelity simulator for autonomous driving. In 2020 IEEE 23rd International con- ference on intelligent transportation systems ...
2020
-
[73]
R.; Lagun, D.; Fei-Fei, L.; Sun, D.; et al
Sargent, K.; Li, Z.; Shah, T.; Herrmann, C.; Yu, H.-X.; Zhang, Y .; Chan, E. R.; Lagun, D.; Fei-Fei, L.; Sun, D.; et al
-
[74]
Sauer, A.; Schwarz, K.; and Geiger, A. 2022. Stylegan-xl: Scaling stylegan to large diverse datasets. In ACM SIG- GRAPH 2022 conference proceedings, 1–10
2022
-
[75]
Podell, D.; English, Z.; Lacey, K.; Blattmann, A.; Dockhorn, T.; M¨uller, J.; Penna, J.; and Rombach, R. 2023. Sdxl: Im- proving latent diffusion models for high-resolution image synthesis. arXiv preprint arXiv:2307.01952
2023 arXiv
-
[76]
Shriram, J.; Trevithick, A.; Liu, L.; and Ramamoorthi, R. 2024. Realmdreamer: Text-driven 3d scene genera- tion with inpainting and depth diffusion. arXiv preprint arXiv:2404.07199
2024 arXiv
-
[77]
W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I
Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021. Learning Transferable Visual Models From Natural Language Supervision. In In- ternational Conference on Machine Learning
2021
-
[78]
Razavi, A.; Van den Oord, A.; and Vinyals, O. 2019. Gener- ating diverse high-fidelity images with vq-vae-2. Advances in neural information processing systems, 32
2019
-
[79]
Ren, X.; Huang, J.; Zeng, X.; Museth, K.; Fidler, S.; and Williams, F. 2024. XCube: Large-Scale 3D Generative Modeling using Sparse V oxel Hierarchies. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition
2024
-
[80]
Song, Y .; Sun, Z.; and Yin, X. 2024. SDXS: Real-Time One- Step Latent Diffusion Models with Image Conditions.arXiv preprint arXiv:2403.16627
2024 arXiv
-
[81]
L.; Taylor, E.; and Loaiza-Ganem, G
Stein, G.; Cresswell, J.; Hosseinzadeh, R.; Sui, Y .; Ross, B.; Villecroze, V .; Liu, Z.; Caterini, A. L.; Taylor, E.; and Loaiza-Ganem, G. 2023. Exposing flaws of generative model evaluation metrics and their unfair treatment of dif- fusion models. Advances in Neural Informat...
2023
-
[82]
Sun, P.; Kretzschmar, H.; Dotiwalla, X.; Chouard, A.; Pat- naik, V .; Tsui, P.; Guo, J.; Zhou, Y .; Chai, Y .; Caine, B.; Vasudevan, V .; Han, W.; Ngiam, J.; Zhao, H.; Timofeev, A.; Ettinger, S.; Krivokon, M.; Gao, A.; Joshi, A.; Zhang, Y .; Shlens, J.; Chen, Z.; and Anguelov,...
2020
-
[83]
Ren, X.; Shen, T.; Huang, J.; Ling, H.; Lu, Y .; Nimier-David, M.; M ¨uller, T.; Keller, A.; Fidler, S.; and Gao, J. 2025. Gen3c: 3d-informed world-consistent video generation with precise camera control. In Proceedings of the Computer Vi- sion and Pattern Recognition Conferen...
2025
-
[84]
Team, A.; Zhu, H.; Wang, Y .; Zhou, J.; Chang, W.; Zhou, Y .; Li, Z.; Chen, J.; Shen, C.; Pang, J.; et al. 2025. Aether: Geometric-aware unified world modeling. arXiv preprint arXiv:2503.18945
2025 arXiv
-
[85]
Team, G. 2024. Mochi 1. https://github.com/genmoai/ models
2024
-
[86]
Team, T. H. 2025. Hunyuan3D 2.0: Scaling Diffusion Mod- els for High Resolution Textured 3D Assets Generation. arXiv:2501.12202
2025 arXiv
-
[87]
arXiv preprint arXiv:2310.17994
Zeronvs: Zero-shot 360-degree view synthesis from a single real image. arXiv preprint arXiv:2310.17994
-
[88]
Tonderski, A.; Lindstr ¨om, C.; Hess, G.; Ljungbergh, W.; Svensson, L.; and Petersson, C. 2023. NeuRAD: Neu- ral rendering for autonomous driving. arXiv preprint arXiv:2311.15260
2023 arXiv
-
[89]
Shah, S.; Dey, D.; Lovett, C.; and Kapoor, A. 2018. Airsim: High-fidelity visual and physical simulation for autonomous vehicles. In Field and Service Robotics: Results of the 11th International Conference, 621–635. Springer
2018
-
[90]
Vahdat, A.; Kreis, K.; and Kautz, J. 2021. Score-based gen- erative modeling in latent space. Advances in neural infor- mation processing systems, 34: 11287–11302
2021
-
[91]
R.; Chan, E
Shue, J. R.; Chan, E. R.; Po, R.; Ankner, Z.; Wu, J.; and Wetzstein, G. 2023. 3d neural field generation using triplane diffusion. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 20875–20886
2023
-
[92]
Singer, U.; Polyak, A.; Hayes, T.; Yin, X.; An, J.; Zhang, S.; Hu, Q.; Yang, H.; Ashual, O.; Gafni, O.; et al. 2022. Make- a-video: Text-to-video generation without text-video data. arXiv preprint arXiv:2209.14792
2022 arXiv
-
[93]
Song, J.; Meng, C.; and Ermon, S. 2020. Denoising Diffu- sion Implicit Models. arXiv:2010.02502
2020 arXiv
-
[94]
Williams, F.; Gojcic, Z.; Khamis, S.; Zorin, D.; Bruna, J.; Fidler, S.; and Litany, O. 2021. Neural Fields as Learnable Kernels for 3D Reconstruction. arXiv:2111.13674
2021 arXiv
-
[95]
Z.; Zhang, Y .; Turki, H.; Ren, X.; Gao, J.; Shou, M
Wu, J. Z.; Zhang, Y .; Turki, H.; Ren, X.; Gao, J.; Shou, M. Z.; Fidler, S.; Gojcic, Z.; and Ling, H. 2025. Di- fix3d+: Improving 3d reconstructions with single-step dif- fusion models. In Proceedings of the Computer Vision and Pattern Recognition Conference, 26024–26035
2025
-
[96]
T.; and Holynski, A
Wu, R.; Gao, R.; Poole, B.; Trevithick, A.; Zheng, C.; Barron, J. T.; and Holynski, A. 2024. CAT4D: Create Anything in 4D with Multi-View Video Diffusion Models. arXiv:2411.18613
2024 arXiv
-
[97]
Talwar, D.; Guruswamy, S.; Ravipati, N.; and Eirinaki, M
-
[98]
In 2020 IEEE International Conference On Artificial Intelligence Testing (AITest) , 73–
Evaluating validity of synthetic data in perception tasks for autonomous vehicles. In 2020 IEEE International Conference On Artificial Intelligence Testing (AITest) , 73–
2020
-
[99]
Xiang, J.; Lv, Z.; Xu, S.; Deng, Y .; Wang, R.; Zhang, B.; Chen, D.; Tong, X.; and Yang, J. 2024. Structured 3d la- tents for scalable and versatile 3d generation.arXiv preprint arXiv:2412.01506
2024 arXiv
-
[100]
Xie, H.; Chen, Z.; Hong, F.; and Liu, Z. 2024. Citydreamer: Compositional generative model of unbounded 3d cities. In Proceedings of the IEEE/CVF conference on computer vi- sion and pattern recognition, 9666–9675
2024
-
[101]
Xie, K.; Lorraine, J.; Cao, T.; Gao, J.; Lucas, J.; Torralba, A.; Fidler, S.; and Zeng, X. 2024. LATTE3D: Large-scale Amortized Text-To-Enhanced3D Synthesis. arXiv preprint arXiv:2403.15385
2024 arXiv
-
[102]
Thies, J.; Zollh ¨ofer, M.; and Nießner, M. 2019. Deferred neural rendering: Image synthesis using neural textures. Acm Transactions on Graphics (TOG), 38(4): 1–12
2019
-
[103]
Yang, J.; Gao, S.; Qiu, Y .; Chen, L.; Li, T.; Dai, B.; Chitta, K.; Wu, P.; Zeng, J.; Luo, P.; Zhang, J.; Geiger, A.; Qiao, Y .; and Li, H. 2024. Generalized Predictive Model for Au- tonomous Driving. In Proceedings of the IEEE/CVF Con- ference on Computer Vision and Pattern R...
2024
-
[104]
Vahdat, A.; and Kautz, J. 2020. NV AE: A deep hierarchi- cal variational autoencoder. Advances in neural information processing systems, 33: 19667–19679
2020
-
[105]
Yang, S.; Hou, L.; Huang, H.; Ma, C.; Wan, P.; Zhang, D.; Chen, X.; and Liao, J. 2024. Direct-a-Video: Customized Video Generation with User-Directed Camera Movement and Object Motion. arXiv preprint arXiv:2402.03162
2024 arXiv
-
[106]
Wan, T.; Wang, A.; Ai, B.; Wen, B.; Mao, C.; Xie, C.-W.; Chen, D.; Yu, F.; Zhao, H.; Yang, J.; et al. 2025. Wan: Open and advanced large-scale video generative models. arXiv preprint arXiv:2503.20314
2025 arXiv
-
[107]
Wang, P.; Xu, D.; Fan, Z.; Wang, D.; Mohan, S.; Iandola, F.; Ranjan, R.; Li, Y .; Liu, Q.; Wang, Z.; et al. 2023. Taming Mode Collapse in Score Distillation for Text-to-3D Genera- tion. arXiv preprint arXiv:2401.00909
2023 arXiv
-
[108]
Wang, X.; Zhu, Z.; Huang, G.; Chen, X.; and Lu, J. 2023. Drivedreamer: Towards real-world-driven world models for autonomous driving. arXiv preprint arXiv:2309.09777
2023 arXiv
-
[109]
Yi, T.; Fang, J.; Wang, J.; Wu, G.; Xie, L.; Zhang, X.; Liu, W.; Tian, Q.; and Wang, X. 2024. GaussianDreamer: Fast Generation from Text to 3D Gaussians by Bridging 2D and 3D Diffusion Models. In CVPR
2024
-
[110]
T.; and Wu, J
Yu, H.-X.; Duan, H.; Herrmann, C.; Freeman, W. T.; and Wu, J. 2024. WonderWorld: Interactive 3D Scene Genera- tion from a Single Image. arXiv preprint arXiv:2406.09394
2024 arXiv
-
[111]
T.; Cole, F.; Sun, D.; Snavely, N.; Wu, J.; et al
Yu, H.-X.; Duan, H.; Hur, J.; Sargent, K.; Rubinstein, M.; Freeman, W. T.; Cole, F.; Sun, D.; Snavely, N.; Wu, J.; et al
-
[112]
P.; Verbin, D.; Barron, J
Wu, R.; Mildenhall, B.; Henzler, P.; Park, K.; Gao, R.; Wat- son, D.; Srinivasan, P. P.; Verbin, D.; Barron, J. T.; Poole, B.; et al. 2024. Reconfusion: 3d reconstruction with diffu- sion priors. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognit...
2024
-
[113]
Wu, Z.; Liu, T.; Luo, L.; Zhong, Z.; Chen, J.; Xiao, H.; Hou, C.; Lou, H.; Chen, Y .; Yang, R.; et al. 2023. Mars: An instance-aware, modular and realistic simulator for au- tonomous driving. In CAAI International Conference on Ar- tificial Intelligence, 3–15. Springer
2023
-
[114]
A.; Shechtman, E.; and Wang, O
Zhang, R.; Isola, P.; Efros, A. A.; Shechtman, E.; and Wang, O. 2018. The Unreasonable Effectiveness of Deep Features as a Perceptual Metric. In CVPR
2018
-
[115]
Zhang, S.; Zhang, Y .; Zheng, Q.; Ma, R.; Hua, W.; Bao, H.; Xu, W.; and Zou, C. 2024. 3D-SceneDreamer: Text- Driven 3D-Consistent Scene Generation. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 10170–10180
2024
-
[116]
Zhang, Z.; Long, F.; Pan, Y .; Qiu, Z.; Yao, T.; Cao, Y .; and Mei, T. 2024. TRIP: Temporal Residual Learning with Im- age Noise Prior for Image-to-Video Diffusion Models.arXiv preprint arXiv:2403.17005
2024 arXiv
-
[117]
Xu, Y .; Chai, M.; Shi, Z.; Peng, S.; Skorokhodov, I.; Siaro- hin, A.; Yang, C.; Shen, Y .; Lee, H.-Y .; Zhou, B.; et al
-
[118]
In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, 4402–4412
Discoscene: Spatially disentangled generative radi- ance fields for controllable 3d-aware scene synthesis. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Recognition, 4402–4412
-
[119]
Zyrianov, V .; Che, H.; Liu, Z.; and Wang, S. 2024. Li- darDM: Generative LiDAR Simulation in a Generated World. arXiv preprint arXiv:2404.02903
2024
-
[120]
Yang, J.; Huang, J.; Chen, Y .; Wang, Y .; Li, B.; You, Y .; Igl, M.; Sharma, A.; Karkus, P.; Xu, D.; Ivanovic, B.; Wang, Y .; and Pavone, M. 2025. STORM: Spatio-Temporal Re- construction Model for Large-scale Outdoor Scenes. arXiv preprint arXiv:2501.00602
2025 arXiv
-
[122]
Yang, Y .; Yang, Y .; Guo, H.; Xiong, R.; Wang, Y .; and Liao, Y . 2023. Urbangiraffe: Representing urban scenes as com- positional generative neural feature fields. In Proceedings of the IEEE/CVF International Conference on Computer Vi- sion, 9199–9210
2023
-
[123]
J.; and Urtasun, R
Yang, Z.; Chen, Y .; Wang, J.; Manivasagam, S.; Ma, W.- C.; Yang, A. J.; and Urtasun, R. 2023. UniSim: A Neural Closed-Loop Sensor Simulator. In CVPR
2023
-
[124]
Yang, Z.; Teng, J.; Zheng, W.; Ding, M.; Huang, S.; Xu, J.; Yang, Y .; Hong, W.; Zhang, X.; Feng, G.; et al. 2024. CogVideoX: Text-to-Video Diffusion Models with An Ex- pert Transformer. arXiv preprint arXiv:2408.06072
2024 arXiv
-
[128]
In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6658–6667
Wonderjourney: Going from anywhere to everywhere. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 6658–6667
-
[129]
R.; Liu, G.; and Zhou, B
Zhang, J.; Zhang, Q.; Zhang, L.; Kompella, R. R.; Liu, G.; and Zhou, B. 2024. Urban Scene Diffusion through Seman- tic Occupancy Map. arXiv preprint arXiv:2403.11697
2024 arXiv
-
[130]
Zhang, L.; Rao, A.; and Agrawala, M. 2023. Adding condi- tional control to text-to-image diffusion models. InProceed- ings of the IEEE/CVF international conference on computer vision, 3836–3847
2023
-
[134]
Zhengwentai, S. 2023. clip-score: CLIP Score for PyTorch. https://github.com/taited/clip-score. Version 0.2.1
2023
-
[135]
Zhou, L.; Du, Y .; and Wu, J. 2021. 3D Shape Generation and Completion Through Point-V oxel Diffusion. InProceedings of the IEEE/CVF International Conference on Computer Vi- sion (ICCV), 5826–5835
2021
-
[2014]
Advances in neural in- formation processing systems, 27
Generative adversarial nets. Advances in neural in- formation processing systems, 27
-
[2016]
In Proceedings of the IEEE conference on com- puter vision and pattern recognition, 3213–3223
The cityscapes dataset for semantic urban scene un- derstanding. In Proceedings of the IEEE conference on com- puter vision and pattern recognition, 3213–3223
-
[2020]
nuScenes: A multimodal dataset for autonomous driv- ing. In CVPR
-
[2021]
In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2856–2865
Neural scene graphs for dynamic scenes. In Proceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2856–2865
-
[2022]
arXiv preprint arXiv:2210.02303
Imagen video: High definition video generation with diffusion models. arXiv preprint arXiv:2210.02303
-
[2023]
ACM Transactions on Graphics, 42(4)
3D Gaussian Splatting for Real-Time Radiance Field Rendering. ACM Transactions on Graphics, 42(4)
-
[2024]
arXiv preprint arXiv:2405.14475
MagicDrive3D: Controllable 3D Generation for Any-View Rendering in Street Scenes. arXiv preprint arXiv:2405.14475
-
[2025]
Cosmos-Drive-Dreams: Scalable Synthetic Driving Data Generation with World Foundation Models
Reviewed August 15, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.