REVIEW 5 major objections 4 minor 59 references
SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model
T0 review · 5 major / 4 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read This paper claims that a single diffusion-based generative world model, trained end-to-end on one denoising loss, can simulate a full trip from point A to point B at city scale, jointly generating agent trajectories, agent validity…
desk verdict A real step forward for trip-level traffic simulation with a clever sparse-tensor validity trick, but the 'point A-to-B' claim outruns the experiments and the validity channel's sensor-relative semantics deserve closer scrutiny. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The load-bearing object is the multi-tensor scene representation $X = \{x_{\mathrm{agent}}, x_{\mathrm{light}}\}$, where each $x_i \in \mathbb{R}^{E_i \times T \times D_i}$ stacks $E_i$ elements over $T$ timesteps and each element's last feature is a validity bit. The denoiser is an axial-attention transformer (the SceneDiffuser backbone) whose input is the element-concatenated, hidden-dimension-homogenized tensor, trained with the single v-prediction diffusion loss. Two sparse-tensor mechanisms carry the argument: during training, invalid entries are zeroed and the loss is masked so only the validity channel is supervised at invalid steps; during inference, soft clipping multiplies the denoised feature values by the predicted validity, so invalid entries decay toward zero instead of drifting out of distribution. That combination is what lets one model output a time-varying, sparse set of valid agents and lights, and it is the mechanism behind the paper's claimed ability to handle spawning, removal, occlusion, and traffic signals jointly.
What would settle it
Replay 60-second SceneDiffuser++ rollouts with a fixed route-unconditioned planner, then ray-cast from the simulated ego pose against the map geometry to compute the physically visible agent set at each step; if the model's predicted validity (or its entering and exiting events) matches ray-cast visibility no better than a baseline that ignores map occlusion, the claim that the validity channel encodes occlusion and visibility would fail.
Extended reading notes
Core claim
The central claim is that agent validity — whether an agent appears in the AV's detection output at a given timestep — can be treated as just another channel of the scene tensor and learned jointly with positions, sizes, and object types. Because validity is generated by the diffusion model itself rather than taken from the log, the model can emit sparse tensors without a prespecified sparsity structure: it learns when to insert a new agent, when to let an existing agent leave, and when an agent is occluded or disoccluded, by interpolating between valid and zero-invalid states during sampling. The same multi-tensor construction also carries a second tensor for traffic lights, with states and positions denoised alongside agents, so the entire trip is one inpainting problem. The paper reports state-of-the-art trip-level realism in 60-second rollouts, with substantially lower Jensen-Shannon divergence than IDM and SceneDiffuser for the distributions that depend directly on validity, while also matching logged traffic-light transition probabilities.
Load-bearing premise
The validity channel is supervised using the logged AV's detection output, so 'valid' means visible to that particular sensor trajectory; in route-unconditioned rollouts the simulated AV drives elsewhere, and nothing in the training objective ties predicted validity to actual visibility from the simulated ego, so the spawned and removed agents may not be the ones the simulated AV could actually see.
Editorial extensions
If this is right
- Trip-level statistics such as travel time, pick-up and drop-off behavior, and long-horizon interactions with emergency vehicles can be generated from logged data alone, without hand-written insertion and removal rules.
- Because the world model predicts traffic-light state transitions, violations of light rules can be measured in simulation rather than assumed from map priors.
- The validity channel removes the hard limit that previously tied simulation duration to the length of the log: agents can be recycled into previously vacated slots, so arbitrarily many agents can appear over long horizons.
- More frequent planner-world-model interaction lowers collision rates and improves speed realism, suggesting the model is most realistic when the AV and background agents continuously react to each other.
- Rollout realism degrades as the horizon grows, especially for the timing of agent insertion, which indicates that autoregressive error accumulation, not map coverage, is the next bottleneck for trip-level simulation.
Reading between the lines
- We would expect the sparse-tensor validity trick to transfer to other generative settings where outputs are naturally sparse, such as point-cloud completion with missing measurements or simultaneous generation of landmarks and objects; the paper does not test this.
- A natural extension, not pursued in the paper, is to condition validity on the simulated ego's sensor footprint rather than on the logged AV's detection output, so that spawn and removal events remain physically meaningful when the simulated route diverges from the log.
- The histogram-based Jensen-Shannon divergence protocol is the paper's own proposal; a reasonable next step for the field is to make such correspondence-free, sliding-window distributional metrics standard for long-horizon simulators.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes SceneDiffuser++, a diffusion-based generative world model for city-scale traffic simulation. The model jointly predicts, from a single denoising objective, agent trajectories, a validity channel that encodes agent presence/visibility, and traffic light states, via an autoregressive rollout over a sparse scene tensor. The authors introduce soft clipping at inference to stabilize the generation of sparse tensors, a multi-tensor architecture to jointly handle agents and traffic lights with different feature dimensions, and a map-extended version of the Waymo Open Motion Dataset (WOMD-XLMap) for trip-level simulation. They evaluate realism using JS divergence between histograms of simulated and logged metrics (agent counts, entering/exiting distances, offroad and collision rates, speeds, traffic-light transitions), reporting quantitative gains over IDM and SceneDiffuser baselines on 60s rollouts.
Significance. If the central claims are upheld, the paper would make a notable contribution: a single end-to-end model that handles scene generation, dynamic agent insertion/removal, occlusion, and traffic-light simulation without task-specific heads or heuristics. The sparse-tensor learning method is an interesting technical idea, and the joint modeling of agents and traffic lights is a useful step beyond prior work. The paper also provides an enlarged dataset and a nontrivial evaluation protocol for long-horizon simulation, which is valuable to the community. However, the claim of 'point A-to-B simulation' is not actually exercised in the experiments, and the validity channel, which is central to the new capabilities, has a train/inference semantic mismatch. These issues currently prevent the paper's strongest claims from being supported.
major comments (5)
- [Abstract and Appendix A.4] The abstract and Section 1 claim that SceneDiffuser++ is capable of 'point A-to-B simulation on a city scale', but Appendix A.4 explicitly states: 'SceneDiffuser and SceneDiffuser++ do not use goal-oriented routing; in other words, they do not use or ingest a goal location in any way, shape, or form.' Figure 1 depicts a 'trip end' star but no conditioning mechanism is described or tested. The route-unconditioned rollouts used in Section 4.1 do not implement point-to-point trips, so this central claim is unsupported as stated. Either the model must be extended with goal/route conditioning, or the paper should be reframed around route-unconditioned world-model simulation and the point-to-point claim removed from the abstract and contributions.
- [Appendix A.4 (Validity Definition) and Section 3 (Learning Sparse Tensors)] The validity channel is trained to reproduce whether an agent appears in the logged AV's detection output, and all positions are normalized by the logged AV's ego pose (Section 3). In route-unconditioned rollouts, the simulated AV follows an arbitrary path that diverges from the logged trajectory, and nothing in the loss, architecture, or sampling procedure ties predicted validity to the simulated AV's actual sensor frustum or detection range. The generated validity can therefore encode 'agent exists on the map' rather than 'agent is visible to this AV'. This undermines the paper's occlusion-reasoning claims and makes the validity-based metrics (# valid agents, entering/exiting agents and distances) an unvalidated proxy for sensor-relative realism. The authors should either constrain validity at inference with the simulated AV's sensor geometry (e.g., raycasting or frustum culling) or provide evidence that the learned validity remains consistent with the simulated AV's visibility.
- [Section 4.1 (Metrics)] The evaluation relies on JS divergence between histograms of simulated and logged metric values, but the paper does not specify histogram bin width or number of bins, report error bars or confidence intervals, state the number of independent rollout seeds, or validate that the divergence is sensitive to the sample sizes used. The numbers in Table 1, such as 0.3132 vs. 0.2206 for # Valid Agents, may not be statistically distinguishable. In addition, the Composite score averages metrics with different units and scales without justification. Without these details, the headline 'state-of-the-art trip-level simulation realism' is not established. Adding standard errors over seeds and explicit histogram construction would make the comparison interpretable.
- [Table 1 and Section 4.2] Several comparison cells are missing for the IDM and SceneDiffuser baselines: Entering and Exiting Distances are not reported for these world models, and Traffic Light Violation and Traffic Light Transition are not reported for any baseline. The text states that SceneDiffuser++ 'achieves significantly better performance in all metrics that relate to agent insertion and removal', but two of the four insertion/removal metrics lack baseline values. Moreover, because IDM and SceneDiffuser never insert agents, their validity distributions are degenerate by construction, which can inflate or deflate JS divergence in ways not discussed. The comparison should either report the missing values (with a principled definition for non-inserting baselines) or restrict the superiority claim to the metrics actually measured against both baselines.
- [Section 4.3, Table 3] The text in Section 3 states that imputing invalid values with zeros 'cannot work', yet Table 3 includes a 'No Clipping' row described as 'a model trained to directly predict invalid agents' features to be 0', which yields a lower JS divergence for # Valid Agents (0.2426) than Soft Clipping (0.3053) and a comparable # Entering Agents (0.2035 vs. 0.2120). The claim that this approach 'cannot work' is internally inconsistent with the reported data. The ablation should present the full trade-off (e.g., no clipping is better on some agent-count metrics, worse on collision/offroad/TL), and the choice of soft clipping should be justified by the metric priorities of the target use case rather than by a blanket statement.
minor comments (4)
- [Figure 7] The legend in the bottom plot uses 'SceneDiffusion', which should be 'SceneDiffuser' for consistency with the rest of the paper.
- [Section 3, Eq. (2)] The loss weight w is described only as 'a loss weighting term'; its exact form (including the loss mask described later in the same section) should be specified in the equation or immediately after, since it is central to the sparse-tensor training.
- [Figure 6] The 'Ground-truth Log' panel covers only 91 steps but the predicted panels show 600 steps; the caption should state this difference explicitly to avoid implying direct temporal alignment beyond step 91.
- [Section 4.2] The phrase 'Average Speed likelihood' appears to mix a distributional metric with a probabilistic notion; consider rewording to 'Average Speed distribution realism'.
Circularity Check
No significant circularity: the validity-channel concern is a train/deployment mismatch, not an input-output reduction.
full rationale
The paper's derivation chain is self-contained against external data: it defines scene tensors, trains a denoiser with the v-prediction loss in Eq. (2) on logged WOMD-XLMap data, rolls out autoregressively, and evaluates realism via JS divergence between simulated and held-out log histograms. No parameter is fitted to the evaluation metrics; the soft-clipping sampler is an inference-time decoding choice ablated on validation behavior, not a quantity derived from the metrics it later ranks. The SceneDiffuser backbone is a self-citation, but it is used as an architectural component rather than as a premise that forces the validity, traffic-light, or spawning results. The closest concern is Appendix A.4's sensor-relative validity definition combined with route-unconditioned rollout: log validity means 'appears in the logged AV's detection output,' while deployment uses a different simulated route with no explicit sensor model. This is a genuine generalization and correctness risk, but it is not circularity under the paper's equations, because predicted validity is not defined in terms of the simulated AV's visible set, nor is it forced to equal an input by construction. The model could in principle learn an implicit proxy from the simulated AV's history. No load-bearing argument reduces to a self-citation, no fitted input is renamed as a prediction, and no uniqueness claim is imported from the authors' prior work. Therefore the correct finding is no significant circularity.
Assumptions & free parameters
free parameters (2)
- Validity soft-clipping threshold =
0.5
- Feature normalization constants =
1/80 for x,y,z; mu_l=4.5, mu_w=2.0, mu_h=1.75, mu_k=0.5; sigma_l=2.5, sigma_w=0.8, sigma_h=0.6, sigma_k=0.5
assumptions (4)
- domain assumption The WOMD-XLMap map expansion covers a sufficient road network for 1km-radius trip-level simulation.
- domain assumption Agent validity defined by appearance in the AV's detection output is a sufficient training target for learning occlusion and spawning behavior.
- domain assumption A model trained on 9.1s WOMD clips generalizes to 60-300s rollouts and unseen map regions.
- standard math Standard diffusion v-prediction and cosine schedule are valid for multi-tensor denoising.
Cite this review
Pith. "Pith review of SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model." pith.science (2026). https://pith.science/paper/6GWOMZEH
@misc{pith2026250621976,
author = {Pith},
title = {Pith review of: SceneDiffuser++: City-Scale Traffic Simulation via a Generative World Model},
year = {2026},
howpublished = {\url{https://pith.science/paper/6GWOMZEH}},
note = {Machine review of arXiv:2506.21976}
}
read the original abstract
The goal of traffic simulation is to augment a potentially limited amount of manually-driven miles that is available for testing and validation, with a much larger amount of simulated synthetic miles. The culmination of this vision would be a generative simulated city, where given a map of the city and an autonomous vehicle (AV) software stack, the simulator can seamlessly simulate the trip from point A to point B by populating the city around the AV and controlling all aspects of the scene, from animating the dynamic agents (e.g., vehicles, pedestrians) to controlling the traffic light states. We refer to this vision as CitySim, which requires an agglomeration of simulation technologies: scene generation to populate the initial scene, agent behavior modeling to animate the scene, occlusion reasoning, dynamic scene generation to seamlessly spawn and remove agents, and environment simulation for factors such as traffic lights. While some key technologies have been separately studied in various works, others such as dynamic scene generation and environment simulation have received less attention in the research community. We propose SceneDiffuser++, the first end-to-end generative world model trained on a single loss function capable of point A-to-B simulation on a city scale integrating all the requirements above. We demonstrate the city-scale traffic simulation capability of SceneDiffuser++ and study its superior realism under long simulation conditions. We evaluate the simulation quality on an augmented version of the Waymo Open Motion Dataset (WOMD) with larger map regions to support trip-level simulation.
Figures
Figures from the paper (6 more)
Reference graph
Works this paper leans on
-
[1]
A survey of statistical model checking
Gul Agha and Karl Palmskog. A survey of statistical model checking. ACM Transactions on Modeling and Computer Simulation (TOMACS), 28(1):1–39, 2018. 2
work page 2018
-
[2]
SimNet: Learning reactive self-driving sim- ulations from real-world observations
Luca Bergamini, Yawei Ye, Oliver Scheel, Long Chen, Chih Hu, Luca Del Pero, Bła ˙zej Osi´nski, Hugo Grimmet, and Pe- ter Ondruska. SimNet: Learning reactive self-driving sim- ulations from real-world observations. In ICRA, 2021. 2, 3
work page 2021
-
[3]
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luh- man, Eric Luhman, Clarence Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators
-
[4]
nuScenes: A multi- modal dataset for autonomous driving
Holger Caesar, Varun Bankiti, Alex H Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom. nuScenes: A multi- modal dataset for autonomous driving. In CVPR, 2020. 2
work page 2020
-
[5]
Argo- verse: 3d tracking and forecasting with rich maps
Ming-Fang Chang, John Lambert, Patsorn Sangkloy, Jag- jeet Singh, Slawomir Bak, Andrew Hartnett, De Wang, Peter Carr, Simon Lucey, Deva Ramanan, and James Hays. Argo- verse: 3d tracking and forecasting with rich maps. In CVPR,
-
[6]
Controllable safety- critical closed-loop traffic simulation via guided diffusion,
Wei-Jer Chang, Francesco Pittaluga, Masayoshi Tomizuka, Wei Zhan, and Manmohan Chandraker. Controllable safety- critical closed-loop traffic simulation via guided diffusion,
-
[7]
Safe-sim: Safety-critical closed-loop traffic simulation with diffusion- controllable adversaries
Wei-Jer Chang, Francesco Pittaluga, Masayoshi Tomizuka, Wei Zhan, and Manmohan Chandraker. Safe-sim: Safety-critical closed-loop traffic simulation with diffusion- controllable adversaries. In ECCV, 2024. 3
work page 2024
-
[8]
Sledge: Synthesizing driving environments with generative models and rule-based traffic
Kashyap Chitta, Daniel Dauner, and Andreas Geiger. Sledge: Synthesizing driving environments with generative models and rule-based traffic. In ECCV, 2024. 2
work page 2024
Show all 59 references
-
[9]
Dice: Diverse diffu- sion model with scoring for trajectory prediction, 2023
Younwoo Choi, Ray Coden Mercurius, Soheil Mohamad Al- izadeh Shabestary, and Amir Rasouli. Dice: Diverse diffu- sion model with scoring for trajectory prediction, 2023. 2
2023
-
[10]
A survey of algorithms for black-box safety validation of cyber-physical systems
Anthony Corso, Robert Moss, Mark Koren, Ritchie Lee, and Mykel Kochenderfer. A survey of algorithms for black-box safety validation of cyber-physical systems. Journal of Arti- ficial Intelligence Research, 72:377–428, 2021. 2
2021
-
[11]
CARLA: An open urban driving simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. CARLA: An open urban driving simulator. In CoRL, 2017. 3
2017
-
[12]
Qi, Yin Zhou, Zoey Yang, Aur´elien Chouard, Pei Sun, Jiquan Ngiam, Vijay Vasudevan, Alexander McCauley, Jonathon Shlens, and Dragomir Anguelov
Scott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu, Hang Zhao, Sabeek Pradhan, Yuning Chai, Ben Sapp, Charles R. Qi, Yin Zhou, Zoey Yang, Aur´elien Chouard, Pei Sun, Jiquan Ngiam, Vijay Vasudevan, Alexander McCauley, Jonathon Shlens, and Dragomir Anguelov. Large scale i...
2021
-
[13]
Trafficgen: Learning to generate diverse and realistic traffic scenarios
Lan Feng, Quanyi Li, Zhenghao Peng, Shuhan Tan, and Bolei Zhou. Trafficgen: Learning to generate diverse and realistic traffic scenarios. In ICRA, 2023. 3
2023
-
[14]
Co-Reyes, Rishabh Agarwal, Re- becca Roelofs, Yao Lu, Nico Montali, Paul Mougin, Zoey Yang, Brandyn White, Aleksandra Faust, Rowan McAllister, Dragomir Anguelov, and Benjamin Sapp
Cole Gulino, Justin Fu, Wenjie Luo, George Tucker, Eli Bronstein, Yiren Lu, Jean Harb, Xinlei Pan, Yan Wang, Xiangyu Chen, John D. Co-Reyes, Rishabh Agarwal, Re- becca Roelofs, Yao Lu, Nico Montali, Paul Mougin, Zoey Yang, Brandyn White, Aleksandra Faust, Rowan McAllister, Dra...
-
[15]
SceneDM: Scene-level multi-agent trajectory gen- eration with consistent diffusion models, 2023
Zhiming Guo, Xing Gao, Jianlan Zhou, Xinyu Cai, and Bo- tian Shi. SceneDM: Scene-level multi-agent trajectory gen- eration with consistent diffusion models, 2023. 2
2023
-
[16]
Recurrent world models facilitate policy evolution
David Ha and J ¨urgen Schmidhuber. Recurrent world models facilitate policy evolution. In NeurIPS, 2018. 2
2018
-
[17]
Dream to control: Learning behaviors by la- tent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Moham- mad Norouzi. Dream to control: Learning behaviors by la- tent imagination. In ICLR, 2020
2020
-
[18]
Mastering atari with discrete world models
Danijar Hafner, Timothy P Lillicrap, Mohammad Norouzi, and Jimmy Ba. Mastering atari with discrete world models. In ICLR, 2021. 2
2021
-
[19]
A holistic framework towards vision-based traffic sig- nal control with microscopic simulation
Pan He, Quanyi Li, Xiaoyong Yuan, and Bolei Zhou. A holistic framework towards vision-based traffic sig- nal control with microscopic simulation. arXiv preprint arXiv:2403.06884, 2024. 3
2024
-
[20]
sim- ple diffusion: End-to-end diffusion for high resolution im- ages
Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. sim- ple diffusion: End-to-end diffusion for high resolution im- ages. In ICML. PMLR, 2023. 3, 4
2023
-
[21]
Gaia-1: A generative world model for au- tonomous driving
Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, and Gian- luca Corrado. Gaia-1: A generative world model for au- tonomous driving. arXiv preprint arXiv:2309.17080, 2023. 2
2023 arXiv
-
[22]
Versatile scene- consistent traffic scenario generation as optimization with diffusion, 2024
Zhiyu Huang, Zixu Zhang, Ameya Vaidya, Yuxiao Chen, Chen Lv, and Jaime Fern ´andez Fisac. Versatile scene- consistent traffic scenario generation as optimization with diffusion, 2024. 2, 3
2024
-
[23]
Motion- diffuser: Controllable multi-agent motion prediction using diffusion
Chiyu “Max” Jiang, Andre Cornman, Cheolho Park, Ben- jamin Sapp, Yin Zhou, and Dragomir Anguelov. Motion- diffuser: Controllable multi-agent motion prediction using diffusion. In CVPR, 2023. 2
2023
-
[24]
Scenediffuser: Efficient and controllable driving simulation initialization and rollout
Chiyu Max Jiang, Yijing Bai, Andre Cornman, Christopher Davis, Xiukun Huang, Hong Jeon, Sakshum Kulshrestha, John Lambert, Shuangyu Li, Xuanyu Zhou, Carlos Fuertes, Chang Yuan, Mingxing Tan, Yin Zhou, and Dragomir Anguelov. Scenediffuser: Efficient and controllable driving sim...
2024
-
[25]
Towards learning- based planning:the nuplan benchmark for real-world au- tonomous driving
Napat Karnchanachari, Dimitris Geromichalos, Kok Seang Tan, Nanxiang Li, Christopher Eriksen, Shakiba Yaghoubi, Noushin Mehdipour, Gianmarco Bernasconi, Whye Kit Fong, Yiluan Guo, and Holger Caesar. Towards learning- based planning:the nuplan benchmark for real-world au- tonom...
2024
-
[26]
Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning
Quanyi Li, Zhenghao Peng, Lan Feng, Qihang Zhang, Zhenghai Xue, and Bolei Zhou. Metadrive: Composing diverse driving scenarios for generalizable reinforcement learning. TPAMI, 2022. 3
2022
-
[27]
A deep reinforcement learning network for traffic light cycle control
Xiaoyuan Liang, Xunsheng Du, Guiling Wang, and Zhu Han. A deep reinforcement learning network for traffic light cycle control. IEEE Transactions on Vehicular Technology, 68(2):1243–1253, 2019. 3
2019
-
[28]
Divergence measures based on the shannon en- tropy
Jianhua Lin. Divergence measures based on the shannon en- tropy. IEEE Transactions on Information theory, 37(1):145– 151, 1991. 5
1991
-
[29]
Microscopic traffic simulation using sumo
Pablo Alvarez Lopez, Michael Behrisch, Laura Bieker-Walz, Jakob Erdmann, Yun-Pang Fl¨otter¨od, Robert Hilbrich, Leon- hard L ¨ucken, Johannes Rummel, Peter Wagner, and Eva- marie Wießner. Microscopic traffic simulation using sumo. In 2018 21st international conference on intel...
2018
-
[30]
Scenecontrol: Diffusion for controllable traffic scene generation
Jack Lu, Kelvin Wong, Chris Zhang, Simon Suo, and Raquel Urtasun. Scenecontrol: Diffusion for controllable traffic scene generation. In ICRA, 2024. 2
2024
-
[31]
Unigen: Unified modeling of initial agent states and trajectories for generating autonomous driving scenarios
Reza Mahjourian, Rongbing Mu, Valerii Likhosherstov, Paul Mougin, Xiukun Huang, Joao Messias, and Shimon White- son. Unigen: Unified modeling of initial agent states and trajectories for generating autonomous driving scenarios. In ICRA, 2024. 3
2024
-
[32]
The waymo open sim agents challenge
Nico Montali, John Lambert, Paul Mougin, Alex Kuefler, Nick Rhinehart, Michelle Li, Cole Gulino, Tristan Em- rich, Zoey Yang, Shimon Whiteson, Brandyn White, and Dragomir Anguelov. The waymo open sim agents challenge. In Advances in Neural Information Processing Systems Track ...
2023
-
[33]
Scene transformer: A unified architecture for pre- dicting future trajectories of multiple agents
Jiquan Ngiam, Vijay Vasudevan, Benjamin Caine, Zheng- dong Zhang, Hao-Tien Lewis Chiang, Jeffrey Ling, Re- becca Roelofs, Alex Bewley, Chenxi Liu, Ashish Venugopal, David J Weiss, Benjamin Sapp, Zhifeng Chen, and Jonathon Shlens. Scene transformer: A unified architecture for p...
2022
-
[34]
A diffusion-model of joint interactive nav- igation
Matthew Niedoba, Jonathan Lavington, Yunpeng Liu, Vasileios Lioutas, Justice Sefas, Xiaoxuan Liang, Dylan Green, Setareh Dabiri, Berend Zwartsenberg, Adam Scibior, and Frank Wood. A diffusion-model of joint interactive nav- igation. In NeurIPS, 2023. 2
2023
-
[35]
Improving agent behaviors with rl fine-tuning for autonomous driving
Zhenghao Peng, Wenjie Luo, Yiren Lu, Tianyi Shen, Cole Gulino, Ari Seff, and Justin Fu. Improving agent behaviors with rl fine-tuning for autonomous driving. In ECCV, 2024. 5
2024
-
[36]
Trajeglish: Learning the language of driving scenarios
Jonah Philion, Xue Bin Peng, and Sanja Fidler. Trajeglish: Learning the language of driving scenarios. In ICLR, 2024. 3
2024
-
[37]
Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du
Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih- Yao Ma, Ching-Yao Chuang, David Yan, Dhruv Choudhary, Dingkang Wang, Geet Sethi, Guan Pang, Haoyu Ma, Ishan Misra, Ji Hou, Jialiang Wang, Kiran Jagadeesh, Kunpeng Li, L...
2024
-
[38]
Scenario diffusion: Controllable driving scenario gen- eration with diffusion
Ethan Pronovost, Meghana Reddy Ganesina, Noureldin Hendy, Zeyu Wang, Andres Morales, Kai Wang, and Nick Roy. Scenario diffusion: Controllable driving scenario gen- eration with diffusion. In NeurIPS, 2023. 2, 3
2023
-
[39]
Generating driv- ing scenes with diffusion
Ethan Pronovost, Kai Wang, and Nick Roy. Generating driv- ing scenes with diffusion. In ICRA Workshop on Scalable Autonomous Driving, 2023. 2
2023
-
[40]
Introducing general world models, 2024
Runway ML. Introducing general world models, 2024. Ac- cessed on 2024-10-24. 2
2024
-
[41]
Progressive distillation for fast sampling of diffusion models
Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In ICLR, 2022. 4
2022
-
[42]
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In ICML, 2018. 11
2018
-
[43]
Drivescenegen: Generating diverse and realistic driving sce- narios from scratch
Shuo Sun, Zekai Gu, Tianchen Sun, Jiawei Sun, Chen- gran Yuan, Yuhang Han, Dongen Li, and Marcelo H Ang. Drivescenegen: Generating diverse and realistic driving sce- narios from scratch. IEEE Robotics and Automation Letters,
-
[44]
Trafficsim: Learning to simulate realistic multi- agent behaviors
Simon Suo, Sebastian Regalado, Sergio Casas, and Raquel Urtasun. Trafficsim: Learning to simulate realistic multi- agent behaviors. In CVPR, 2021. 3
2021
-
[45]
Scenegen: Learning to generate realistic traffic scenes
Shuhan Tan, Kelvin Wong, Shenlong Wang, Sivabalan Mani- vasagam, Mengye Ren, and Raquel Urtasun. Scenegen: Learning to generate realistic traffic scenes. In CVPR, 2021. 3
2021
-
[46]
Language conditioned traffic gen- eration
Shuhan Tan, Boris Ivanovic, Xinshuo Weng, Marco Pavone, and Philipp Kr ¨ahenb¨uhl. Language conditioned traffic gen- eration. 7th Annual Conference on Robot Learning (CoRL),
-
[47]
Promptable closed-loop traffic simulation
Shuhan Tan, Boris Ivanovic, Yuxiao Chen, Boyi Li, Xinshuo Weng, Yulong Cao, Philipp Kr¨ahenb¨uhl, and Marco Pavone. Promptable closed-loop traffic simulation. 8th Annual Con- ference on Robot Learning (CoRL), 2024. 3
2024
-
[48]
Srinivasan, Jonathan T
Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Prad- han, Ben Mildenhall, Pratul P. Srinivasan, Jonathan T. Bar- ron, and Henrik Kretzschmar. Block-nerf: Scalable large scene neural view synthesis. In CVPR, 2022. 3
2022
-
[49]
Con- gested traffic states in empirical observations and micro- scopic simulations
Martin Treiber, Ansgar Hennecke, and Dirk Helbing. Con- gested traffic states in empirical observations and micro- scopic simulations. Physical review E , 62(2):1805, 2000. 5
2000
-
[50]
Diffusion models are real-time game engines
Dani Valevski, Yaniv Leviathan, Moab Arar, and Shlomi Fruchter. Diffusion models are real-time game engines. In ICLR, 2025. 2
2025
-
[51]
A survey on traffic signal control methods
Hua Wei, Guanjie Zheng, Vikash Gayah, and Zhenhui Li. A survey on traffic signal control methods. arXiv preprint arXiv:1904.08117, 2019. 3
1904 arXiv
-
[52]
Argoverse 2: Next generation datasets for self-driving perception and fore- casting
Benjamin Wilson, William Qi, Tanmay Agarwal, John Lam- bert, Jagjeet Singh, Siddhesh Khandelwal, Bowen Pan, Rat- nesh Kumar, Andrew Hartnett, Jhony Kaesemodel Pontes, Deva Ramanan, Peter Carr, and James Hays. Argoverse 2: Next generation datasets for self-driving perception an...
2021
-
[53]
Smart: Scalable multi-agent real-time motion generation via next- token prediction
Wei Wu, Xiaoxin Feng, Ziyan Gao, and Yuheng Kan. Smart: Scalable multi-agent real-time motion generation via next- token prediction. In NeurIPS, 2024. 2
2024
-
[54]
Wcdt: World-centric diffusion trans- former for traffic scene generation, 2024
Chen Yang, Aaron Xuxiang Tian, Dong Chen, Tianyu Shi, and Arsalan Heydarian. Wcdt: World-centric diffusion trans- former for traffic scene generation, 2024. 2
2024
-
[55]
Unisim: A neural closed-loop sensor simulator
Ze Yang, Yun Chen, Jingkang Wang, Sivabalan Mani- vasagam, Wei-Chiu Ma, Anqi Joyce Yang, and Raquel Ur- tasun. Unisim: A neural closed-loop sensor simulator. In CVPR, 2023. 3
2023
-
[56]
Learning to drive via asymmetric self-play
Chris Zhang, Sourav Biswas, Kelvin Wong, Kion Fallah, Lunjun Zhang, Dian Chen, Sergio Casas, and Raquel Urta- sun. Learning to drive via asymmetric self-play. In ECCV,
-
[57]
Tedi: Temporally-entangled diffusion for long-term motion synthesis
Zihan Zhang, Richard Liu, Rana Hanocka, and Kfir Aber- man. Tedi: Temporally-entangled diffusion for long-term motion synthesis. In ACM SIGGRAPH 2024 Conference Pa- pers, 2024. 2
2024
-
[58]
Language-guided traffic simulation via scene-level diffusion
Ziyuan Zhong, Davis Rempe, Yuxiao Chen, Boris Ivanovic, Yulong Cao, Danfei Xu, Marco Pavone, and Baishakhi Ray. Language-guided traffic simulation via scene-level diffusion. In CoRL, 2023. 2
2023
-
[59]
trip end
Xiaoyu Zhou, Zhiwei Lin, Xiaojun Shan, Yongtao Wang, Deqing Sun, and Ming-Hsuan Yang. Drivinggaussian: Composite gaussian splatting for surrounding dynamic au- tonomous driving scenes. In CVPR, 2024. 3 A. Appendix A.1. Video We provide a video 2 that features a brief and intui...
2024
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.