Pith. sign in

REVIEW 5 major objections 5 minor 50 references

Jointly denoising future states and low-level controls lets one diffusion model generate controllable, realistic traffic scenarios, and a collision-seeking variant produces challenging interactions for autonomous driving systems.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

T0 review · deepseek-v4-flash

2026-08-01 14:46 UTC pith:NJBEAMVA

load-bearing objection Realism claims fail on their own metrics; the state-action diffusion system is coherent and worth a referee, but not as-is. the 5 major comments →

arxiv 2607.18637 v1 pith:NJBEAMVA submitted 2026-07-21 cs.RO cs.LG

End-to-end Conditional Diffusion for Realistic and Controllable Visual Traffic Scenario Generation

classification cs.RO cs.LG
keywords diffusion modelscenario generationautonomous drivingsafety-critical scenarioscontrollable traffic simulationstate-action diffusionclosed-loop simulationconditional guidance
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The paper tries to establish that the usual trade-off between controllability and realism in learned traffic scenario generation can be avoided by generating, in one diffusion process, both the future motion states of a critical background vehicle and the low-level controls that realize them, conditioned on front-view visual context. It argues that this end-to-end state-action generation removes the planning-control mismatch of trajectory-then-PID pipelines, while differentiable guidance over speed, drivable area, and collision behavior gives a single trained model the ability to produce either naturalistic or safety-critical interactions at inference time. If right, simulation-based evaluation of autonomous driving could switch from hand-crafted scenarios to generated ones that are simultaneously behaviorally plausible and precisely steered toward rare risks.

Core claim

E2E-CDiff's central claim is that a clean-trajectory diffusion model with a decoupled action-state head—predicting controls first, then states from detached control estimates—can jointly denoise executable throttle, steering, and braking commands together with future vehicle states for route-interacting background vehicles, and that gradient guidance applied to the predicted clean trajectory steers generation toward target speeds, drivable-area compliance, and either collision avoidance or collision seeking. The paper reports that this yields a better controllability-realism balance than reinforcement- and imitation-learning baselines across two ego planners, that the collision-guided varian

What carries the argument

The central object is a state-action trajectory: a sequence of 4-dimensional vehicle states (position, heading, speed) and 3-dimensional control actions (throttle, steering, braking) over a fixed horizon, treated as a single object in a diffusion denoising process. The denoiser uses a clean-trajectory parameterization with two heads—an action head that predicts controls and a state head that predicts states from detached action estimates to avoid unstable gradient coupling—and during guided sampling the state head recomputes states from guided actions without detaching, preserving gradients to the control sequence. Guidance objectives (speed deviation, signed-distance drivable-area penalty,

Load-bearing premise

The realism claims rest on treating normality of speed and acceleration distributions, plus closeness to an unspecified target speed distribution, as sufficient evidence of behavioral realism; the paper asserts this normality is supported by empirical traffic studies but cites none, so if real traffic distributions are skewed or multimodal, the measured realism advantage may be an artifact.

What would settle it

Run the paper's realism metrics on real naturalistic driving trajectories: if real speed and acceleration distributions fail the Shapiro-Wilk normality test at typical sample sizes, then high normality scores are not evidence of realism. One could then recompute the paper's comparisons using a direct distributional distance (for example, Wasserstein distance between generated and real speed and acceleration distributions) and check whether the claimed controllability-realism trade-off still favors E2E-CDiff over the baselines.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

If this is right

  • Scenario generation no longer needs a separate trajectory-tracking controller; the model directly outputs executable control signals, so closed-loop execution should align with the predicted states and reduce planning-control mismatch.
  • Users can compose speed, drivable-area, and collision guidance at inference time to generate either naturalistic or adversarial scenarios from a single trained model, without task-specific retraining.
  • Collision-seeking guidance, gated by a trigger distance, can stress-test multiple ego planners with controllable risk while preserving plausible behavior outside the trigger zone.
  • The same diffusion model can serve as an ego planner with competitive driving score and route completion, although its iterative denoising makes it too slow for real-time deployment in its current form.
  • The method is positioned as a closed-loop scenario construction tool rather than a real-time planner; reducing sampling cost is the main obstacle to broader planner use.

Where Pith is reading between the lines

These are editorial extensions of the paper, not claims the author makes directly.

  • The paper's realism evidence rests on statistical proxies; a stronger validation would compare generated state-action distributions directly against real naturalistic driving trajectories, since normality of speed and acceleration is a weak proxy for behavioral plausibility.
  • The approach could naturally extend from single critical vehicle control to jointly generating multiple interacting background vehicles, enabling scene-level guidance over traffic density, cut-in timing, or coordinated adversarial maneuvers.
  • Guidance weights could be tuned automatically through quality-diversity or search-based methods to produce diverse safety-critical scenarios, linking this work to adversarial stress testing more tightly.
  • If latency were reduced by distillation or fewer denoising steps, the same state-action diffusion model might become a viable real-time ego planner, making the controllability guidance available for online driving as well as offline scenario generation.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

5 major / 5 minor

Summary. The paper proposes E2E-CDiff, a conditional diffusion framework that jointly generates future states and low-level control actions for critical background vehicles in closed-loop traffic simulation. Conditioned on a front-view camera image, a diffusion denoiser predicts a state-action rollout, and differentiable guidance objectives for speed, drivable-area compliance, and collision avoidance/seeking steer sampling toward controllable yet naturalistic behavior. Experiments on Bench2Drive with PDM-Lite, PlanT, UniAD, and VAD as ego planners compare against PPO, Pluto, and RIFT, reporting controllability/realism metrics, safety-critical collision rates, ablations, and an ego-planning evaluation.

Significance. The end-to-end state-action diffusion idea and the adaptive guidance mechanism are interesting and potentially useful for scenario generation. If the claims were fully validated, the paper would contribute a flexible method for controllable and safety-critical traffic simulation. However, the current evidence is insufficient: the realism metrics are founded on an unsupported normality assumption and an unspecified target distribution; comparisons omit the most closely related diffusion-based scenario generators; and the controllability results partly reflect the exact quantities being optimized. Because these issues bear directly on the central claim of a "favorable controllability-realism trade-off," substantial revision is required before publication.

major comments (5)
  1. [Section IV-A2] The realism metrics are not valid proxies for behavioral realism. S-SW and A-SW are Shapiro-Wilk normality statistics, yet the manuscript asserts this assumption is "supported by empirical traffic studies" without citing any. Naturalistic speed distributions are often skewed or multimodal and acceleration distributions heavy-tailed, so a high W-statistic does not indicate realism. S-WD is defined against a "target speed distribution" that is never specified, making the metric unfalsifiable and its numerical values uninterpretable. This is load-bearing because the abstract's "favorable controllability-realism trade-off" rests on Tables I and II, where E2E-CDiff is not consistently better than RIFT: ORR is worse (1.26 vs 0.83 in Table I; 5.96 vs 0.38 in Table II), and S-WD and A-SW are worse in Table I.
  2. [Section III-D] The controllability metrics are exactly the quantities optimized by the guidance objectives: TTC is affected by collision guidance, ORR by drivable-area guidance, and speed-related metrics by speed guidance. Thus favorable controllability numbers are partly by construction. The paper does not report how accurately user-specified targets (e.g., v* or δ_margin) are achieved, nor does it show that varying guidance weights produces predictable changes. Similarly, the CPK increase for the collision-seeking variant (Fig. 4) follows directly from negating J_safe_coll in Eq. (12); no DS/RC results for E2E-CDiff w/ coll are reported, so the claim that these collisions actually challenge the ego AD systems is not demonstrated. A downstream evaluation on ego driving quality is needed.
  3. [Section II-A and Section IV-B] The related work lists CTG, SceneControl, Scenario Diffusion, DiffScene, CCDiff, and SceneDiffuser++ as conditional diffusion scenario generators, but the experiments compare only against PPO, Pluto, and RIFT. RIFT is an IL/RL method, not a diffusion-based scene generator. Without direct comparison to at least one of these diffusion baselines, the claimed advantage over "existing learning-based methods" and the specific contribution of the guidance mechanism are not established. Please add experiments against the most relevant diffusion baselines on the same metrics.
  4. [Section III-B] The CBV identification module is underspecified. The "route-level interaction score" is not defined, and the synthesis of the "conflict-aware route" is not described. This component determines which background vehicles are controlled and is essential for reproducibility. Provide the exact scoring function, the route-anchoring procedure, and any threshold or selection criterion.
  5. [Section IV-A2] The paper justifies avoiding trajectory-comparison metrics by stating that "CARLA lacks expert demonstrations," but Bench2Drive training routes include expert driving data used to train the ego planners. Realism could be validated by comparing generated speed/acceleration distributions against those expert logs via Wasserstein distance or density-ratio estimates. Additionally, the test set is only 10 routes, and no significance tests are reported; the mean±std values in Tables I–VI may not support reliable differences.
minor comments (5)
  1. [Section I] Missing citation after "image- or video-based generative models [?]." Please fix the placeholder.
  2. [References] Several reference entries contain inline descriptive annotations (e.g., RIFT, SceneControl, CCDiff), making the bibliography nonstandard. Clean these up.
  3. [Figure 3] The speed/acceleration distribution plots would be more informative if overlayed with a reference distribution from expert/training data, so the reader can visually assess realism.
  4. [Section III-D1c] The trigger gate η(d_ref) in Eq. (12) is not specified. Give its functional form (hard or smooth) and the threshold d_th, as well as the chosen values of κ, λ, σ, and δ_margin.
  5. [Section IV-E] The term "end-to-end" is used loosely: the method takes a front-view image and outputs controls, but it does not include full perception training typical of end-to-end driving. Clarify the intended meaning.

Circularity Check

2 steps flagged

Controllability metrics are the guidance objectives, and the collision-guided CPK increase is literally the negative of the collision-avoidance loss; realism proxies are independently unsupported.

specific steps
  1. self definitional [Sec. IV-A2a (Controllability Evaluation); Sec. III-D1b Eq. (10); Sec. III-D1a Eq. (9); Sec. III-D1c Eq. (11)]
    "To assess controllability, we focus on parameters directly affected by our guidance objectives, including relative speed and time-to-collision (TTC) between the ego and the CBV . ... Off-Road Rate (ORR): the percentage of time CBVs spend off-road on average."

    The controllability evaluation is explicitly over 'parameters directly affected by our guidance objectives.' Eq. (10) minimizes a penalty that grows when the CBV leaves the drivable area, while ORR measures the fraction of time off-road; Eq. (11) penalizes low ego-CBV separation, while 2D-TTC measures closeness; Eq. (9) drives CBV speed toward v*, while S-WD measures speed distributions. The favorable controllability numbers in Tables I/II/V therefore show that the denoising process optimized its own losses, not that an independent metric of controllability was satisfied. This is partial circularity: the comparisons to PPO/Pluto/RIFT and the diffusion prior are still independent content.

  2. self definitional [Sec. III-D1c, Eq. (12); Sec. IV-A2c (CPK); Fig. 4]
    "The risk-seeking loss is: J^risk_coll(τ0) = −η(d_ref) J^safe_coll(τ0).(12) Minimizing J^risk_coll encourages closer interaction only in the triggered distance range, which helps generate safety-critical scenarios."

    Eq. (12) defines the collision-seeking objective as the negative of the collision-avoidance loss, i.e., an explicit instruction to reduce ego-CBV separation. CPK counts collisions per kilometer. The reported result that 'adding collision-oriented guidance consistently increases the CPK' (e.g., 5.68 to 20.77 for PDM-Lite) is therefore the expected effect of the optimizer following its own loss, not an independent finding about scenario challenge. The claim that the 'collision-guided variant induces challenging interactions across multiple autonomous driving systems' is thus a restatement of the objective rather than a prediction validated by external evidence.

full rationale

The paper is not a self-citation loop: the diffusion training loss Eq. (7) is a standard DDPM objective, the end-to-end state-action representation is a genuine architectural choice, and the ego-planner results in Table VI are externally meaningful. However, two load-bearing evaluation claims reduce by construction. First, the controllability metrics are explicitly selected as 'parameters directly affected by our guidance objectives'; ORR is the exact quantity minimized by Eq. (10), and 2D-TTC/S-WD are closely aligned with the collision and speed objectives. Second, the safety-critical CPK increase of E2E-CDiff w/ coll is the direct result of Eq. (12) being the negative of the collision-avoidance loss. The realism evaluation also lacks support: the paper says the normality assumption is 'supported by empirical traffic studies' but cites none, and the S-WD target speed distribution is never specified. Those are missing-validity problems rather than circularity, because no equation-level reduction can be exhibited for them. The comparability of baselines (PPO, Pluto, RIFT) provides some independent content, so this is partial circularity, not a fully tautological derivation.

Axiom & Free-Parameter Ledger

4 free parameters · 6 axioms · 0 invented entities

The central claims rest on standard diffusion math plus several domain assumptions about simulation fidelity and realism metrics. The free parameters are mostly learned model weights and unreported guidance weights that directly set the reported trade-off.

free parameters (4)
  • denoiser parameters θ = trained on 220 Bench2Drive routes; no checkpoint released
    The diffusion model's action/state heads are fit to training trajectories; all downstream claims depend on this learned mapping.
  • guidance weights w_speed, w_area, w_coll = not reported
    Eqs. (13)-(14) combine guidance objectives with scalar weights; the reported controllability-realism trade-off depends on their values, but they are not given in the paper.
  • guidance shaping constants κ, λ, σ, δ_margin, d_th = not reported
    Eqs. (10)-(12) include safety margin, softplus scale, exponential shaping, and trigger range that determine collision-seeking behavior; values are not stated.
  • target speed v* = user-specified per scenario
    Speed guidance (Eq. 9) penalizes deviation from v*, which sets the speed controllability metric; chosen by user rather than learned.
axioms (6)
  • standard math DDPM forward/reverse process and clean-trajectory parameterization are valid for state-action trajectories
    Eqs. (3)-(7) follow Ho et al. [40]; assumed standard.
  • domain assumption Front-view camera image provides sufficient context to predict the selected CBV's future state-action rollout
    The entire conditioning c is a vehicle-centric front-view image (Sec. III-E); no ablation on context modality or evidence that interaction-relevant information is fully visible.
  • domain assumption Normality of speed/acceleration distributions is a valid realism proxy
    Sec. IV-A2 uses Shapiro-Wilk tests as realism metrics; no supporting traffic-study citation is provided, and natural traffic distributions are not necessarily normal.
  • domain assumption Bench2Drive/CARLA closed-loop simulation is a faithful enough proxy for real-world traffic behavior to support realism claims
    All evaluations are in simulation; transfer to real-world naturalistic driving is assumed rather than demonstrated.
  • ad hoc to paper The unspecified route-level interaction score selects exactly the critical background vehicles
    Sec. III-B defines the score only verbally; its correctness is load-bearing for intervention on a small subset of agents.
  • ad hoc to paper Collision-seeking trigger gate η(d_ref) avoids globally aggressive behavior
    Eq. (12) gates risk-seeking by current distance; this design choice is not derived from data or theory.

pith-pipeline@v1.3.0-alltime-deepseek · 15344 in / 15443 out tokens · 136711 ms · 2026-08-01T14:46:29.087888+00:00 · methodology

0 comments
read the original abstract

Generating closed-loop traffic scenarios that are both realistic and controllable is crucial for evaluating autonomous driving systems, especially under rare safety-critical interactions. Existing learning-based methods often struggle to balance controllability and realism, offering either limited fine-grained control over traffic behavior or controllable scenarios at the expense of behavioral plausibility. This paper presents E2E-CDiff, an end-to-end conditional diffusion framework for controllable and realistic scenario generation. Conditioned on front-view visual observations, E2E-CDiff jointly denoises future motion states and executable low-level controls for route-interacting background vehicles. This unified state-action generation mitigates the planning-control mismatch in conventional two-stage trajectory-then-controller pipelines. Differentiable guidance further regulates speed, enforces drivable-area compliance, and supports collision-avoidance or collision-seeking behaviors, enabling both naturalistic and safety-critical scenario generation. Experiments on Bench2Drive show that E2E-CDiff achieves a favorable controllability-realism trade-off compared with representative reinforcement- and imitation-learning baselines, while its collision-guided variant induces challenging interactions across multiple autonomous driving systems. E2E-CDiff also performs competitively as a learning-based ego planner, demonstrating the generality of end-to-end state-action diffusion.

Figures

Figures reproduced from arXiv: 2607.18637 by Baochang Zhang, Bing Li, Binhang Qi, Jingzheng Li, Keyu Chen, Philip S Yu, Qianren Mao, Xianglong Liu, Yufei Ge, Zhijun Chen, Zizhe Wang.

Figure 1
Figure 1. Figure 1: Comparison between prior two-stage control and our end-to-end conditional diffusion planner. (a) Prior work predicts trajectories [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Conditional guidance in reverse diffusion. At each denoising [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Speed and acceleration distribution of CBVs. [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figure 4
Figure 4. Figure 4: CPK comparison between E2E-CDiff and E2E-CDiff w/ coll. [PITH_FULL_IMAGE:figures/full_fig_p006_4.png] view at source ↗
Figure 5
Figure 5. Figure 5: Qualitative comparison across CBV planners. For each planner, three temporal frames are shown. In each frame, the CBV is marked [PITH_FULL_IMAGE:figures/full_fig_p007_5.png] view at source ↗
Figure 6
Figure 6. Figure 6: Qualitative results of safety-critical scenario generation between E2E-CDiff and E2E-CDiff w/ coll. [PITH_FULL_IMAGE:figures/full_fig_p007_6.png] view at source ↗
Figure 8
Figure 8. Figure 8: Latency comparison. control signals with the end-to-end conditional diffusion model by comparing it with a two-stage trajectory-then-PID pipeline [PITH_FULL_IMAGE:figures/full_fig_p008_8.png] view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

50 extracted references · 8 linked inside Pith

  1. [1]

    Intelligent driving in- telligence test for autonomous vehicles with naturalistic and adversarial environment,

    S. Feng, X. Yan, H. Sun, Y . Feng, and H. X. Liu, “Intelligent driving in- telligence test for autonomous vehicles with naturalistic and adversarial environment,” Nature communications, vol. 12, no. 1, p. 748, 2021

  2. [2]

    Safebench: A benchmarking platform for safety evaluation of autonomous vehicles,

    C. Xu, W. Ding, W. Lyu, Z. Liu, S. Wang, Y . He, H. Hu, D. Zhao, and B. Li, “Safebench: A benchmarking platform for safety evaluation of autonomous vehicles,” Advances in Neural Information Processing Systems, vol. 35, pp. 25 667–25 682, 2022

  3. [3]

    Coda: A real-world road corner case dataset for object detection in autonomous driving,

    K. Li, K. Chen, H. Wang, L. Hong, C. Ye, J. Han, Y . Chen, W. Zhang, C. Xu, D.-Y . Yeunget al., “Coda: A real-world road corner case dataset for object detection in autonomous driving,” in European Conference on Computer Vision. Springer, 2022, pp. 406–423

  4. [4]

    Bench2drive: Towards multi-ability benchmarking of closed-loop end-to-end autonomous driv- ing,

    X. Jia, Z. Yang, Q. Li, Z. Zhang, and J. Yan, “Bench2drive: Towards multi-ability benchmarking of closed-loop end-to-end autonomous driv- ing,” in The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2024

  5. [5]

    Rift: Closed-loop rl fine-tuning for realistic and controllable traffic simulation,

    K. Chen, W. Sun, H. Cheng, and S. Zheng, “Rift: Closed-loop rl fine-tuning for realistic and controllable traffic simulation,” arXiv preprint arXiv:2505.03344, 2025, dual-stage imitation pre-training + closed-loop RL fine-tuning for multimodal, adaptive agents. [Online]. Available: https://arxiv.org/abs/2505.03344

  6. [6]

    Foundation models in autonomous driving: A survey on scenario generation and scenario analysis,

    Y . Gao, M. Piccinini, Y . Zhang, D. Wang, K. Moller, R. Brusnicki, B. Zarrouki, A. Gambi, J. F. Totz, K. Storms et al., “Foundation models in autonomous driving: A survey on scenario generation and scenario analysis,” arXiv preprint arXiv:2506.11526, 2025

  7. [7]

    Chatscene: Knowledge-enabled safety- critical scenario generation for autonomous vehicles,

    J. Zhang, C. Xu, and B. Li, “Chatscene: Knowledge-enabled safety- critical scenario generation for autonomous vehicles,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 15 459–15 469

  8. [8]

    Talk2traffic: In- teractive and editable traffic scenario generation for autonomous driving with multimodal large language model,

    Z. Sheng, Z. Huang, Y . Qu, Y . Leng, and S. Chen, “Talk2traffic: In- teractive and editable traffic scenario generation for autonomous driving with multimodal large language model,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 3788–3797

  9. [9]

    King: Generating safety-critical driving scenarios for robust imitation via kinematics gradients,

    N. Hanselmann, K. Renz, K. Chitta, A. Bhattacharyya, and A. Geiger, “King: Generating safety-critical driving scenarios for robust imitation via kinematics gradients,” in European Conference on Computer Vision. Springer, 2022, pp. 335–352

  10. [10]

    Adversarial safety-critical scenario generation using naturalistic human driving priors,

    K. Hao, W. Cui, Y . Luo, L. Xie, Y . Bai, J. Yang, S. Yan, Y . Pan, and Z. Yang, “Adversarial safety-critical scenario generation using naturalistic human driving priors,” IEEE Transactions on Intelligent Vehicles, 2023

  11. [11]

    Frea: Feasibility-guided generation of safety-critical scenarios with reasonable adversariality,

    K. Chen, Y . Lei, H. Cheng, H. Wu, W. Sun, and S. Zheng, “Frea: Feasibility-guided generation of safety-critical scenarios with reasonable adversariality,” arXiv preprint arXiv:2406.02983, 2024

  12. [12]

    Generating useful accident-prone driving scenarios via a learned traffic prior,

    D. Rempe, J. Philion, L. J. Guibas, S. Fidler, and O. Litany, “Generating useful accident-prone driving scenarios via a learned traffic prior,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 17 305–17 315

  13. [13]

    Generating traffic scenarios via in-context learning to learn better motion planner,

    A. Aiersilan, “Generating traffic scenarios via in-context learning to learn better motion planner,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 14, 2025, pp. 14 539–14 547

  14. [14]

    Drivescenegen: Generating diverse and realistic driving scenarios from scratch,

    S. Sun, Z. Gu, T. Sun, J. Sun, C. Yuan, Y . Han, D. Li, and M. H. Ang, “Drivescenegen: Generating diverse and realistic driving scenarios from scratch,” IEEE Robotics and Automation Letters, vol. 9, no. 8, pp. 7007–7014, 2024

  15. [15]

    Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving,

    B. Liao, S. Chen, H. Yin, B. Jiang, C. Wang, S. Yan, X. Zhang, X. Li, Y . Zhang, Q. Zhang et al., “Diffusiondrive: Truncated diffusion model for end-to-end autonomous driving,” in Proceedings of the Computer Vision and Pattern Recognition Conference, 2025, pp. 12 037–12 047

  16. [16]

    Diffscene: Diffusion- based safety-critical scenario generation for autonomous vehicles,

    C. Xu, A. Petiushko, D. Zhao, and B. Li, “Diffscene: Diffusion- based safety-critical scenario generation for autonomous vehicles,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 8, 2025, pp. 8797–8805

  17. [17]

    Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong baseline,

    P. Wu, X. Jia, L. Chen, J. Yan, H. Li, and Y . Qiao, “Trajectory-guided control prediction for end-to-end autonomous driving: A simple yet strong baseline,” Advances in Neural Information Processing Systems, vol. 35, pp. 6119–6132, 2022

  18. [18]

    Guided conditional diffusion for controllable traffic simulation,

    Z. Zhong, D. Rempe, D. Xu, Y . Chen, S. Veer, T. Che, B. Ray, and M. Pavone, “Guided conditional diffusion for controllable traffic simulation,” in IEEE International Conference on Robotics and Automation (ICRA), 2023, cTG: Controllable Traffic Generation. [Online]. Available: https://research.nvidia.com/labs/avg/ publication/zhong.rempe.etal.icra2023/

  19. [19]

    Scenecontrol: Diffusion for controllable traffic scene generation,

    J. Lu, K. Wong, C. Zhang, S. Suo, and R. Urtasun, “Scenecontrol: Diffusion for controllable traffic scene generation,” in Technical report / preprint (Waabi Research), 2024, pp. 1–7, diffusion model + guided sampling to satisfy arbitrary high-level constraints; interactive scene generation. [Online]. Available: https://research-assets.waabi.ai/ SceneContr...

  20. [20]

    Scenario diffusion: Controllable driving scenario generation with diffusion,

    E. Pronovost, M. R. Ganesina, N. Hendy, Z. Wang, A. Morales, K. Wang, and N. Roy, “Scenario diffusion: Controllable driving scenario generation with diffusion,” in Advances in Neural Information Processing Systems (NeurIPS), 2023. [Online]. Available: https: //arxiv.org/abs/2311.02738

  21. [21]

    Realgen: Retrieval augmented generation for controllable traffic scenarios,

    W. Ding, Y . Cao, D. Zhao, C. Xiao, and M. Pavone, “Realgen: Retrieval augmented generation for controllable traffic scenarios,” in European Conference on Computer Vision (ECCV) Workshops / arXiv preprint, 2024, retrieval-augmented in-context composition for controllable scenario synthesis. [Online]. Available: https://arxiv.org/abs/2312.13303

  22. [22]

    Chat2scenario: Scenario extraction from dataset through utilization of large language model,

    Y . Zhao, W. Xiao, T. Mihalj, J. Hu, and A. Eichberger, “Chat2scenario: Scenario extraction from dataset through utilization of large language model,” in 2024 IEEE Intelligent Vehicles Symposium (IV). IEEE, 2024, pp. 559–566

  23. [23]

    Cadre: Controllable and diverse generation of safety-critical driving scenarios using real-world trajectories,

    P. Huang, W. Ding, B. Stoler, J. Francis, B. Chen, and D. Zhao, “Cadre: Controllable and diverse generation of safety-critical driving scenarios using real-world trajectories,” arXiv preprint, 2024, quality-diversity (QD) + archive optimization for controllable, diverse safety-critical scenarios. [Online]. Available: https://arxiv.org/abs/2403.13208

  24. [24]

    Adaptive stress testing for autonomous vehicles,

    M. Koren and M. J. Kochenderfer, “Adaptive stress testing for autonomous vehicles,” in IEEE Intelligent Vehicles Symposium (IV), 2021, pp. 682–689. [Online]. Available: https://ieeexplore.ieee.org/ document/9575638

  25. [25]

    Deepscenario: Efficient search-based generation of safety-critical driving scenarios,

    J. Huang, Y . Qi, Z. Huang, W. Ding, and D. Zhao, “Deepscenario: Efficient search-based generation of safety-critical driving scenarios,” in IEEE International Conference on Intelligent Transportation Systems (ITSC), 2022, pp. 3514–3521. [Online]. Available: https: //arxiv.org/abs/2208.13258

  26. [26]

    Waymax: An accelerated, data-driven simulator for large-scale autonomous driving research,

    C. Gulino et al., “Waymax: An accelerated, data-driven simulator for large-scale autonomous driving research,” in NeurIPS, 2024, data-driven simulation platform referenced in the review

  27. [27]

    nuplan: A closed-loop planning benchmark for autonomous vehicles,

    H. Caesar, V . Bankiti, A. H. Lang, S. V ora, V . Narayanan, A. Krishnan, and O. Beijbom, “nuplan: A closed-loop planning benchmark for autonomous vehicles,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2021. [Online]. Available: https://www.nuscenes.org/nuplan

  28. [28]

    Scenegen: Learning to generate realistic traffic scenes,

    S. Tan, D. Jayaraman, A. Gupta, and B. Zhou, “Scenegen: Learning to generate realistic traffic scenes,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, pp. 14 185–14 195. [Online]. Available: https://openaccess. thecvf.com/content/CVPR2021/papers/Tan SceneGen Learning To Generate Realistic Traffic Scenes ...

  29. [29]

    Trafficgen: Learn- ing to generate diverse and realistic traffic scenarios,

    L. Feng, Q. Li, Z. Peng, S. Tan, and B. Zhou, “Trafficgen: Learn- ing to generate diverse and realistic traffic scenarios,” in 2023 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2023, pp. 3567–3575

  30. [30]

    Scenediffuser++: City-scale traffic simulation via a generative world model,

    S. Tan and et al., “Scenediffuser++: City-scale traffic simulation via a generative world model,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, city-scale generative world model for long-horizon closed-loop simulation. [Online]. Available: https://openaccess.thecvf.com/content/ CVPR2025/html/Tan SceneDi...

  31. [31]

    Ccdiff: Causal composition diffusion model for closed-loop traffic generation,

    H. Lin and et al., “Ccdiff: Causal composition diffusion model for closed-loop traffic generation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2025, causal composition + diffusion guidance to improve controllability and long-horizon realism. [Online]. Available: https://openaccess.thecvf. com/content/CVPR2025...

  32. [32]

    Ctrl-sim: Reactive and controllable driving agents with offline reinforcement learning,

    L. Rowe, R. Girgis, A. Gosselin, B. Carrez, F. Golemo, F. Heide, L. Paull, and C. Pal, “Ctrl-sim: Reactive and controllable driving agents with offline reinforcement learning,” in arXiv preprint arXiv:2403.19918, 2024, return-conditioned multi-agent behaviour models via offline RL. [Online]. Available: https://arxiv.org/abs/2403. 19918

  33. [33]

    Sce- narionet: Open-source platform for large-scale traffic scenario simulation and modeling,

    Q. Li, Z. M. Peng, L. Feng, Z. Liu, C. Duan, W. Mo, and B. Zhou, “Sce- narionet: Open-source platform for large-scale traffic scenario simulation and modeling,” Advances in neural information processing systems, vol. 36, pp. 3894–3920, 2023

  34. [34]

    Editable scene simulation for autonomous driving via collaborative llm-agents,

    Y . Wei, Z. Wang, Y . Lu, C. Xu, C. Liu, H. Zhao, S. Chen, and Y . Wang, “Editable scene simulation for autonomous driving via collaborative llm-agents,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2024, pp. 15 077–15 087. JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2015 11

  35. [35]

    A survey on safety-critical driving scenario generation—a methodological perspec- tive,

    W. Ding, C. Xu, M. Arief, H. Lin, B. Li, and D. Zhao, “A survey on safety-critical driving scenario generation—a methodological perspec- tive,” IEEE Transactions on Intelligent Transportation Systems, vol. 24, no. 7, pp. 6971–6988, 2023

  36. [36]

    Waymo simulated driving behavior in reconstructed fatal crashes within an autonomous vehicle operating domain,

    J. M. Scanlon, K. D. Kusano, T. Daniel, C. Alderson, A. Ogle, and T. Victor, “Waymo simulated driving behavior in reconstructed fatal crashes within an autonomous vehicle operating domain,” Accident Analysis & Prevention, vol. 163, p. 106454, 2021

  37. [37]

    Adversarial evaluation of autonomous vehicles in lane-change scenarios,

    B. Chen, X. Chen, Q. Wu, and L. Li, “Adversarial evaluation of autonomous vehicles in lane-change scenarios,” IEEE transactions on intelligent transportation systems, vol. 23, no. 8, pp. 10 333–10 342, 2021

  38. [38]

    Cat: Closed-loop adversarial training for safe end-to-end driving,

    L. Zhang, Z. Peng, Q. Li, and B. Zhou, “Cat: Closed-loop adversarial training for safe end-to-end driving,” in Conference on Robot Learning. PMLR, 2023, pp. 2357–2372

  39. [39]

    Summit: A simulator for urban driving in massive mixed traffic,

    P. Cai, Y . Lee, Y . Luo, and D. Hsu, “Summit: A simulator for urban driving in massive mixed traffic,” in2020 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2020, pp. 4023–4029

  40. [40]

    Denoising diffusion probabilistic models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising diffusion probabilistic models,” Advances in Neural Information Processing Systems (NeurIPS), vol. 33, pp. 6840–6851, 2020

  41. [41]

    Planning with diffusion for flexible behavior synthesis,

    M. Janner, Y . Du, J. B. Tenenbaum, and S. Levine, “Planning with diffusion for flexible behavior synthesis,” in Proceedings of the 39th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, K. Chaudhuri, S. Jegelka, L. Song, C. Szepesv ´ari, G. Niu, and S. Sabato, Eds., vol. 162. PMLR, 2022, pp. 9902–9915

  42. [42]

    PDM-Lite: A rule-based planner for carla leaderboard 2.0,

    J. Beißwenger, “PDM-Lite: A rule-based planner for carla leaderboard 2.0,” https://github.com/OpenDriveLab/DriveLM/blob/ DriveLM-CARLA/docs/report.pdf, 2024

  43. [43]

    Plant: Explainable planning transformers via object-level representa- tions,

    K. Renz, K. Chitta, O.-B. Mercea, A. Koepke, Z. Akata, and A. Geiger, “Plant: Explainable planning transformers via object-level representa- tions,” in 6th Annual Conference on Robot Learning. MLResearch- Press, 2022, pp. 459–470

  44. [44]

    Planning-oriented autonomous driving,

    Y . Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wang et al., “Planning-oriented autonomous driving,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2023, pp. 17 853–17 862

  45. [45]

    Vad: Vectorized scene representation for efficient autonomous driving,

    B. Jiang, S. Chen, Q. Xu, B. Liao, J. Chen, H. Zhou, Q. Zhang, W. Liu, C. Huang, and X. Wang, “Vad: Vectorized scene representation for efficient autonomous driving,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2023, pp. 8340–8350

  46. [46]

    Pluto: Pushing the limit of imita- tion learning-based planning for autonomous driving,

    J. Cheng, Y . Chen, and Q. Chen, “Pluto: Pushing the limit of imita- tion learning-based planning for autonomous driving,” arXiv preprint arXiv:2404.14327, 2024

  47. [47]

    Prox- imal policy optimization algorithms,

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Prox- imal policy optimization algorithms,” arXiv preprint arXiv:1707.06347, 2017

  48. [48]

    trajdata: A unified interface to multiple human trajectory datasets,

    B. Ivanovic, G. Song, I. Gilitschenski, and M. Pavone, “trajdata: A unified interface to multiple human trajectory datasets,” Advances in Neural Information Processing Systems, vol. 36, pp. 27 582–27 593, 2023

  49. [49]

    An analysis of variance test for normality,

    S. Shaphiro, M. Wilk et al., “An analysis of variance test for normality,” Biometrika, vol. 52, no. 3, pp. 591–611, 1965

  50. [50]

    Markov processes over denumerable products of spaces, describing large systems of automata,

    L. N. Vaserstein, “Markov processes over denumerable products of spaces, describing large systems of automata,” Problemy Peredachi Informatsii, vol. 5, no. 3, pp. 64–72, 1969