Pith. sign in

REVIEW 4 major objections 4 minor 47 references

Integrating Decision-Making Into Differentiable Optimization Guided Learning for End-to-End Planning of Autonomous Vehicles

T0 review · 4 major / 4 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read This paper claims that integrating lane-selection decisions into a differentiable optimization-guided learning framework produces safer and more efficient autonomous driving plans than imitation learning baselines.

desk verdict A genuine DIPP extension with promising closed-loop results, but the central claim about preserving one-hot decision constraints is not supported by the paper's own penalty equations. read the letter →

arxiv 2412.01234 v1 pith:YLEZILDB submitted 2024-12-02 cs.RO

classification cs.RO
keywords autonomousdrivingend-to-endplanningdifferentiableoptimizationdecision-makingtrajectorylaneselectionimitationlearningmotionprediction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper aims to show that autonomous vehicles can be trained end-to-end to make lane-selection decisions and plan trajectories by optimizing explicit safety, efficiency, and comfort objectives rather than merely imitating human demonstrations. The authors build a differentiable optimizer that consumes predicted futures from a transformer-based predictor and jointly solves for the best lane choice and the best trajectory; the whole pipeline is trainable by backpropagation through the optimizer. If the central claim is right, end-to-end planning systems can go beyond expert trajectories, perform discretionary lane changes that avoid obstacles, improve progress, and retain interpretability through the optimization objectives. The reported results show a 4.50% open-loop collision rate versus 7.55% for the DIPP baseline and 72.77 m closed-loop progress versus 47.58 m for DIPP.

What carries the argument

The load-bearing machinery is a differentiable constrained nonlinear optimization problem solved with the Gauss-Newton algorithm inside a bilevel optimization loop: the inner loop solves for the ego vehicle's lane-selection decision variables $b_\alpha(\tau) \in \{0,1\}$ (relaxed to continuous), states $x(\tau)$, and controls $u(\tau)$; the outer loop trains the transformer-based predictor, whose outputs initialize the decision and trajectory for the inner loop. The cost function combines position tracking, decision-dependent safety costs relative to leading and neighboring vehicles, traveling efficiency, comfort, and hinge-loss penalties for collisions and traffic-light violations. The discrete nature of the decision is handled by relaxing $b_\alpha$ and adding penalty terms $\ell_{\mathrm{binary}}$ and $\ell_{\mathrm{equality}}$ with large weights, with the claimed effect of preserving the integer and equality constraints throughout learning. The entire pipeline is trained end-to-end, with the cost weights themselves also learnable, using a combination of prediction, score, decision, planning, and imitation losses.

What would settle it

Run the trained optimizer on a diverse set of held-out scenes and record the optimized decision variables $b_\alpha$: the claim that the formulation enforces a single lane choice is falsified if any converged solution mixes two lanes with strictly fractional values while the penalty terms are zero, or if removing the penalty weights $w_{\mathrm{bi}}$ and $w_{\mathrm{eq}}$ leaves all optimized decisions unchanged.

Watch

Extended reading notes

Core claim

The paper's central discovery is that lane-selection decisions and trajectory plans can be jointly optimized in a differentiable constrained nonlinear program, and that training a transformer predictor end-to-end with this optimizer yields driving behavior that is safer, more efficient, and more comfortable than imitation-based planning. The optimization objectives explicitly encode tracking, longitudinal and lateral safety, traveling efficiency, riding comfort, and collision/traffic-light compliance, with the decision variable relaxed from binary to continuous but accompanied by penalty terms intended to enforce the one-hot and equality constraints. On the Waymo Open Motion Dataset, open-loop testing gives the method the lowest collision rate (4.50%) among the compared methods, and closed-loop testing gives the lowest collision rate (4%) and the largest progress (72.77 m) against baselines including DIPP. The paper also claims that the learned initialization of decisions and actions from the prediction module is essential for the optimizer to converge and for overall performance, as demonstrated by ablations.

Load-bearing premise

The argument assumes that the penalty terms actually force the planner to settle on exactly one lane; in fact those penalties are zero for any fractional combination of lanes that does not exceed one lane in total, so the guarantee really comes from the optimizer happening to land on clean choices, not from the math of the penalties.

Editorial extensions

If this is right

  • End-to-end planners can be built without a predefined reference route, with lane selection emerging from the optimization objectives instead of from HD-map waypoints.
  • Joint training of prediction, decision-making, and trajectory planning through a differentiable optimizer can reduce closed-loop collision rates and increase travel progress relative to imitation-only baselines.
  • Providing learned initialization for both the decision and the control inputs is a critical design choice: the ablation shows convergence drops to 15% without decision initialization.
  • Because the cost weights are learnable, the trade-off between safety, efficiency, and comfort can be tuned from data during end-to-end training.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper's reported 100% constraint compliance is likely not guaranteed by the penalty terms alone: for any continuous $b_\alpha \in [0,1]$ with $\sum_\alpha b_\alpha \le 1$, both $\ell_{\mathrm{binary}}$ and $\ell_{\mathrm{equality}}$ are zero, so the one-hot property depends on the optimizer and learned initialization landing on near-binary values rather than on the formulation.
  • An immediate testable extension is to inspect the optimized decision variables across many held-out scenes; if any converged solution mixes two lanes with comparable weights at zero penalty, the hard-constraint claim would be weakened.
  • The higher open-loop planning error at 3 s and 5 s relative to DIPP suggests the gains are concentrated in closed-loop re-planning; in settings where the planner is not re-run frequently, the benefit may be smaller.
  • The closed-loop evaluation replays recorded agent trajectories (non-reactive), so the collision-rate improvement may not fully reflect performance in fully interactive traffic where other vehicles respond to the ego vehicle's decisions.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. The paper proposes an end-to-end trainable framework for autonomous vehicle planning that couples a transformer-based motion predictor with a differentiable nonlinear optimizer. The optimizer jointly solves lane-selection decisions and trajectory planning using costs for safety, traveling efficiency, riding comfort, and penalties intended to enforce the integer and equality constraints of the decision problem. The framework is trained on the Waymo Open Motion Dataset and evaluated in open-loop and closed-loop settings against Vanilla IL, IL+OPT, and DIPP. The authors report the lowest collision rate (4.50% open-loop, 4% closed-loop) and the largest closed-loop progress (72.77 m) among the compared methods, and they argue that the integrated decision-making capability enables planning that goes beyond imitation of expert demonstrations. Ablation studies examine the role of learned initialization and learnable cost weights.

Significance. If the results are reproducible and the decision-constraint issue is resolved, the integration of discrete lane-selection decisions into differentiable optimization-guided learning is a useful step toward end-to-end planning that goes beyond imitation. The closed-loop progress improvement over DIPP (72.77 m vs 47.58 m) and the low collision rates are notable, and the closed-loop log-replay evaluation is appropriate for the claimed driving-performance gains. The ablation study for decision initialization is also informative. However, the paper's central structural claim that the decision-making constraints are 'preserved throughout the learning process' is not supported by the penalty formulation as written, and the open-loop 'consistent improvement' claim is contradicted by parts of Table II. These issues must be addressed before the results can be taken at face value.

major comments (4)
  1. [IV-A6, Eqs. (20)-(21), Table V] The penalty terms do not enforce the claimed decision constraints. For any continuous b in [0,1], b(b-1) <= 0, so max(0, b(b-1)) is identically zero; for any b with b_-1 + b_0 + b_1 <= 1, max(0, sum(b)-1) is also zero. The relaxed feasible set therefore contains all fractional mixtures with sum at most 1, including b=(0,0,0), and neither penalty creates any gradient toward integrality. The statement in Section IV-A6 that constraints (3d) and (3c) are 'still respected' is not justified by these equations. Table V's 100% compliance must come from the learned initialization, a rounding/thresholding evaluation, or some other mechanism, not from the formulation as written. This is load-bearing because the paper's distinguishing contribution is that discrete decision-making constraints are preserved through learning. Please either replace the penalties with a form that actually penalizes fractional values (e.g., a penalty on b(1-b) plus a two-sided equality penalty), or demonstrate that the optimizer's solutions are binary by construction, and report the compliance metric accordingly.
  2. [Abstract and Section V-B, Table II] The abstract's claim that open-loop outcomes 'consistently outperform' baselines is contradicted by Table II. At 3s and 5s the planning error is larger than DIPP (2.105 vs 1.715 m and 4.763 vs 4.630 m), and the off-route rate is higher (8.85% vs 7.68%). The statement in Section V-B that the method 'consistently outperforming both Vanilla IL and IL+OPT, while closely matching DIPP' is more accurate. Please revise the abstract and the summary of open-loop results to reflect the mixed outcome on planning accuracy, or supply an explicit weighting/aggregation that justifies the 'consistent' claim.
  3. [Section V, first paragraph; Section V-B] The DIPP comparison may be unfair. The authors state that a notable modification in their processed data is the absence of a reference route, while DIPP [15] is a framework that plans with respect to a preconfigured reference route. If DIPP is evaluated without the reference route it was designed to use, its off-route and control-smoothness numbers in Table II and Fig. 3 could be degraded. Please clarify how the DIPP baseline obtains its reference line under the shared data pipeline, or include an additional DIPP variant with a reference route.
  4. [Section V, Tables II and III] All results are reported as single point estimates without error bars, repeated-seed statistics, or significance tests. Collision rates such as 4%, 5%, and 9% and progress differences of several meters may be within scenario-level noise. Because the central support is empirical, please report means and standard deviations (or bootstrap confidence intervals) over multiple training runs and/or over scenario subsets.
minor comments (4)
  1. [Eq. (21)] There is an extra closing parenthesis in Eq. (21): 'max(0, b_-1 + b_0 + b_1 - 1))' should be 'max(0, b_-1 + b_0 + b_1 - 1)'.
  2. [Table III] The ablation row header 'No learnbale cost function' contains a typo; it should read 'No learnable cost function'.
  3. [Algorithm 1, lines 12-13] The algorithm says gradients are computed with respect to θ and {ω_i}, but θ is the optimization variable; line 13 updates ϕ and {ω_i}. The gradient should be with respect to network parameters ϕ (and cost weights), not θ. Please correct this inconsistency.
  4. [Section V, first paragraph] The description of dataset sampling is slightly confusing: 'randomly sample 10% (i.e., 100 data files)' followed by filtering 'resulting in a total of 88,123 frames' would benefit from stating whether the 88,123 frames are frames after filtering from the 100 files or from the full dataset.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity; the empirical benchmark against external baselines supports the central claim, with only routine self-citations to prior work.

full rationale

The paper's central contribution—joint differentiable optimization of lane-selection decisions and trajectories with a learned predictor—is supported by open-loop and closed-loop experiments against external baselines (Vanilla IL, IL+OPT, and DIPP) on the Waymo Open Motion Dataset. DIPP [15] is prior work by overlapping authors, but it is an externally published benchmark and is used both as a comparison and as a source of standard cost terms, not as the justification for the claimed decision-making capability. The added decision variables and one-hot constraints are formulated in this paper (Eqs. (3), (20)-(21)) rather than imported as a theorem from a self-citation. The decision loss (28) supervises the predictor with the optimizer's own output, which is an endogenous training loop; however, it does not reduce the reported driving-performance numbers to a fitted input, and the ablation without decision initialization shows materially different optimizer behavior and progress, so the loop is not degenerate. The concern that penalties (20)-(21) do not mathematically force integrality is a constraint-enforcement flaw, not circularity: the 100% compliance in Table V is an empirical claim about the optimized values, not a quantity that equals the input by construction. The only self-citation elements—building on [3] and [15]—are normal incremental-research references and are not load-bearing for the validity of the measured comparison. Hence no significant circularity; score 2 reflects the routine self-citation overlap, not a reduction of the central claim to its inputs.

Assumptions & free parameters 5 free parameters · 6 assumptions · 0 invented entities

The central claim rests on a hand-designed cost function with many weights (most values undisclosed, some learned), a relaxed binary decision variable whose penalties do not mathematically enforce integrality, predicted agent futures, and the Waymo dataset. No new physical entities are introduced.

free parameters (5)
  • Objective cost weights w_tr,x, w_tr,y, w_v-lon, w_d-lon, w_v-lat, w_d-lat, w_velo, w_rc1, w_rc2 = not reported (some learned)
    Eqs. (4)-(16) define all costs in terms of these constants; Algorithm 1 says cost weights {omega_i} are updated during training, but final values are not disclosed.
  • Penalty weights w_safe, w_stop, w_bi=10, w_eq=1000 = w_bi=10, w_eq=1000; others unlisted
    Section IV-A.5 and IV-A.6; Table V shows that performance depends on these hand-set weights.
  • Loss weights lambda_1..lambda_5 = 0.5, 1, 1, 1, 0.1
    Eq. (24) and Table I; chosen by hand, affect training balance between prediction, decision, and planning losses.
  • Optimizer step size beta and planning iterations = beta=0.4 train, 0.5 inference; iterations 2 train, 10 inference
    Table I; training with only 2 Gauss-Newton iterations may bias the gradients that guide upstream learning.
  • Numerical stability constant epsilon and safety distance eps = not specified
    Eqs. (6), (11), (17); small nonzero constants prevent division by zero and define safety margins.
assumptions (6)
  • domain assumption Kinematic bicycle model (Eq. 2) accurately represents the ego vehicle over the 5s planning horizon
    Used in all dynamics constraints and cost computations; ignores tire slip, actuation limits, and model error.
  • domain assumption The weighted sum cost (Eq. 22) correctly encodes safety, traveling efficiency, and riding comfort
    The paper's optimality claim rests on this hand-designed objective; no validation that the weights correlate with human preferences or true risk.
  • ad hoc to paper The finite penalty relaxations (Eqs. 20-21) preserve the binary and equality decision constraints
    As written, penalties are zero for any b in [0,1] with sum at most 1, so they do not force integrality; 100% compliance is only an empirical observation on the evaluated frames.
  • domain assumption Predicted trajectories of surrounding agents are reliable enough to plan against
    The optimizer uses network predictions as if they were true future positions; no uncertainty weighting or safety margin beyond the hinge loss.
  • domain assumption Waymo Open Motion Dataset and the closed-loop log-replay protocol provide a valid measure of driving performance
    The regime is open-loop and log-replay closed-loop, not a full closed-loop simulator with interactive agents; results may not transfer to interactive traffic.
  • standard math Gauss-Newton / Theseus provides correct gradients for end-to-end training with two inner iterations
    Differentiability of the optimizer is standard, but unrolling only 2 iterations is an approximation whose effect on training is not analyzed.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Integrating Decision-Making Into Differentiable Optimization Guided Learning for End-to-End Planning of Autonomous Vehicles." pith.science (2026). https://pith.science/paper/YLEZILDB

@misc{pith2026241201234,
  author       = {Pith},
  title        = {Pith review of: Integrating Decision-Making Into Differentiable Optimization Guided Learning for End-to-End Planning of Autonomous Vehicles},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YLEZILDB}},
  note         = {Machine review of arXiv:2412.01234}
}
read the original abstract

We address the decision-making capability within an end-to-end planning framework that focuses on motion prediction, decision-making, and trajectory planning. Specifically, we formulate decision-making and trajectory planning as a differentiable nonlinear optimization problem, which ensures compatibility with learning-based modules to establish an end-to-end trainable architecture. This optimization introduces explicit objectives related to safety, traveling efficiency, and riding comfort, guiding the learning process in our proposed pipeline. Intrinsic constraints resulting from the decision-making task are integrated into the optimization formulation and preserved throughout the learning process. By integrating the differentiable optimizer with a neural network predictor, the proposed framework is end-to-end trainable, aligning various driving tasks with ultimate performance goals defined by the optimization objectives. The proposed framework is trained and validated using the Waymo Open Motion dataset. The open-loop testing reveals that while the planning outcomes using our method do not always resemble the expert trajectory, they consistently outperform baseline approaches with improved safety, traveling efficiency, and riding comfort. The closed-loop testing further demonstrates the effectiveness of optimizing decisions and improving driving performance. Ablation studies demonstrate that the initialization provided by the learning-based prediction module is essential for the convergence of the optimizer as well as the overall driving performance.

Figures

Figures reproduced from arXiv: 2412.01234 by the authors.

Figure 1
Figure 1. Illustration of our proposed approach. (a) The end-to-end planning [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Pipeline of learning-based predictor and the differentiable optimizer for the integrated decision-making and trajectory planning tasks. The proposed [PITH_FULL_IMAGE:figures/full_fig_p004_2.png] view at source ↗
Figure 3
Figure 3. Comparison of the planning outcomes by DIPP and our proposed framework in the open-loop testing. The top figures show the planned trajectory, [PITH_FULL_IMAGE:figures/full_fig_p010_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Representative scenarios of the proposed framework in closed-loop testing. The red solid lines are the planned trajectories for the AV. Top: optimized [PITH_FULL_IMAGE:figures/full_fig_p012_4.png]
Figure 5
Figure 5. Figure 5: Comparison of the planned trajectory with and without the learned [PITH_FULL_IMAGE:figures/full_fig_p013_5.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

47 extracted references · 40 canonical work pages

  1. [15]

    Differentiable Integrated Motion Prediction and Planning with Learnable Cost Function for Autonomous Driving,

    Z. Huang, H. Liu, J. Wu, and C. Lv, “Differentiable Integrated Motion Prediction and Planning with Learnable Cost Function for Autonomous Driving,” IEEE Transactions on Neural Networks and Learning Systems, 2023

  2. [1]

    Recent Advancements in End-to-End Au- tonomous Driving Using Deep Learning: A Survey,

    P. S. Chib and P. Singh, “Recent Advancements in End-to-End Au- tonomous Driving Using Deep Learning: A Survey,” IEEE Transactions on Intelligent Vehicles, 2023

  3. [2]

    Autonomous Driving: A Bird’s Eye View,

    M. Mart ´ınez-D´ıaz, F. Soriguera, and I. P ´erez, “Autonomous Driving: A Bird’s Eye View,” IET intelligent transport systems , vol. 13, no. 4, pp. 563–579, 2019

  4. [3]

    Synergizing Decision Making and Trajectory Planning Using Two-Stage Optimization for Autonomous Vehicles

    W. Liu, H. Liu, L. Zeng, Z. Huang, and J. Ma, “Synergizing decision making and trajectory planning using two-stage optimization for au- tonomous vehicles,” arXiv preprint arXiv:2411.18974 , 2024

  5. [4]

    End-to-End Autonomous Driving: Challenges and Frontiers,

    L. Chen, P. Wu, K. Chitta, B. Jaeger, A. Geiger, and H. Li, “End-to-End Autonomous Driving: Challenges and Frontiers,” IEEE Transactions on Pattern Analysis and Machine Intelligence , 2024

  6. [5]

    Milestones in Autonomous Driving and Intelligent Vehicles: Survey of Surveys,

    L. Chen, Y . Li, C. Huang, B. Li, Y . Xing, D. Tian, L. Li, Z. Hu, X. Na, Z. Li, et al., “Milestones in Autonomous Driving and Intelligent Vehicles: Survey of Surveys,”IEEE Transactions on Intelligent Vehicles, vol. 8, no. 2, pp. 1046–1056, 2022

  7. [6]

    Rethinking integration of prediction and planning in deep learning-based automated driving systems: a review,

    S. Hagedorn, M. Hallgarten, M. Stoll, and A. Condurache, “Rethinking integration of prediction and planning in deep learning-based automated driving systems: a review,” arXiv preprint arXiv:2308.05731 , 2023. 14

  8. [7]

    A Path Towards Autonomous Machine Intelligence Version 0.9.2, 2022-06-27,

    Y . LeCun, “A Path Towards Autonomous Machine Intelligence Version 0.9.2, 2022-06-27,” Open Review, vol. 62, no. 1, pp. 1–62, 2022

Show all 47 references
  1. [8]

    Is It Safe to Drive? An Overview of Factors, Metrics, and Datasets for Driveability Assessment in Au- tonomous Driving,

    J. Guo, U. Kurup, and M. Shah, “Is It Safe to Drive? An Overview of Factors, Metrics, and Datasets for Driveability Assessment in Au- tonomous Driving,” IEEE Transactions on Intelligent Transportation Systems, vol. 21, no. 8, pp. 3135–3151, 2019

  2. [9]

    A Survey of Deep RL and IL for Autonomous Driving Policy Learning,

    Z. Zhu and H. Zhao, “A Survey of Deep RL and IL for Autonomous Driving Policy Learning,” IEEE Transactions on Intelligent Transporta- tion Systems, vol. 23, no. 9, pp. 14043–14065, 2021

  3. [10]

    Local Learning Enabled Iterative Linear Quadratic Regulator for Con- strained Trajectory Planning,

    J. Ma, Z. Cheng, X. Zhang, Z. Lin, F. L. Lewis, and T. H. Lee, “Local Learning Enabled Iterative Linear Quadratic Regulator for Con- strained Trajectory Planning,” IEEE Transactions on Neural Networks and Learning Systems , vol. 34, no. 9, pp. 5354–5365, 2022

  4. [11]

    Game-Theoretic Driver Modeling and Decision-Making for Autonomous Driving with Temporal-Spatial Attention-Based Deep Q-Learning,

    X. Zhou, Z. Peng, Y . Xie, M. Liu, and J. Ma, “Game-Theoretic Driver Modeling and Decision-Making for Autonomous Driving with Temporal-Spatial Attention-Based Deep Q-Learning,” IEEE Transac- tions on Intelligent Vehicles , 2024

  5. [12]

    A Two- Stage Optimization-Based Motion Planner for Safe Urban Driving,

    F. Eiras, M. Hawasly, S. V . Albrecht, and S. Ramamoorthy, “A Two- Stage Optimization-Based Motion Planner for Safe Urban Driving,” IEEE Transactions on Robotics , vol. 38, no. 2, pp. 822–834, 2021

  6. [13]

    Integrated Decision Making and Trajectory Planning for Autonomous Driving Under Multimodal Uncer- tainties: A Bayesian Game Approach,

    Z. Huang, T. Li, S. Shen, and J. Ma, “Integrated Decision Making and Trajectory Planning for Autonomous Driving Under Multimodal Uncer- tainties: A Bayesian Game Approach,” arXiv preprint arXiv:2409.13993, 2024

  7. [14]

    Improved Consensus ADMM for Cooperative Motion Planning of Large-Scale Connected Autonomous Vehicles with Limited Communication,

    H. Liu, Z. Huang, Z. Zhu, Y . Li, S. Shen, and J. Ma, “Improved Consensus ADMM for Cooperative Motion Planning of Large-Scale Connected Autonomous Vehicles with Limited Communication,” IEEE Transactions on Intelligent Vehicles, 2024

  8. [16]

    A Survey of Deep Learning Techniques for Autonomous Driving,

    S. Grigorescu, B. Trasnea, T. Cocias, and G. Macesanu, “A Survey of Deep Learning Techniques for Autonomous Driving,” Journal of Field Robotics, vol. 37, no. 3, pp. 362–386, 2020

  9. [17]

    Imitation Learning for Agile Autonomous Driving,

    Y . Pan, C.-A. Cheng, K. Saigol, K. Lee, X. Yan, E. A. Theodorou, and B. Boots, “Imitation Learning for Agile Autonomous Driving,” The International Journal of Robotics Research , vol. 39, no. 2-3, pp. 286– 302, 2020

  10. [18]

    Hybrid Trajectory Planning for Autonomous Driving in On-Road Dynamic Scenarios,

    W. Lim, S. Lee, M. Sunwoo, and K. Jo, “Hybrid Trajectory Planning for Autonomous Driving in On-Road Dynamic Scenarios,” IEEE Transac- tions on Intelligent Transportation Systems, vol. 22, no. 1, pp. 341–355, 2019

  11. [19]

    Alternating Direction Method of Multipliers for Constrained Iterative LQR in Autonomous Driving,

    J. Ma, Z. Cheng, X. Zhang, M. Tomizuka, and T. H. Lee, “Alternating Direction Method of Multipliers for Constrained Iterative LQR in Autonomous Driving,” IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 12, pp. 23031–23042, 2022

  12. [20]

    Safe Nonlinear Trajectory Generation for Parallel Autonomy with a Dynamic Vehicle Model,

    W. Schwarting, J. Alonso-Mora, L. Paull, S. Karaman, and D. Rus, “Safe Nonlinear Trajectory Generation for Parallel Autonomy with a Dynamic Vehicle Model,” IEEE Transactions on Intelligent Transportation Sys- tems, vol. 19, no. 9, pp. 2994–3008, 2017

  13. [21]

    Bonnans, J

    J.-F. Bonnans, J. C. Gilbert, C. Lemar ´echal, and C. A. Sagastiz ´abal, Numerical Optimization: Theoretical and Practical Aspects . Springer Science & Business Media, 2006

  14. [22]

    Maplite: Autonomous Intersection Navigation Without a Detailed Prior Map,

    T. Ort, K. Murthy, R. Banerjee, S. K. Gottipati, D. Bhatt, I. Gilitschenski, L. Paull, and D. Rus, “Maplite: Autonomous Intersection Navigation Without a Detailed Prior Map,” IEEE Robotics and Automation Letters , vol. 5, no. 2, pp. 556–563, 2019

  15. [23]

    Planning-Oriented Autonomous Driving,

    Y . Hu, J. Yang, L. Chen, K. Li, C. Sima, X. Zhu, S. Chai, S. Du, T. Lin, W. Wang, et al. , “Planning-Oriented Autonomous Driving,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 17853–17862, 2023

  16. [24]

    V AD: Vectorized Scene Representation for Efficient Autonomous Driving,

    B. Jiang, S. Chen, Q. Xu, B. Liao, J. Chen, H. Zhou, Q. Zhang, W. Liu, C. Huang, and X. Wang, “V AD: Vectorized Scene Representation for Efficient Autonomous Driving,” in Proceedings of the IEEE/CVF International Conference on Computer Vision , pp. 8340–8350, 2023

  17. [25]

    An End-to-End Curriculum Learning Approach for Autonomous Driving Scenarios,

    L. Anzalone, P. Barra, S. Barra, A. Castiglione, and M. Nappi, “An End-to-End Curriculum Learning Approach for Autonomous Driving Scenarios,” IEEE Transactions on Intelligent Transportation Systems , vol. 23, no. 10, pp. 19817–19826, 2022

  18. [26]

    A Survey on Imitation Learning Techniques for End-to-End Autonomous Vehicles,

    L. Le Mero, D. Yi, M. Dianati, and A. Mouzakitis, “A Survey on Imitation Learning Techniques for End-to-End Autonomous Vehicles,” IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 9, pp. 14128–14147, 2022

  19. [27]

    Drive Anywhere: Generalizable End-to-End Autonomous Driving with Multi-Modal Foundation Models,

    T.-H. Wang, A. Maalouf, W. Xiao, Y . Ban, A. Amini, G. Rosman, S. Karaman, and D. Rus, “Drive Anywhere: Generalizable End-to-End Autonomous Driving with Multi-Modal Foundation Models,” in 2024 IEEE International Conference on Robotics and Automation (ICRA) , pp. 6687–6694, IEEE, 2024

  20. [28]

    IR-STP: Enhancing Autonomous Driving with Interaction Reasoning in Spatio-Temporal Planning,

    Y . Chen, J. Cheng, L. Gan, S. Wang, H. Liu, X. Mei, and M. Liu, “IR-STP: Enhancing Autonomous Driving with Interaction Reasoning in Spatio-Temporal Planning,” IEEE Transactions on Intelligent Trans- portation Systems, 2024

  21. [29]

    Multimodal End-to-End Autonomous Driving,

    Y . Xiao, F. Codevilla, A. Gurram, O. Urfalioglu, and A. M. L ´opez, “Multimodal End-to-End Autonomous Driving,” IEEE Transactions on Intelligent Transportation Systems, vol. 23, no. 1, pp. 537–547, 2020

  22. [30]

    End-to-End Autonomous Driving: An Angle Branched Network Approach,

    Q. Wang, L. Chen, B. Tian, W. Tian, L. Li, and D. Cao, “End-to-End Autonomous Driving: An Angle Branched Network Approach,” IEEE Transactions on Vehicular Technology, vol. 68, no. 12, pp. 11599–11610, 2019

  23. [31]

    End-to-End Interpretable Neural Motion Planner,

    W. Zeng, W. Luo, S. Suo, A. Sadat, B. Yang, S. Casas, and R. Urtasun, “End-to-End Interpretable Neural Motion Planner,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 8660–8669, 2019

  24. [32]

    Efficient Deep Learning of Robust Policies from MPC Using Imitation and Tube-Guided Data Augmentation,

    A. Tagliabue and J. P. How, “Efficient Deep Learning of Robust Policies from MPC Using Imitation and Tube-Guided Data Augmentation,”IEEE Transactions on Robotics , 2024

  25. [33]

    Evolutionary Decision-Making and Planning for Autonomous Driving: A Hybrid Augmented Intelligence Framework,

    K. Yuan, Y . Huang, S. Yang, M. Wu, D. Cao, Q. Chen, and H. Chen, “Evolutionary Decision-Making and Planning for Autonomous Driving: A Hybrid Augmented Intelligence Framework,” IEEE Transactions on Intelligent Transportation Systems, 2024

  26. [34]

    ChauffeurNet: Learning to drive by imitating the best and synthesizing the worst,

    M. Bansal, A. Krizhevsky, and A. Ogale, “ChauffeurNet: Learning to drive by imitating the best and synthesizing the worst,” arXiv preprint arXiv:1812.03079, 2018

  27. [35]

    BEV-TP: End-to-End Visual Perception and Trajectory Prediction for Autonomous Driving,

    B. Lang, X. Li, and M. C. Chuah, “BEV-TP: End-to-End Visual Perception and Trajectory Prediction for Autonomous Driving,” IEEE Transactions on Intelligent Transportation Systems , 2024

  28. [36]

    Differen- tiable MPC for end-to-end planning and control,

    B. Amos, I. Jimenez, J. Sacks, B. Boots, and J. Z. Kolter, “Differen- tiable MPC for end-to-end planning and control,” Advances in neural information processing systems , vol. 31, 2018

  29. [37]

    Exploring Imitation Learning for Autonomous Driving with Feedback Synthesizer and Differentiable Rasterization,

    J. Zhou, R. Wang, X. Liu, Y . Jiang, S. Jiang, J. Tao, J. Miao, and S. Song, “Exploring Imitation Learning for Autonomous Driving with Feedback Synthesizer and Differentiable Rasterization,” in 2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pp...

  30. [38]

    Deep Kinematic Models for Kinematically Feasible Vehicle Trajectory Predictions,

    H. Cui, T. Nguyen, F.-C. Chou, T.-H. Lin, J. Schneider, D. Bradley, and N. Djuric, “Deep Kinematic Models for Kinematically Feasible Vehicle Trajectory Predictions,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) , pp. 10563–10569, IEEE, 2020

  31. [39]

    LeTO: Learning Constrained Visuomotor Policy with Differentiable Trajectory Optimization,

    Z. Xu and Y . She, “LeTO: Learning Constrained Visuomotor Policy with Differentiable Trajectory Optimization,” arXiv preprint arXiv:2401.17500, 2024

  32. [40]

    DC3: A learning method for optimization with hard constraints,

    P. Donti, D. Rolnick, and J. Z. Kolter, “DC3: A learning method for optimization with hard constraints,” in International Conference on Learning Representations, 2021

  33. [41]

    Dif- ferentiable Constrained Imitation Learning for Robot Motion Planning and Control,

    C. Diehl, J. Adamek, M. Kr ¨uger, F. Hoffmann, and T. Bertram, “Dif- ferentiable Constrained Imitation Learning for Robot Motion Planning and Control,” arXiv preprint arXiv:2210.11796 , 2022

  34. [42]

    DTPP: Differentiable joint conditional prediction and cost evalua- tion for tree policy planning in autonomous driving,

    Z. Huang, P. Karkus, B. Ivanovic, Y . Chen, M. Pavone, and C. Lv, “DTPP: Differentiable joint conditional prediction and cost evalua- tion for tree policy planning in autonomous driving,” arXiv preprint arXiv:2310.05885, 2023

  35. [43]

    Diffstack: A differentiable and modular control stack for autonomous vehicles,

    P. Karkus, B. Ivanovic, S. Mannor, and M. Pavone, “Diffstack: A differentiable and modular control stack for autonomous vehicles,” in Conference on robot learning , pp. 2170–2180, PMLR, 2023

  36. [44]

    Theseus: A library for differentiable nonlinear optimization,

    L. Pineda, T. Fan, M. Monge, S. Venkataraman, P. Sodhi, R. T. Chen, J. Ortiz, D. DeTone, A. Wang, S. Anderson, et al. , “Theseus: A library for differentiable nonlinear optimization,” Advances in Neural Information Processing Systems , vol. 35, pp. 3801–3818, 2022

  37. [45]

    Decentralized iLQR for Coopera- tive Trajectory Planning of Connected Autonomous Vehicles via Dual Consensus ADMM,

    Z. Huang, S. Shen, and J. Ma, “Decentralized iLQR for Coopera- tive Trajectory Planning of Connected Autonomous Vehicles via Dual Consensus ADMM,” IEEE Transactions on Intelligent Transportation Systems, vol. 24, no. 11, pp. 12754–12766, 2023

  38. [46]

    Scalability in Perception for Autonomous Driving: Waymo Open Dataset,

    P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V . Patnaik, P. Tsui, J. Guo, Y . Zhou, Y . Chai, B. Caine, et al. , “Scalability in Perception for Autonomous Driving: Waymo Open Dataset,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition ,...

  39. [47]

    An Integrated Decision and Motion Planning Framework for Automated Driving on Highway,

    P. Wu, F. Gao, X. Tang, and K. Li, “An Integrated Decision and Motion Planning Framework for Automated Driving on Highway,”IEEE Transactions on Vehicular Technology, vol. 72, no. 12, pp. 15574–15584, 2023

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.