Pith. sign in

REVIEW 3 major objections 4 minor 4 cited by

Synthesis of Model Predictive Control and Reinforcement Learning: Survey and Classification

T0 review · 3 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read This survey claims the MPC+RL combination literature can be organized by where MPC sits inside an actor-critic framework: as expert actor, as part of the deployed policy, or as critic.

desk verdict A solid, useful survey: the actor-critic role taxonomy is new and works as a map, but it is multi-label by design and should be described as such. read the letter →

arxiv 2502.02133 v1 pith:VQMYEWGY submitted 2025-02-04 eess.SY cs.AIcs.LGcs.SY

classification eess.SYcs.AIcs.LGcs.SY
keywords modelpredictivecontrolreinforcementlearningactor-criticsurveyclassificationimitationclosed-loopsynthesis
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper tries to establish a systematic way of dividing the growing literature that combines model predictive control (MPC) and reinforcement learning (RL). Instead of sorting papers by application or algorithm, it sorts them by the algorithmic role MPC plays inside an RL actor-critic framework. It claims three roles cover the field: MPC as an expert actor that guides or demonstrates for training, MPC inside the deployed policy (learned as an optimization layer or used for pre/postprocessing), and MPC as a critic that supplies value estimates. Alongside this, it defines four parameterized-MPC architectures (integrated, hierarchical, parallel, parameterized) and distinguishes aligned learning from closed-loop learning within the deployed-policy role. If the taxonomy is right, the field gains a common vocabulary and a structured map of roughly fifty works, plus a clear statement of the main methodological fork for training MPC-based policies.

What carries the argument

The organizing device is the actor-critic decomposition of RL, used as a coordinate system. MPC is treated as a modular optimization layer that can be placed in the expert, actor, or critic position, and parameterized MPC is classified by how a learned function approximator enters the optimization problem: integrated, where the learned function depends on decision variables inside the optimization; hierarchical, where a learned function is evaluated beforehand and supplies parameters; parallel, where an NN and MPC are evaluated side by side and their outputs are combined; and parameterized, where learned parameters are constant and do not depend on state or decision variables. This machinery carries the classification: every surveyed combination is described by one role plus one architecture, and the aligned-versus-closed-loop distinction then splits the actor role.

What would settle it

Locate a published MPC+RL combination whose MPC component both provides training targets and remains in the deployed loop, forcing placement into two of the three roles; the paper's own placement of [263] in both Table V and Table VII indicates the categories are not exclusive.

Watch

Extended reading notes

Core claim

The central claim is that every way of combining MPC and RL can be understood as inserting the MPC optimization problem into one of three algorithmic parts of an actor-critic RL agent: as an expert actor whose fixed policy is mimicked or used to guide exploration (Section VII); as part of the deployed policy, either aligned with the MDP by learning model, terminal value function, or stage cost, trained for closed-loop optimality, or used for pre/postprocessing (Section VIII); or as part of the critic, where MPC computes or parameterizes value/action-value estimates (Section IX). The paper supports this with a modular diagram (Fig. 2), a survey of roughly 50 works in Tables IV through IX, and a theoretical section that links MPC-MDP equivalence results to these roles. It further argues that within the deployed-policy role the crucial methodological divide is between MDP-aligned learning, which makes each MPC component approximate the corresponding MDP component, and closed-loop learning, which changes model, cost, or constraints purely to improve closed-loop performance.

Load-bearing premise

The classification assumes that the three roles (expert actor, part of the deployed policy, part of the critic) are jointly exhaustive and that each surveyed work can be assigned cleanly to one role and one architecture; the paper itself places some works in multiple tables, so the map can overlap.

Editorial extensions

If this is right

  • Researchers can place new MPC+RL algorithms into one of three roles and one of four architectures, making method comparisons and transfers easier.
  • The aligned-versus-closed-loop distinction clarifies the main design choice: either make each MPC component approximate the MDP (model, value, cost) or tune the whole optimization layer for closed-loop performance.
  • The theoretical results reviewed in Section X, especially the MPC-MDP equivalence theorems, justify using MPC as actor or critic by showing when a parameterized MPC can reproduce the optimal policy and value functions of a discounted MDP.
  • The survey's software section indicates that practical combination work is currently limited by solver support for differentiation and integration with learning frameworks; closing that gap would accelerate the field.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The taxonomy could function as a generative design tool: choosing a role and an architecture enumerates the combination space, exposing under-explored cells such as MPC as a learnable actor-critic.
  • As differentiation tools mature, closed-loop optimal learning may displace MDP-aligned learning for performance-critical applications because it optimizes the actual deployment objective rather than individual components.
  • The role-based map also unifies imitation learning, guided policy search, and safe RL: many safe-RL schemes are MPC-as-filter or MPC-as-expert, and labeling them by role could reveal redundant designs.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 4 minor

Summary. The manuscript surveys methods that combine model predictive control (MPC) and reinforcement learning (RL). It unifies the notation of the two fields, compares their practical and theoretical properties, and proposes a taxonomy based on the algorithmic role of MPC within an actor-critic RL framework: MPC as an expert actor (Sect. VII), MPC within the deployed policy (Sect. VIII), and MPC as a critic (Sect. IX), together with four parameterized-MPC architectures (integrated, hierarchical, parallel, parameterized; Sect. VI-A). The survey covers roughly fifty works in Tables IV–IX, includes a theory section (Sect. X) centered on an MPC–MDP equivalence theorem, and closes with software tools and implementation aspects. The paper is notably honest about its own limitations, explicitly flagging the sampling-notation issue in Remark 3.1 and marking the convergence target of the terminal-value-function update with a question mark in Sect. VIII-A2.

Significance. If the taxonomy is accepted, the paper provides a genuinely useful organizational contribution: a role-based vocabulary for the MPC+RL combination literature, an extensive tabulation of applications and algorithms, and a careful notation unification. The paper does not overclaim its theoretical grounding, and its software discussion is a practical asset. The central classification claim, however, is currently underspecified because the categories are not shown to be mutually exclusive or cleanly assignable, and the tables do not state a convention for works that realize multiple roles. Until this is resolved, the claimed 'shared vocabulary' is not yet a stable map of the field. The comparison tables and the honest treatment of open theoretical points are valuable independently of this issue.

major comments (3)
  1. [§VII-A, §IX-A, Tables IV–VIII] The role taxonomy is presented as the paper's central organizational contribution, but the categories are not stated to be mutually exclusive, and the text gives conflicting guidance on assignment. Section VII-A explicitly moves Q-function-based imitation into the critic category ('we categorize such methods in a class named MPC as a critic'), while Section IX-A states that using an MPC as an expert critic 'leads to imitating the MPC expert policy.' A fixed-MPC method that trains a policy from MPC demonstrations can therefore be placed in either the expert-actor category (if the training loss uses actions) or the expert-critic category (if the loss uses Q-values), even though the MPC occupies the same position in the RL architecture. The tables reflect this ambiguity: reference [263] appears in Tables V and VII, and reference [162] appears in Tables III, IV, and VII, with no stated convention for multi-role entries. Please either declare the taxonomy explicitly multi-label and define membership criteria, or adjust the category definitions so that the classification is a partition of the surveyed works.
  2. [§VIII-B] The passage defining closed-loop learning refers to 'the MPC policy µMPC according to Definition 4,' but no Definition 4 exists anywhere in the manuscript. Since the aligned-versus-closed-loop distinction is one of the two main axes of the deployed-policy category, the missing definition is load-bearing: the reader cannot verify what 'closed-loop optimal' means at the point where the taxonomy is introduced. Please add the missing definition at the point of use, or rephrase the passage to refer to the characterization later given in Sect. X (condition 1 of Theorem 10.1).
  3. [§VIII-A and §X] The term 'MDP aligned' is used inconsistently. In §VIII-A, 'aligned learning' requires the MPC's model, costs, constraints, and terminal value function to 'individually approximate the MDP and the optimal value function'; in §X, 'the MDP aligned property ... demands that the MPC model exactly matches the state-transition kernel of the MDP' and is said to be 'rarely satisfied in practice.' These are different properties (component-wise approximation versus exact transition-kernel match), and the latter definition would make most of the Table V entries not 'MDP aligned' at all. Since the aligned-versus-closed-loop axis is central to the deployed-policy classification, the terminology should be reconciled, or two distinct names should be introduced.
minor comments (4)
  1. [Fig. 1] Figure 1 lists both 'Tab. IX: Literature using MPC as a critic' and 'Tab. IX: Theoretical results for combining MPC and RL'; the second of these should be Table X.
  2. [Table VII and Fig. 1] Table VII's header reads 'MPC for Postrocessing' and Fig. 1 contains 'Literature wiht MPC as filter'; both are typos that should be corrected.
  3. [§XI-B1] The statement that solver differentiation support 'is currently ongoing work in acados' and 'will be detailed in a future publication' is acceptable for a survey but should be time-stamped or softened to avoid ageing quickly.
  4. [§V-10] The sentence reporting that 'most authors report superiority of MPC' with cost reductions of 'less than 4%' is vague because the baseline (RL variant, task, and metric) is not stated in the surrounding text; please make the comparison explicit or point to the table entries that support it.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: survey classification is an independent organizational contribution; theory self-citations are external theorems with proofs.

full rationale

This paper is a survey and classification; it does not claim to derive predictions from fitted parameters. Its central contribution is a categorization of MPC+RL combination methods by the algorithmic role of MPC (expert actor, deployed policy, critic). This categorization is defined directly from the literature and does not presuppose the conclusions it organizes. The theoretical section (Sect. X) cites several papers from the same research group (e.g., [224], [262], [286], [316], [317], [328], [288]), but these are external mathematical results with proofs (e.g., Theorem 10.1 from Kordabad et al., proof referenced as 'See [262]'), not ansatz or uniqueness assertions that force the taxonomy. The overlap noted between Sect. VII-A (Q-function-based imitation categorized as critic) and Sect. IX-A (expert critic imitates expert policy) is a classification ambiguity, not a derivation that reduces to its own inputs: no equation is defined in terms of the conclusion, and no fitted parameter is renamed as a prediction. The paper even flags an open question in the terminal-value learning update (footnote 1) and notes missing solver capabilities (Sect. XI-B1), consistent with an honest survey rather than a circular argument.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper's contribution is an organizational taxonomy, not a derivation, so there are no fitted free parameters and no new postulated entities. The taxonomy rests on (i) standard dynamic programming mathematics, and (ii) a set of cited theoretical results, mainly from the authors' own prior work (Gros, Zanon, Kordabad, Diehl et al.), which are asserted rather than derived in this text, most notably Theorem 10.1. The architectural typology in Fig. 3 is an original scheme but is a categorization, not an empirically testable entity.

assumptions (4)
  • domain assumption Solving the MDP in (4) is the common normative problem for both MPC and RL; MPC is treated as an approximation of (4).
    Adopted in Sects. II and IV to unify the two fields; the comparison framework weakens if MPC is not usefully viewed as an approximate MDP solver, though the paper notes economic MPC (Sect. IV-A2) builds the bridge.
  • standard math Bellman operators T^pi (7) and T (8) are contraction mappings for gamma < 1, giving unique fixed points.
    Invoked in Sect. III-A1 to justify value and policy iteration; standard result from [23], not proved in the text.
  • domain assumption Assumption 10.1 from [262]: there is a non-empty set of states for which the expected optimal value under the MPC model is finite up to a given horizon.
    Quoted in Sect. X as the standing assumption for Theorem 10.1; it bounds the model mismatch that the equivalence result tolerates.
  • domain assumption Theorem 10.1 from [262]: there exist terminal and stage costs such that the MPC policy, value function and Q-function equal the MDP optima for all discount factors and horizons, even under model mismatch.
    Stated in Sect. X with 'Proof 10.1: See [262]'; the survey does not derive it, so the theoretical foundation of closed-loop learning rests on a cited result from the authors' own group.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Synthesis of Model Predictive Control and Reinforcement Learning: Survey and Classification." pith.science (2026). https://pith.science/paper/VQMYEWGY

@misc{pith2026250202133,
  author       = {Pith},
  title        = {Pith review of: Synthesis of Model Predictive Control and Reinforcement Learning: Survey and Classification},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VQMYEWGY}},
  note         = {Machine review of arXiv:2502.02133}
}
read the original abstract

The fields of MPC and RL consider two successful control techniques for Markov decision processes. Both approaches are derived from similar fundamental principles, and both are widely used in practical applications, including robotics, process control, energy systems, and autonomous driving. Despite their similarities, MPC and RL follow distinct paradigms that emerged from diverse communities and different requirements. Various technical discrepancies, particularly the role of an environment model as part of the algorithm, lead to methodologies with nearly complementary advantages. Due to their orthogonal benefits, research interest in combination methods has recently increased significantly, leading to a large and growing set of complex ideas leveraging MPC and RL. This work illuminates the differences, similarities, and fundamentals that allow for different combination algorithms and categorizes existing work accordingly. Particularly, we focus on the versatile actor-critic RL approach as a basis for our categorization and examine how the online optimization approach of MPC can be used to improve the overall closed-loop performance of a policy.

Figures

Figures reproduced from arXiv: 2502.02133 by the authors.

Figure 1
Figure 1. Paper structure. The main sections are highlighted in gray, subchapters in white, [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. Modular view on combinations of MPC and RL. The combinations are aligned [PITH_FULL_IMAGE:figures/full_fig_p014_2.png] view at source ↗
Figure 3
Figure 3. Parameterized MPC architectures: Proposed architectures of actors (or potentially [PITH_FULL_IMAGE:figures/full_fig_p016_3.png] view at source ↗
Figures from the paper (4 more)
Figure 4
Figure 4. Figure 4: Combinations: MPC as an expert actor. The plot is split into the learning and [PITH_FULL_IMAGE:figures/full_fig_p017_4.png]
Figure 5
Figure 5. Figure 5: Combinations: MPC within the deployed policy. The plot is split into the learning [PITH_FULL_IMAGE:figures/full_fig_p019_5.png]
Figure 6
Figure 6. Figure 6: MPC can be used as a reference generator for an RL policy. The distribution of [PITH_FULL_IMAGE:figures/full_fig_p021_6.png]
Figure 7
Figure 7. Figure 7: MPC can be used to smooth or filter reference trajectories. For example, MPC [PITH_FULL_IMAGE:figures/full_fig_p022_7.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 4 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bridging Reinforcement Learning and Optimal Control via Feasible Action Mapping

    eess.SY 2026-07 conditional novelty 6.0 of 10

    FAOC maps RL actions into a state-dependent feasible parameter set, guaranteeing optimal-control feasibility and improving RL-MPC performance on table tennis.

  2. Mini Amusement Parks (MAPs): A Testbed for Modelling Business Decisions

    cs.AI 2025-11 conditional novelty 6.0 of 10

    MAPs is a new amusement-park simulator benchmark on which frontier LLM agents score 7–15% of human performance, exposing persistent gaps in long-horizon planning, active learning, spatial reasoning, and handling stoch...

  3. Reference-Free Iterative Learning Model Predictive Control with Neural Certificates

    eess.SY 2025-07 conditional novelty 6.0 of 10

    Reference-free iterative MPC replaces mixed-integer terminal constraints with a learned neural CLBF terminal set and cost, giving conditional recursive feasibility, stability, and non-increasing cost over iterations.

  4. Model-free Reinforcement Learning for Model-based Control: Towards Safe, Interpretable and Sample-efficient Agents

    cs.LG 2025-07 conditional novelty 3.0 of 10

    A perspective paper argues that model predictive control can be used as a learned policy in model-free reinforcement learning and reviews the methods and open problems.

Reference graph

Works this paper leans on

300 extracted references · 72 canonical work pages · cited by 4 Pith papers

  1. [8]

    Relations between Model Predictive Control and Reinforce- ment Learning,

    D. G ¨orges, “Relations between Model Predictive Control and Reinforce- ment Learning,” IFAC-PapersOnLine, vol. 50, no. 1, pp. 4920–4928, Jul. 2017

  2. [16]

    Learning- based model predictive control: Toward safe learning in control,

    L. Hewing, K. P. Wabersich, M. Menner, and M. N. Zeilinger, “Learning- based model predictive control: Toward safe learning in control,” Annual Review of Control, Robotics, and Autonomous Systems , vol. 3, no. 1, pp. 269–296, 2020

  3. [17]

    Fusion of Machine Learning and MPC under Uncertainty: What Advances Are on the Horizon?

    A. Mesbah, K. P. Wabersich, A. P. Schoellig, M. N. Zeilinger, S. Lucia, T. A. Badgwell, and J. A. Paulson, “Fusion of Machine Learning and MPC under Uncertainty: What Advances Are on the Horizon?” in American Control Conference (ACC) , Jun. 2022, pp. 342–357, iSSN: 2378-5861

  4. [162]

    Safe Imitation Learning of Nonlinear Model Predictive Control for Flexible Robots,

    S. Mamedov, R. Reiter, S. M. B. Azad, R. Viljoen, J. Boedecker, M. Diehl, and J. Swevers, “Safe Imitation Learning of Nonlinear Model Predictive Control for Flexible Robots,” in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , Oct. 2024, pp. 3613–3619, iSSN: 2153-0866

  5. [263]

    AC4MPC: Actor-Critic Reinforcement Learning for Nonlinear Model Predictive Control,

    R. Reiter, A. Ghezzi, K. Baumg ¨artner, J. Hoffmann, R. D. McAllister, and M. Diehl, “AC4MPC: Actor-Critic Reinforcement Learning for Nonlinear Model Predictive Control,” Jun. 2024

  6. [1]

    R. S. Sutton and A. G. Barto, Reinforcement learning: an introduction , second edition ed., ser. Adaptive computation and machine learning series. Cambridge, Massachusetts: The MIT Press, 2018

  7. [2]

    J. B. Rawlings, D. Q. Mayne, and M. M. Diehl, Model Predictive Control: Theory, Computation, and Design , 2nd ed. Santa Barbara, California: Nob Hill Publishing, 2017

  8. [3]

    Bertsekas, Reinforcement Learning and Optimal Control , first edition ed

    D. Bertsekas, Reinforcement Learning and Optimal Control , first edition ed. Belmont, Massachusetts: Athena Scientific, Jul. 2019

Show all 300 references
  1. [4]

    G. B. Dantzig, The Simplex Method. Santa Monica, CA: RAND Corporation, 1956

  2. [5]

    Review on model predictive control: an engineering perspective,

    M. Schwenzer, M. Ay, T. Bergs, and D. Abel, “Review on model predictive control: an engineering perspective,” The International Journal of Advanced Manufacturing Technology , vol. 117, pp. 1327 – 1349, 2021

  3. [6]

    Human-level control through deep reinforcement learning,

    V . Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Riedmiller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Kumaran, D. Wierstra, S. Legg, and D. Hassabis, “Human-level control through deep...

  4. [7]

    Outracing champion Gran Turismo drivers with deep reinforcement learning,

    P. R. Wurman, S. Barrett, K. Kawamoto, J. MacGlashan, K. Subra- manian, T. J. Walsh, R. Capobianco, A. Devlic, F. Eckert, F. Fuchs, L. Gilpin, P. Khandelwal, V . Kompella, H. Lin, P. MacAlpine, D. Oller, T. Seno, C. Sherstan, M. D. Thomure, H. Aghabozorgi, L. Barrett, R. Dougl...

  5. [9]

    Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning,

    L. Brunke, M. Greeff, A. W. Hall, Z. Yuan, S. Zhou, J. Panerati, and A. P. Schoellig, “Safe Learning in Robotics: From Learning-Based Control to Safe Reinforcement Learning,” Annual Review of Control, Robotics, and Autonomous Systems , vol. 5, no. 1, pp. 411–444, 2022

  6. [10]

    A Tour of Reinforcement Learning: The View from Contin- uous Control,

    B. Recht, “A Tour of Reinforcement Learning: The View from Contin- uous Control,” Annual Review of Control, Robotics, and Autonomous Systems, vol. 2, no. V olume 2, 2019, pp. 253–279, May 2019, publisher: Annual Reviews

  7. [11]

    Machine learning for combinato- rial optimization: A methodological tour d’horizon,

    Y . Bengio, A. Lodi, and A. Prouvost, “Machine learning for combinato- rial optimization: A methodological tour d’horizon,” European Journal of Operational Research , vol. 290, no. 2, pp. 405–421, Apr. 2021

  8. [12]

    Bertsekas, Lessons from AlphaZero for Optimal, Model Predictive, and Adaptive Control

    D. Bertsekas, Lessons from AlphaZero for Optimal, Model Predictive, and Adaptive Control . Athena Scientific, 2022

  9. [13]

    Newton’s method for reinforcement learning and model predic- tive control,

    ——, “Newton’s method for reinforcement learning and model predic- tive control,” Results in Control and Optimization , vol. 7, p. 100121, Jun. 2022

  10. [14]

    Mastering Chess and Shogi by Self-Play with a Gen- eral Reinforcement Learning Algorithm,

    D. Silver, T. Hubert, J. Schrittwieser, I. Antonoglou, M. Lai, A. Guez, M. Lanctot, L. Sifre, D. Kumaran, T. Graepel, T. Lillicrap, K. Simonyan, and D. Hassabis, “Mastering Chess and Shogi by Self-Play with a Gen- eral Reinforcement Learning Algorithm,” Dec. 2017, arXiv:1712.0...

  11. [15]

    Mastering the game of Go without human knowledge,

    D. Silver, J. Schrittwieser, K. Simonyan, I. Antonoglou, A. Huang, A. Guez, T. Hubert, L. Baker, M. Lai, A. Bolton, Y . Chen, T. Lillicrap, F. Hui, L. Sifre, G. van den Driessche, T. Graepel, and D. Hassabis, “Mastering the game of Go without human knowledge,” Nature, vol. 550...

  12. [18]

    Integrating Machine Learning and Model Predictive Control for automotive applications: A review and future directions,

    A. Norouzi, H. Heidarifar, H. Borhan, M. Shahbakhti, and C. R. Koch, “Integrating Machine Learning and Model Predictive Control for automotive applications: A review and future directions,” Engineering Applications of Artificial Intelligence , vol. 120, p. 105878, Apr. 2023

  13. [19]

    Building Energy Management With Reinforcement Learning and Model Predictive Control: A Survey,

    H. Zhang, S. Seal, D. Wu, F. Bouffard, and B. Boulet, “Building Energy Management With Reinforcement Learning and Model Predictive Control: A Survey,” IEEE Access, vol. 10, pp. 27 853–27 862, 2022

  14. [20]

    Benchmarking Model-Based Rein- forcement Learning,

    T. Wang, X. Bao, I. Clavera, J. Hoang, Y . Wen, E. Langlois, S. Zhang, G. Zhang, P. Abbeel, and J. Ba, “Benchmarking Model-Based Rein- forcement Learning,” Jul. 2019, arXiv:1907.02057 [cs]

  15. [21]

    M. L. Puterman, Markov decision processes: discrete stochastic dynamic programming, ser. Wiley series in probability and statistics. Hoboken, NJ: Wiley-Interscience, 2005, oCLC: 254152847

  16. [22]

    Dynamic programming,

    R. Bellman, “Dynamic programming,” American Association for the Advancement of Science , vol. 153, no. 3731, pp. 34–37, 1966

  17. [23]

    D. P. Bertsekas and J. N. Tsitsiklis, Neuro-dynamic programming, ser. Optimization and neural computation series. Belmont, Mass: Athena Scientific, 1996

  18. [24]

    W. B. Powell, Approximate Dynamic Programming: Solving the curses of dimensionality. John Wiley & Sons, 2007, vol. 703

  19. [25]

    Q-learning,

    C. J. C. H. Watkins and P. Dayan, “Q-learning,” Machine Learning, vol. 8, no. 3, pp. 279–292, May 1992

  20. [26]

    Simple statistical gradient-following algorithms for connectionist reinforcement learning,

    R. J. Williams, “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning, vol. 8, no. 3, pp. 229–256, May 1992

  21. [27]

    Deterministic policy gradient algorithms,

    D. Silver, G. Lever, N. Heess, T. Degris, D. Wierstra, and M. A. Riedmiller, “Deterministic policy gradient algorithms,” in Proceedings of the 31th international conference on machine learning, ICML 2014, beijing, china, 21-26 june 2014 , ser. JMLR workshop and conference proc...

  22. [28]

    Addressing function approximation error in actor-critic methods,

    S. Fujimoto, H. van Hoof, and D. Meger, “Addressing function approximation error in actor-critic methods,” in Proceedings of the 35th international conference on machine learning, ICML 2018 , ser. Proceedings of machine learning research, J. G. Dy and A. Krause, Eds., vol. 80....

  23. [29]

    Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,

    T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochastic actor,” in Proceedings of the 35th international conference on machine learning, ICML, ser. Proceedings of machine learning research, J...

  24. [30]

    Revisiting the Gumbel- Softmax in MADDPG,

    C. R. Tilbury, F. Christianos, and S. V . Albrecht, “Revisiting the Gumbel- Softmax in MADDPG,” 2023, arXiv:2302.11793 [cs]

  25. [31]

    Efficient BackProp,

    Y . LeCun, L. Bottou, G. B. Orr, and K. R. M ¨uller, “Efficient BackProp,” in Neural Networks: Tricks of the Trade , G. B. Orr and K.-R. M ¨uller, Eds. Berlin, Heidelberg: Springer, 1998, pp. 9–50

  26. [32]

    Continuous control with deep reinforcement learning,

    T. P. Lillicrap, J. J. Hunt, A. Pritzel, N. Heess, T. Erez, Y . Tassa, D. Silver, and D. Wierstra, “Continuous control with deep reinforcement learning,” in 4th International Conference on Learning Representations, ICLR , Y . Bengio and Y . LeCun, Eds., 2016

  27. [33]

    Proximal policy optimization algorithms,

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal policy optimization algorithms,” CoRR, vol. abs/1707.06347, 2017

  28. [34]

    Trust region policy optimization,

    J. Schulman, S. Levine, P. Moritz, M. I. Jordan, and P. Abbeel, “Trust region policy optimization,” CoRR, vol. abs/1502.05477, 2015

  29. [35]

    High- Dimensional Continuous Control Using Generalized Advantage Esti- 29 mation

    J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel, “High- Dimensional Continuous Control Using Generalized Advantage Esti- 29 mation.” in 4th International Conference on Learning Representations, ICLR, 2016

  30. [36]

    Mirror learning: A unifying framework of policy optimisation,

    J. Grudzien, C. A. S. De Witt, and J. Foerster, “Mirror learning: A unifying framework of policy optimisation,” in International Conference on Machine Learning . PMLR, 2022, pp. 7825–7844

  31. [37]

    Soft Actor- Critic Algorithms and Applications,

    T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V . Kumar, H. Zhu, A. Gupta, P. Abbeel, and S. Levine, “Soft Actor- Critic Algorithms and Applications,” Jan. 2019, arXiv:1812.05905 [cs]

  32. [38]

    Maximum entropy inverse reinforcement learning,

    B. D. Ziebart, A. L. Maas, J. A. Bagnell, and A. K. Dey, “Maximum entropy inverse reinforcement learning,” in Proceedings of the 23rd AAAI conference on artificial intelligence . AAAI Press, 2008, pp. 1433–1438

  33. [39]

    OffCon3: What is state of the art anyway?

    P. J. Ball and S. J. Roberts, “OffCon3: What is state of the art anyway?” Mar. 2021

  34. [40]

    Maximum Entropy RL (Provably) Solves Some Robust RL Problems

    B. Eysenbach and S. Levine, “Maximum Entropy RL (Provably) Solves Some Robust RL Problems.” in 10th International Conference on Learning Representations, ICLR , 2022

  35. [41]

    Essentially Sharp Estimates on the Entropy Regularization Error in Discounted Markov Decision Processes,

    J. M ¨uller and S. Cayci, “Essentially Sharp Estimates on the Entropy Regularization Error in Discounted Markov Decision Processes,” Jun. 2024

  36. [42]

    Gr ¨une and J

    L. Gr ¨une and J. Pannek, Nonlinear Model Predictive Control. Theory and Algorithms, 2nd ed. Springer, 2017

  37. [43]

    Economic receding horizon control without terminal constraints,

    L. Gr ¨une, “Economic receding horizon control without terminal constraints,” Automatica, vol. 49, no. 3, pp. 725–734, 2013. [Online]. Available: https://www.sciencedirect.com/science/article/pii/ S0005109812006024

  38. [44]

    Stabilizing Receding- Horizon control of nonlinear time varying systems,

    G. D. Nicolao, L. Magni, and R. Scattolini, “Stabilizing Receding- Horizon control of nonlinear time varying systems,” IEEE Transactions on Automatic Control , vol. AC-43, no. 7, pp. 1030–1036, 1998

  39. [45]

    Stabilizing nonlinear receding horizon control via a nonquadratic terminal state penalty,

    ——, “Stabilizing nonlinear receding horizon control via a nonquadratic terminal state penalty,” in Symposium on Control, Optimization and Supervision, CESA’96 IMACS Multiconference , Lille, 1996, pp. 185– 187

  40. [46]

    Online NMPC of a looping kite using approximate infinite horizon closed loop costing,

    M. Diehl, L. Magni, and G. D. Nicolao, “Online NMPC of a looping kite using approximate infinite horizon closed loop costing,” in Proceedings of the IFAC Conference on Control Systems Design . Bratislava, Slovak Republic: IFAC, Sep. 2003

  41. [47]

    Efficient NMPC of unstable periodic systems using approximate infinite horizon closed loop costing,

    ——, “Efficient NMPC of unstable periodic systems using approximate infinite horizon closed loop costing,” Annual Reviews in Control, vol. 28, no. 1, pp. 37–45, 2004

  42. [48]

    B. D. O. Anderson and J. B. Moore, Optimal Control - Linear Quadratic Methods. Dover, 1990

  43. [49]

    Nocedal and S

    J. Nocedal and S. J. Wright, Numerical Optimization , 2nd ed., ser. Springer Series in Operations Research and Financial Engineering. Springer, 2006

  44. [50]

    Robust Constraint Satisfaction: Invariant Sets and Predictive Control,

    E. C. Kerrigan, “Robust Constraint Satisfaction: Invariant Sets and Predictive Control,” PhD Thesis, University of Cambridge, UK, 2000

  45. [51]

    Model predictive control with discrete actuators: Theory and application,

    J. B. Rawlings and M. J. Risbeck, “Model predictive control with discrete actuators: Theory and application,” Automatica, vol. 78, 2017

  46. [52]

    Stochastic Model Predictive Control: An Overview and Perspectives for Future Research,

    A. Mesbah, “Stochastic Model Predictive Control: An Overview and Perspectives for Future Research,” IEEE Control Systems Magazine , vol. 36, no. 6, pp. 30–44, Dec. 2016, conference Name: IEEE Control Systems Magazine

  47. [53]

    Kouvaritakis and M

    B. Kouvaritakis and M. Cannon, Model Predictive Control: Classical, Robust and Stochastic , 1st ed. Cham Heidelberg New York Dordrecht London: Springer, Dec. 2015

  48. [54]

    Shapiro, D

    A. Shapiro, D. Dentcheva, and A. Ruszczynski, Lectures on Stochastic Programming: Modelling and Theory . SIAM, 2009

  49. [55]

    The Scenario Approach to Robust Control Design,

    G. C. Calafiore and M. C. Campi, “The Scenario Approach to Robust Control Design,” IEEE Trans. Automat. Control , 2006

  50. [56]

    Robust model predictive control using tubes,

    W. Langson, S. R. I. Chryssochoos, and D. Q. Mayne, “Robust model predictive control using tubes,” Automatica, vol. 40, no. 1, pp. 125–133, 2004

  51. [57]

    Tube-based robust nonlinear model predictive control,

    D. Mayne, E. Kerrigan, E. J. v. Wyk, and P. Falugi, “Tube-based robust nonlinear model predictive control,” International Journal of Robust and Nonlinear Control , vol. 21, pp. 1341–1353, 2011

  52. [58]

    Neural network dynamics for model-based deep reinforcement learning with model- free fine-tuning,

    A. Nagabandi, G. Kahn, R. S. Fearing, and S. Levine, “Neural network dynamics for model-based deep reinforcement learning with model- free fine-tuning,” in IEEE international conference on robotics and automation (ICRA). IEEE, 2018, pp. 7559–7566

  53. [59]

    A dual Newton strategy for tree-sparse quadratic programs and its implementation in the open-source software treeQP,

    D. Kouzoupis, E. Klintberg, G. Frison, S. Gros, and M. Diehl, “A dual Newton strategy for tree-sparse quadratic programs and its implementation in the open-source software treeQP,” International Jounal of Robust and Nonlinear Control , 2019

  54. [60]

    A. B. Kurzhanski and P. Valyi, Ellipsoidal Calculus for Estimation and Control. Birkh ¨auser Boston, 1997

  55. [61]

    Robust Optimization of Dynamic Systems,

    B. Houska, “Robust Optimization of Dynamic Systems,” PhD Thesis, KU Leuven, 2011

  56. [62]

    A Positive Definiteness Preserving Discretization Method for nonlinear Lyapunov Differential Equations,

    J. Gillis and M. Diehl, “A Positive Definiteness Preserving Discretization Method for nonlinear Lyapunov Differential Equations,” in CDC, 2013

  57. [63]

    Robust MPC via min-max differential inequalities,

    M. E. Villanueva, R. Quirynen, M. Diehl, B. Chachuat, and B. Houska, “Robust MPC via min-max differential inequalities,” Automatica, vol. 77, pp. 311–321, Mar. 2017

  58. [64]

    Parameterized Tube Model Predictive Control,

    S. Rakovic, B. Kouvaritakis, M. Cannon, C. Panos, and R. Findeisen, “Parameterized Tube Model Predictive Control,” Automatic Control, IEEE Transactions on , vol. 57, no. 11, pp. 2746–2761, Nov. 2012

  59. [65]

    Robust Model Predictive Control,

    S. V . Rakovi´c, “Robust Model Predictive Control,” in Encyclopedia of Systems and Control, J. Baillieul and T. Samad, Eds. London: Springer London, 2019, pp. 1–11

  60. [66]

    Robust Tube MPC for Linear Systems With Multiplicative Uncertainty,

    J. Fleming, B. Kouvaritakis, and M. Cannon, “Robust Tube MPC for Linear Systems With Multiplicative Uncertainty,” IEEE Trans. Automat. Control, vol. 60, no. 4, 2015

  61. [67]

    Homothetic Tube Model Predictive Control for Nonlinear Systems,

    S. V . Rakovi´c, L. Dai, and Y . Xia, “Homothetic Tube Model Predictive Control for Nonlinear Systems,” IEEE Trans. Automat. Control, vol. 68, no. 8, 2022

  62. [68]

    Configuration- Constrained Tube MPC,

    M. E. Villanueva, M. A. M ¨uller, and B. Houska, “Configuration- Constrained Tube MPC,” Automatica, vol. 163, 2024

  63. [69]

    Ad- justable robust solutions of uncertain linear programs,

    A. Ben-Tal, A. Goryashko, E. Guslitzer, and A. Nemirovski, “Ad- justable robust solutions of uncertain linear programs,” Mathematical Programming, vol. 99, no. 2, pp. 351–376, 2004

  64. [70]

    A receding-horizon regulator for nonlinear systems and a neural approximation,

    T. Parisini and R. Zoppoli, “A receding-horizon regulator for nonlinear systems and a neural approximation,” Automatica, vol. 31, no. 10, pp. 1443–1451, Oct. 1995

  65. [71]

    Open-loop and closed-loop robust optimal control of batch processes using distributional and worst-case analysis,

    Z. K. Nagy and R. D. Braatz, “Open-loop and closed-loop robust optimal control of batch processes using distributional and worst-case analysis,” Journal of Process Control , vol. 14, pp. 411–422, 2004

  66. [72]

    An Efficient Algorithm for Tube-based Robust Nonlinear Optimal Control with Optimal Linear Feedback,

    F. Messerer and M. Diehl, “An Efficient Algorithm for Tube-based Robust Nonlinear Optimal Control with Optimal Linear Feedback,” in Proceedings of the IEEE Conference on Decision and Control (CDC) , 2021

  67. [73]

    Optimization over state feedback policies for robust control with constraints,

    P. J. Goulart, E. C. Kerrigan, and J. M. Maciejowski, “Optimization over state feedback policies for robust control with constraints,” Automatica, vol. 42, pp. 523–533, 2006

  68. [74]

    Min-max feedback model predictive control for constrained linear systems,

    P. O. M. Scokaert and D. Q. Mayne, “Min-max feedback model predictive control for constrained linear systems,” IEEE Transactions on Automatic Control , vol. 43, pp. 1136–1142, 1998

  69. [75]

    Formulation of Closed Loop Min-Max MPC as a Quadrati- cally Constrained Quadratic Program,

    M. Diehl, “Formulation of Closed Loop Min-Max MPC as a Quadrati- cally Constrained Quadratic Program,” IEEE Transactions on Automatic Control, vol. 52, no. 2, pp. 339–343, 2007

  70. [76]

    Mixed-integer optimization-based planning for autonomous racing with obstacles and rewards,

    R. Reiter, M. Kirchengast, D. Watzenig, and M. Diehl, “Mixed-integer optimization-based planning for autonomous racing with obstacles and rewards,” IFAC-PapersOnLine, vol. 54, no. 6, pp. 99–106, Jan. 2021

  71. [77]

    Numerical Methods for Optimal Control of Nonsmooth Dynamical Systems,

    A. Nurkanovi´c, “Numerical Methods for Optimal Control of Nonsmooth Dynamical Systems,” PhD Thesis, University of Freiburg, 2023

  72. [78]

    Dynamic Programming and Suboptimal Control: A Survey from {ADP} to {MPC}*,

    D. P. Bertsekas, “Dynamic Programming and Suboptimal Control: A Survey from {ADP} to {MPC}*,” European Journal of Control, vol. 11, no. 4, pp. 310–334, Jan. 2005

  73. [79]

    On the infinite horizon performance of receding horizon controllers,

    L. Gr ¨une and A. Rantzer, “On the infinite horizon performance of receding horizon controllers,” IEEE Trans. Automat. Control , vol. 53, no. 9, 2008, publisher: University of Bayreuth

  74. [80]

    On the Finite-Time Behavior of Suboptimal Linear Model Predictive Control,

    A. Karapetyan, E. C. Balta, A. Iannelli, and J. Lygeros, “On the Finite-Time Behavior of Suboptimal Linear Model Predictive Control,” Proceedings of the IEEE Conference on Decision and Control (CDC) , 2023

  75. [81]

    Performance Bounds of Model Predictive Control for Unconstrained and Constrained Linear Quadratic Problems and Beyond,

    Y . Li, A. Karapetyan, J. Lygeros, K. H. Johansson, and J. M ˚artensson, “Performance Bounds of Model Predictive Control for Unconstrained and Constrained Linear Quadratic Problems and Beyond,” Proceedings of the IFAC World Congress , 2023

  76. [82]

    An Efficient Method to Estimate the Suboptimality of Affine Controllers,

    M. J. Hadjiyiannis, P. J. Goulart, and D. Kuhn, “An Efficient Method to Estimate the Suboptimality of Affine Controllers,” IEEE Trans. Automat. Control, vol. 56, no. 12, 2011

  77. [83]

    Stochastic Control for Small Noise Intensities,

    W. H. Fleming, “Stochastic Control for Small Noise Intensities,” SIAM J. Control, vol. 9, no. 3, 1971

  78. [84]

    Fourth-order suboptimality of nominal model predictive control in the presence of uncertainty,

    F. Messerer, K. Baumg ¨artner, S. Lucia, and M. Diehl, “Fourth-order suboptimality of nominal model predictive control in the presence of uncertainty,” arXiv, 2024

  79. [85]

    Bounded-Regret MPC via Perturbation Analysis: Prediction Error, Constraints, and Nonlinearity,

    Y . Lin, Y . Hu, G. Qu, T. Li, and A. Wierman, “Bounded-Regret MPC via Perturbation Analysis: Prediction Error, Constraints, and Nonlinearity,” NeurIPS, 2022

  80. [86]

    A Quasi-Infinite Horizon Nonlinear Model Predictive Control Scheme with Guaranteed Stability,

    H. Chen and F. Allg ¨ower, “A Quasi-Infinite Horizon Nonlinear Model Predictive Control Scheme with Guaranteed Stability,” Automatica, vol. 34, no. 10, pp. 1205–1218, 1998. 30

  81. [87]

    Con- strained model predictive control: Stability and optimality,

    D. Q. Mayne, J. B. Rawlings, C. V . Rao, and P. O. M. Scokaert, “Con- strained model predictive control: Stability and optimality,” Automatica, vol. 26, no. 6, pp. 789–814, 2000

  82. [88]

    Input-to-State Stability: A Unifying Framework for Robust Model Predictive Control,

    D. Limon and others, “Input-to-State Stability: A Unifying Framework for Robust Model Predictive Control,” in Nonlinear Model Predictive Control. Lecture Notes in Control and Information Sciences , F. A. L. Magni, D. M. Raimondo, Ed. Springer, Berlin, Heidelberg, 2009, vol. 384

  83. [89]

    NMPC without terminal constraints,

    L. Gr ¨une, “NMPC without terminal constraints,” IFAC Proceedings Volumes, vol. 45, no. 17, pp. 1–13, 2012, publisher: Elsevier

  84. [90]

    Oops! I cannot do it again: Testing for recursive feasibility in MPC,

    J. L ¨ofberg, “Oops! I cannot do it again: Testing for recursive feasibility in MPC,” Automatica, vol. 48, no. 3, 2012

  85. [91]

    An apologia for stabilising terminal conditions in model predictive control,

    D. Mayne, “An apologia for stabilising terminal conditions in model predictive control,” Internat. J. Control , vol. 11, 2013

  86. [92]

    Discrete-time Stability with Perturbations: Application to Model Predictive Control,

    P. Scokaert, J. Rawlings, and E. Meadows, “Discrete-time Stability with Perturbations: Application to Model Predictive Control,” Automatica, vol. 33, no. 3, pp. 463–470, 1997

  87. [93]

    On the Robustness of Receding-Horizon Control with Terminal Constraints,

    G. D. Nicolao, L. Magni, and R. Scattolini, “On the Robustness of Receding-Horizon Control with Terminal Constraints,” IEEE Trans. Automat. Control, vol. 41, no. 3, 1996

  88. [94]

    Stability margins of nonlinear receding- horizon control via inverse optimality,

    L. Magni and R. Sepulchre, “Stability margins of nonlinear receding- horizon control via inverse optimality,” Systems & Control Letters , vol. 32, pp. 241–245, 1997

  89. [95]

    Inherent Stochastic Robustness of Model Predictive Control to Large and Infrequent Disturbances,

    R. D. McAllister and J. B. Rawlings, “Inherent Stochastic Robustness of Model Predictive Control to Large and Infrequent Disturbances,” IEEE Trans. Automat. Control , vol. 67, no. 10, 2022

  90. [96]

    Examples when nonlinear model predictive control is nonrobust,

    G. Grimm, M. J. Messina, S. E. Tuna, and A. R. Teel, “Examples when nonlinear model predictive control is nonrobust,” Automatica, vol. 40, pp. 1729–1738, 2004

  91. [97]

    Inherent robustness properties of quasi-infinite horizon nonlinear model predictive control,

    S. Yu, M. Reble, H. Chen, and F. Allg ¨ower, “Inherent robustness properties of quasi-infinite horizon nonlinear model predictive control,” Automatica, vol. 50, 2014

  92. [98]

    On the Inherent Distributional Robustness of Stochastic and Nominal Model Predictive Control,

    R. D. McAllister and J. B. Rawlings, “On the Inherent Distributional Robustness of Stochastic and Nominal Model Predictive Control,” IEEE Trans. Automat. Control, vol. 69, no. 2, 2024

  93. [99]

    Suboptimal Model Predictive Control (Feasibility Implies Stability),

    P. O. M. Scokaert, D. Q. Mayne, and J. Rawlings, “Suboptimal Model Predictive Control (Feasibility Implies Stability),” IEEE Transactions on Automatic Control , vol. 44, no. 3, pp. 648–654, 1999

  94. [100]

    Nominal Stability of the Real-Time Iteration Scheme for Nonlinear Model Predictive Control,

    M. Diehl, R. Findeisen, F. Allg ¨ower, H. G. Bock, and J. P. Schl ¨oder, “Nominal Stability of the Real-Time Iteration Scheme for Nonlinear Model Predictive Control,” IEE Proc.-Control Theory Appl. , vol. 152, no. 3, pp. 296–308, 2005, publisher: IEE

  95. [101]

    Conditions under which suboptimal nonlinear MPC is inherently robust,

    G. Pannocchia, J. Rawlings, and S. Wright, “Conditions under which suboptimal nonlinear MPC is inherently robust,” System & Control Letters, vol. 60, no. 9, pp. 747–755, 2011

  96. [102]

    On the inherent robustness of optimal and suboptimal nonlinear MPC,

    D. A. Allan, C. N. Bates, M. J. Risbeck, and J. B. Rawlings, “On the inherent robustness of optimal and suboptimal nonlinear MPC,” Systems & Control Letters , vol. 106, pp. 68–78, 2017

  97. [103]

    Randomized model predictive control for robot navigation,

    J. L. Piovesan and H. G. Tanner, “Randomized model predictive control for robot navigation,” in 2009 IEEE International Conference on Robotics and Automation . IEEE, 2009, pp. 94–99

  98. [104]

    Optimization of computer simulation models with rare events,

    R. Y . Rubinstein, “Optimization of computer simulation models with rare events,” European Journal of Operational Research , vol. 99, no. 1, pp. 89–112, 1997, publisher: Elsevier

  99. [105]

    Cross-entropy motion planning,

    M. Kobilarov, “Cross-entropy motion planning,” The International Journal of Robotics Research , vol. 31, no. 7, pp. 855–871, 2012, publisher: SAGE Publications Sage UK: London, England

  100. [106]

    Deep reinforce- ment learning in a handful of trials using probabilistic dynamics models,

    K. Chua, R. Calandra, R. McAllister, and S. Levine, “Deep reinforce- ment learning in a handful of trials using probabilistic dynamics models,” Advances in neural information processing systems , vol. 31, 2018

  101. [107]

    Linear Theory for Control of Nonlinear Stochastic Systems,

    H. J. Kappen, “Linear Theory for Control of Nonlinear Stochastic Systems,” Physical Review Letters , vol. 95, no. 20, p. 200201, Nov. 2005, publisher: American Physical Society

  102. [108]

    Relative entropy and free energy dualities: Connections to Path Integral and KL control,

    E. A. Theodorou and E. Todorov, “Relative entropy and free energy dualities: Connections to Path Integral and KL control,” in IEEE 51st IEEE Conference on Decision and Control (CDC) , Dec. 2012, pp. 1466–1473, iSSN: 0743-1546

  103. [109]

    Nonlinear Stochastic Control and Information The- oretic Dualities: Connections, Interdependencies and Thermodynamic Interpretations,

    E. A. Theodorou, “Nonlinear Stochastic Control and Information The- oretic Dualities: Connections, Interdependencies and Thermodynamic Interpretations,” Entropy, vol. 17, no. 5, pp. 3352–3375, May 2015, number: 5 Publisher: Multidisciplinary Digital Publishing Institute

  104. [110]

    Model Predictive Path Integral Control: From Theory to Parallel Computation,

    G. Williams, A. Aldrich, and E. A. Theodorou, “Model Predictive Path Integral Control: From Theory to Parallel Computation,” Journal of Guidance, Control, and Dynamics , vol. 40, no. 2, pp. 344–357, 2017

  105. [111]

    AutoRally: An Open Platform for Aggressive Autonomous Driving,

    B. Goldfain, P. Drews, C. You, M. Barulic, O. Velev, P. Tsiotras, and J. M. Rehg, “AutoRally: An Open Platform for Aggressive Autonomous Driving,” IEEE Control Systems Magazine , vol. 39, no. 1, pp. 26–55, Feb. 2019, conference Name: IEEE Control Systems Magazine

  106. [112]

    TorchRL: A data-driven decision-making library for PyTorch,

    A. Bou, M. Bettini, S. Dittert, V . Kumar, S. Sodhani, X. Yang, G. D. Fabritiis, and V . Moens, “TorchRL: A data-driven decision-making library for PyTorch,” Nov. 2023, arXiv:2306.00577 [cs]

  107. [113]

    MPPI-Generic: A CUDA Library for Stochastic Optimization,

    B. Vlahov, J. Gibson, M. Gandhi, and E. A. Theodorou, “MPPI-Generic: A CUDA Library for Stochastic Optimization,” Sep. 2024

  108. [114]

    Sequential Monte Carlo for model predictive control,

    N. Kantas, J. Maciejowski, and A. Lecchini-Visintini, “Sequential Monte Carlo for model predictive control,” Nonlinear model predictive control: Towards new challenging applications , pp. 263–273, 2009, publisher: Springer

  109. [115]

    A survey of numerical methods for optimal control,

    A. V . Rao, “A survey of numerical methods for optimal control,” Advances in the astronautical Sciences , vol. 135, no. 1, pp. 497–528, 2009, publisher: Univelt, Inc

  110. [116]

    Introduction to Model Based Optimization of Chemical Processes on Moving Horizons,

    T. Binder, L. Blank, H. G. Bock, R. Bulirsch, W. Dahmen, M. Diehl, T. Kronseder, W. Marquardt, J. P. Schl ¨oder, and O. v. Stryk, “Introduction to Model Based Optimization of Chemical Processes on Moving Horizons,” in Online Optimization of Large Scale Systems: State of the Ar...

  111. [117]

    A Multiple Shooting Algorithm for Direct Solution of Optimal Control Problems,

    H. G. Bock and K. J. Plitt, “A Multiple Shooting Algorithm for Direct Solution of Optimal Control Problems,” in Proceedings of the IFAC World Congress. Pergamon Press, 1984, pp. 242–247

  112. [118]

    An accelerated dual gradient-projection algorithm for embedded linear model predictive control,

    P. Patrinos and A. Bemporad, “An accelerated dual gradient-projection algorithm for embedded linear model predictive control,” IEEE Transac- tions on Automatic Control , vol. 59, no. 1, pp. 18–33, 2013, publisher: IEEE

  113. [119]

    Embedded online optimization for model predictive control at megahertz rates,

    J. L. Jerez, P. J. Goulart, S. Richter, G. A. Constantinides, E. C. Kerrigan, and M. Morari, “Embedded online optimization for model predictive control at megahertz rates,” IEEE Transactions on Automatic Control , vol. 59, no. 12, pp. 3238–3251, 2014, publisher: IEEE

  114. [120]

    The gradient based nonlinear model predictive control software GRAMPC,

    B. K ¨apernick and K. Graichen, “The gradient based nonlinear model predictive control software GRAMPC,” in 2014 European Control Conference (ECC). IEEE, 2014, pp. 1170–1175

  115. [121]

    A software framework for embedded nonlinear model predictive control using a gradient-based augmented Lagrangian approach (GRAMPC),

    T. Englert, A. V ¨olz, F. Mesmer, S. Rhein, and K. Graichen, “A software framework for embedded nonlinear model predictive control using a gradient-based augmented Lagrangian approach (GRAMPC),” Optimization and Engineering , vol. 20, no. 3, pp. 769–809, Sep. 2019

  116. [122]

    Decomposition via ADMM for scenario-based model predictive control,

    J. Kang, A. U. Raghunathan, and S. Di Cairano, “Decomposition via ADMM for scenario-based model predictive control,” in American Control Conference (ACC). IEEE, 2015, pp. 1246–1251

  117. [123]

    Efficient Numerical Methods for Nonlinear MPC and Moving Horizon Estimation,

    M. Diehl, H. J. Ferreau, and N. Haverbeke, “Efficient Numerical Methods for Nonlinear MPC and Moving Horizon Estimation,” in Nonlinear model predictive control , ser. Lecture Notes in Control and Information Sciences, L. Magni, M. D. Raimondo, and F. Allg ¨ower, Eds. Springer,...

  118. [124]

    Application of Interior- Point Methods to Model Predictive Control,

    C. V . Rao, S. J. Wright, and J. B. Rawlings, “Application of Interior- Point Methods to Model Predictive Control,” Journal of Optimization Theory and Applications , vol. 99, pp. 723–757, 1998

  119. [125]

    Active-Set based Inexact Interior Point QP Solver for Model Predictive Control,

    J. Frey, S. D. Cairano, and R. Quirynen, “Active-Set based Inexact Interior Point QP Solver for Model Predictive Control,” in Proceedings of the IFAC World Congress , 2020

  120. [126]

    Frison and M

    G. Frison and M. Diehl, “HPIPM: a high-performance quadratic programming framework for model predictive control **This research was supported by the German Federal Ministry for Economic Affairs and Energy (BMWi) via eco4wind (0324125B) and DyConPV (0324166B), and by DFG via Re...

  121. [127]

    Fatrop: A fast constrained optimal control problem solver for robot trajectory optimization and control,

    L. Vanroye, A. Sathya, J. De Schutter, and W. Decr ´e, “Fatrop: A fast constrained optimal control problem solver for robot trajectory optimization and control,” in 2023 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2023, pp. 10 036– 10 043

  122. [128]

    Recent advances in quadratic programming algorithms for nonlinear model predictive control,

    D. Kouzoupis, G. Frison, A. Zanelli, and M. Diehl, “Recent advances in quadratic programming algorithms for nonlinear model predictive control,” Vietnam Journal of Mathematics , vol. 46, no. 4, pp. 863–882, 2018

  123. [129]

    ACADO toolkit—An open- source framework for automatic control and dynamic optimization,

    B. Houska, H. J. Ferreau, and M. Diehl, “ACADO toolkit—An open- source framework for automatic control and dynamic optimization,” Optimal Control Applications and Methods , vol. 32, no. 3, pp. 298–312, 2011, eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1002/oca.939

  124. [130]

    acados—a modular open-source framework for fast embedded optimal control,

    R. Verschueren, G. Frison, D. Kouzoupis, J. Frey, N. v. Duijkeren, A. Zanelli, B. Novoselnik, T. Albin, R. Quirynen, and M. Diehl, “acados—a modular open-source framework for fast embedded optimal control,” Mathematical Programming Computation , vol. 14, no. 1, pp. 147–183, Ma...

  125. [131]

    Diehl, Real-Time Optimization for Large Scale Nonlinear Processes , ser

    M. Diehl, Real-Time Optimization for Large Scale Nonlinear Processes , ser. Fortschritt-Berichte VDI Reihe 8, Meß-, Steuerungs- und Regelung- stechnik. D ¨usseldorf: VDI Verlag, 2002, vol. 920

  126. [132]

    Constrained Optimal Feedback Control of Systems Governed by Large Differential Algebraic Equations,

    H. G. Bock, M. Diehl, E. A. Kostina, and J. P. Schl ¨oder, “Constrained Optimal Feedback Control of Systems Governed by Large Differential Algebraic Equations,” in Real-Time and Online PDE-Constrained Optimization. SIAM, 2007, pp. 3–22

  127. [133]

    From Linear to Nonlinear MPC: bridging the gap via the Real-Time Iteration,

    S. Gros, M. Zanon, R. Quirynen, A. Bemporad, and M. Diehl, “From Linear to Nonlinear MPC: bridging the gap via the Real-Time Iteration,” International Journal of Control , 2016

  128. [134]

    C. F. Gauss, Theoria motus corporum coelestium in sectionibus conicis solem ambientium. Perthes et Besser, 1809, vol. 7

  129. [135]

    Recent Advances in Parameter Identification Techniques for ODE,

    H. G. Bock, “Recent Advances in Parameter Identification Techniques for ODE,” in Numerical Treatment of Inverse Problems in Differential and Integral Equations . Birk \-h¨au\-ser, 1983, pp. 95–121

  130. [136]

    Survey of Sequential Con- vex Programming and Generalized Gauss-Newton Methods,

    F. Messerer, K. Baumg¨artner, and M. Diehl, “Survey of Sequential Con- vex Programming and Generalized Gauss-Newton Methods,” ESAIM: Proceedings and Surveys, vol. 71, pp. 64–88, 2021

  131. [137]

    A Second-order Gradient Method for Determining Optimal Trajectories of Non-linear Discrete-time Systems,

    D. Mayne, “A Second-order Gradient Method for Determining Optimal Trajectories of Non-linear Discrete-time Systems,” Int. J. Control, vol. 3, no. 1, pp. 85–96, 1966

  132. [138]

    Control-Limited Differential Dynamic Programming,

    Y . Tassa, N. Mansard, and E. Todorov, “Control-Limited Differential Dynamic Programming,” in IEEE International Conference on Robotics and Automation, 2014

  133. [139]

    Iterative Linear Quadratic Regulator Design for Nonlinear Biological Movement Systems,

    W. Li and E. Todorov, “Iterative Linear Quadratic Regulator Design for Nonlinear Biological Movement Systems,” in Proceedings of the 1st International Conference on Informatics in Control, Automation and Robotics, 2004

  134. [140]

    A generalized iterative LQG method for locally- optimal feedback control of constrained nonlinear stochastic systems,

    E. Todorov and W. Li, “A generalized iterative LQG method for locally- optimal feedback control of constrained nonlinear stochastic systems,” in Proceedings of the American Control Conference (ACC) , 2005

  135. [141]

    Squash-box feasibility driven differential dynamic programming,

    J. Marti-Saumell, J. Sol `a, C. Mastalli, and A. Santamaria-Navarro, “Squash-box feasibility driven differential dynamic programming,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS). IEEE, 2020, pp. 7637–7644

  136. [142]

    An efficient sequential linear quadratic algorithm for solving unconstrained nonlinear optimal control problems,

    A. Sideris and J. Bodrow, “An efficient sequential linear quadratic algorithm for solving unconstrained nonlinear optimal control problems,” IEEE Transactions on Automatic Control, vol. 50, no. 12, pp. 2043–2047, 2005

  137. [143]

    A Unified Local Conver- gence Analysis of Differential Dynamic Programming, Direct Single Shooting, and Direct Multiple Shooting,

    K. Baumg ¨artner, F. Messerer, and M. Diehl, “A Unified Local Conver- gence Analysis of Differential Dynamic Programming, Direct Single Shooting, and Direct Multiple Shooting,” in Proceedings of the European Control Conference (ECC), 2023

  138. [144]

    The explicit solution of model predictive control via multiparametric quadratic programming,

    A. Bemporad, M. Morari, V . Dua, and E. N. Pistikopoulos, “The explicit solution of model predictive control via multiparametric quadratic programming,” in Proceedings of the 2000 American Control Conference. ACC (IEEE Cat. No. 00CH36334) , vol. 2. IEEE, 2000, pp. 872–876

  139. [145]

    Grancharova and T

    A. Grancharova and T. A. Johansen, Explicit Nonlinear Model Predictive Control: Theory and Applications , ser. Lecture Notes in Control and Information Sciences. Berlin, Heidelberg: Springer, 2012, vol. 429

  140. [146]

    Tree-Sparse Convex Programs,

    M. C. Steinbach, “Tree-Sparse Convex Programs,” Mathematical Methods of Operations Research , vol. 56, no. 3, pp. 347–376, 2002

  141. [147]

    An improved dual Newton strategy for scenario-tree MPC,

    E. Klintberg, J. Dahl, J. Fredriksson, and S. Gros, “An improved dual Newton strategy for scenario-tree MPC,” in CDC, 2016, pp. 3675–3681

  142. [148]

    A high- performance Riccati based solver for tree-structured quadratic programs,

    G. Frison, D. Kouzoupis, M. Diehl, and J. B. Jørgensen, “A high- performance Riccati based solver for tree-structured quadratic programs,” in IFAC, vol. 50, 2017, pp. 14 399–14 405, issue: 1

  143. [149]

    Structure-exploiting numerical methods for tree-sparse optimal control problems,

    D. Kouzoupis, “Structure-exploiting numerical methods for tree-sparse optimal control problems,” PhD Thesis, University of Freiburg, 2019

  144. [150]

    Scaling up Gaussian Belief Space Planning Through Covariance-Free Trajectory Optimization and Automatic Differentiation,

    S. Patil, G. Kahn, M. Laskey, J. Schulman, K. Goldberg, and P. Abbeel, “Scaling up Gaussian Belief Space Planning Through Covariance-Free Trajectory Optimization and Automatic Differentiation,” in Algorithmic Foundations of Robotics XI: Selected Contributions of the Eleventh I...

  145. [151]

    Inexact adjoint-based SQP algorithm for real-time stochastic nonlinear MPC,

    X. Feng, S. Di Cairano, and R. Quirynen, “Inexact adjoint-based SQP algorithm for real-time stochastic nonlinear MPC,” IFAC-PapersOnLine, vol. 53, no. 2, pp. 6529–6535, 2020, publisher: Elsevier

  146. [152]

    Zero-Order Robust Nonlinear Model Predictive Control with Ellipsoidal Uncertainty Sets,

    A. Zanelli, J. Frey, F. Messerer, and M. Diehl, “Zero-Order Robust Nonlinear Model Predictive Control with Ellipsoidal Uncertainty Sets,” Proceedings of the IFAC Conference on Nonlinear Model Predictive Control (NMPC), 2021

  147. [153]

    Efficient Zero-Order Robust Optimization for Real-Time Model Pre- dictive Control with acados,

    J. Frey, Y . Gao, F. Messerer, A. Lahr, M. N. Zeilinger, and M. Diehl, “Efficient Zero-Order Robust Optimization for Real-Time Model Pre- dictive Control with acados,” in Proceedings of the European Control Conference (ECC), 2024

  148. [154]

    Efficient robust optimization for robust control with constraints,

    P. J. Goulart, E. C. Kerrigan, and D. Ralph, “Efficient robust optimization for robust control with constraints,” Mathematical Programming, vol. 114, no. 1, pp. 115–147, 2008, publisher: Springer

  149. [155]

    Fast System Level Synthesis: Robust Model Predictive Con- trol using Riccati Recursions,

    A. P. Leeman, J. K ¨ohler, F. Messerer, A. Lahr, M. Diehl, and M. N. Zeilinger, “Fast System Level Synthesis: Robust Model Predictive Con- trol using Riccati Recursions,” in Proceedings of the IFAC Conference on Nonlinear Model Predictive Control (NMPC) , 2024

  150. [156]

    Sam- pling Complexity of Path Integral Methods for Trajectory Optimization,

    H.-J. Yoon, C. Tao, H. Kim, N. Hovakimyan, and P. V oulgaris, “Sam- pling Complexity of Path Integral Methods for Trajectory Optimization,” in 2022 American Control Conference (ACC) , 2022, pp. 3482–3487

  151. [157]

    Model predictive control: past, present and future,

    M. Morari and J. H. Lee, “Model predictive control: past, present and future,” Computers & Chemical Engineering , vol. 23, no. 4, pp. 667–682, 1999

  152. [158]

    On Average Performance and Stability of Economic Model Predictive Control,

    D. Angeli, R. Amrit, and J. B. Rawlings, “On Average Performance and Stability of Economic Model Predictive Control,” IEEE Transactions on Automatic Control , vol. 57, no. 7, pp. 1615–1626, 2012

  153. [159]

    Reinforcement Learning with Long Short-Term Memory,

    B. Bakker, “Reinforcement Learning with Long Short-Term Memory,” in Advances in Neural Information Processing Systems , vol. 14. MIT Press, 2001

  154. [160]

    Regularized and Distributionally Robust Data-Enabled Predictive Control,

    J. Coulson, J. Lygeros, and F. D ¨orfler, “Regularized and Distributionally Robust Data-Enabled Predictive Control,” in IEEE 58th Conference on Decision and Control (CDC) , Dec. 2019, pp. 2696–2701, iSSN: 2576-2370

  155. [161]

    Model- predictive control and reinforcement learning in multi-energy system case studies,

    G. Ceusters, R. C. Rodr ´ıguez, A. B. Garc ´ıa, R. Franke, G. Deconinck, L. Helsen, A. Now ´e, M. Messagie, and L. R. Camargo, “Model- predictive control and reinforcement learning in multi-energy system case studies,” Applied Energy, vol. 303, p. 117634, Dec. 2021

  156. [163]

    Comparison of online and offline deep reinforcement learning with model predictive control for thermal energy management,

    S. Brandi, M. Fiorentini, and A. Capozzoli, “Comparison of online and offline deep reinforcement learning with model predictive control for thermal energy management,” Automation in Construction , vol. 135, p. 104128, Mar. 2022

  157. [164]

    Champion-level drone racing using deep reinforcement learning,

    E. Kaufmann, L. Bauersfeld, A. Loquercio, M. M ¨uller, V . Koltun, and D. Scaramuzza, “Champion-level drone racing using deep reinforcement learning,” Nature, vol. 620, no. 7976, pp. 982–987, Aug. 2023, number: 7976 Publisher: Nature Publishing Group

  158. [165]

    Recurrent networks, hidden states and beliefs in partially observable environments,

    G. Lambrechts, A. Bolland, and D. Ernst, “Recurrent networks, hidden states and beliefs in partially observable environments,” Transactions on Machine Learning Research , May 2022

  159. [166]

    An Efficient Method for the Joint Estimation of System Parameters and Noise Covariances for Linear Time-Variant Systems,

    L. Simpson, A. Ghezzi, J. Asprion, and M. Diehl, “An Efficient Method for the Joint Estimation of System Parameters and Noise Covariances for Linear Time-Variant Systems,” in 62nd IEEE Conference on Decision and Control (CDC) , 2023, pp. 4524–4529

  160. [167]

    Cautious Model Predictive Control Using Gaussian Process Regression,

    L. Hewing, J. Kabzan, and M. N. Zeilinger, “Cautious Model Predictive Control Using Gaussian Process Regression,” IEEE Transactions on Control Systems Technology, vol. 28, no. 6, pp. 2736–2743, Nov. 2020, conference Name: IEEE Transactions on Control Systems Technology

  161. [168]

    Reinforcement learning in robotic applications: a comprehensive survey,

    B. Singh, R. Kumar, and V . Singh, “Reinforcement learning in robotic applications: a comprehensive survey,” Artificial Intelligence Review , vol. 55, Feb. 2022

  162. [169]

    Challenges of real-world reinforcement learning: definitions, benchmarks and analysis,

    G. Dulac-Arnold, N. Levine, D. J. Mankowitz, J. Li, C. Paduraru, S. Gowal, and T. Hester, “Challenges of real-world reinforcement learning: definitions, benchmarks and analysis,” Machine Learning , vol. 110, no. 9, pp. 2419–2468, Sep. 2021

  163. [170]

    Demonstrating a walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning,

    L. Smith, I. Kostrikov, and S. Levine, “Demonstrating a walk in the park: Learning to walk in 20 minutes with model-free reinforcement learning,” Robotics: Science and Systems (RSS) Demo , vol. 2, no. 3, p. 4, 2023

  164. [171]

    Making deep q-learning methods robust to time discretization,

    C. Tallec, L. Blier, and Y . Ollivier, “Making deep q-learning methods robust to time discretization,” in International Conference on Machine Learning. PMLR, 2019, pp. 6096–6104

  165. [172]

    Ljung, System Identification: Theory for the User

    L. Ljung, System Identification: Theory for the User . Prentice Hall PTR, 1999, google-Books-ID: nHFoQgAACAAJ

  166. [173]

    Reinforcement learning in the presence of rare events,

    J. Frank, S. Mannor, and D. Precup, “Reinforcement learning in the presence of rare events,” in Proceedings of the 25th international conference on Machine learning , 2008, pp. 336–343

  167. [174]

    Robust Reinforcement Learning,

    J. Morimoto and K. Doya, “Robust Reinforcement Learning,” in Advances in Neural Information Processing Systems , vol. 13. MIT Press, 2000

  168. [175]

    Robust Reinforcement Learning: A Review of Foundations and Recent Advances,

    J. Moos, K. Hansel, H. Abdulsamad, S. Stark, D. Clever, and J. Peters, “Robust Reinforcement Learning: A Review of Foundations and Recent Advances,” Machine Learning and Knowledge Extraction , vol. 4, no. 1, 32 pp. 276–315, Mar. 2022, number: 1 Publisher: Multidisciplinary Dig...

  169. [176]

    Model-agnostic meta-learning for fast adaptation of deep networks,

    C. Finn, P. Abbeel, and S. Levine, “Model-agnostic meta-learning for fast adaptation of deep networks,” in International conference on machine learning. PMLR, 2017, pp. 1126–1135

  170. [177]

    Language models are few-shot learners,

    T. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, and A. Askell, “Language models are few-shot learners,” Advances in neural information processing systems, vol. 33, pp. 1877–1901, 2020

  171. [178]

    Distri- butionally Robust Control of Constrained Stochastic Systems,

    B. P. G. Van Parys, D. Kuhn, P. J. Goulart, and M. Morari, “Distri- butionally Robust Control of Constrained Stochastic Systems,” IEEE Transactions on Automatic Control , vol. 61, no. 2, pp. 430–442, 2016

  172. [179]

    Economic MPC of Markov Decision Pro- cesses: Dissipativity in undiscounted infinite-horizon optimal control,

    S. Gros and M. Zanon, “Economic MPC of Markov Decision Pro- cesses: Dissipativity in undiscounted infinite-horizon optimal control,” Automatica, vol. 146, p. 110602, Dec. 2022

  173. [180]

    Altman, Constrained Markov Decision Processes , 1st, Ed

    E. Altman, Constrained Markov Decision Processes , 1st, Ed. New York: Routledge, 1999

  174. [181]

    Deep Inverse Q-learning with Constraints,

    G. Kalweit, M. Huegle, M. Werling, and J. Boedecker, “Deep Inverse Q-learning with Constraints,” in Advances in Neural Information Processing Systems, vol. 33. Curran Associates, Inc., 2020, pp. 14 291– 14 302

  175. [182]

    Benchmarking Safe Exploration in Deep Reinforcement Learning,

    A. Ray, J. Achiam, and D. Amodei, “Benchmarking Safe Exploration in Deep Reinforcement Learning,” 2019. [Online]. Available: https://cdn.openai.com/safexp-short.pdf

  176. [183]

    Constrained Policy Optimization,

    J. Achiam, D. Held, A. Tamar, and P. Abbeel, “Constrained Policy Optimization,” in Proceedings of the 34th International Conference on Machine Learning. PMLR, Jul. 2017, pp. 22–31, iSSN: 2640-3498

  177. [184]

    End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks,

    R. Cheng, G. Orosz, R. M. Murray, and J. W. Burdick, “End-to-end safe reinforcement learning through barrier functions for safety-critical continuous control tasks,” in Proceedings of the AAAI conference on artificial intelligence, vol. 33, 2019, pp. 3387–3395, issue: 01

  178. [185]

    A predictive safety filter for learning-based control of constrained nonlinear dynamical systems,

    K. P. Wabersich and M. N. Zeilinger, “A predictive safety filter for learning-based control of constrained nonlinear dynamical systems,” Automatica, vol. 129, p. 109597, Jul. 2021

  179. [186]

    State Augmented Constrained Reinforcement Learning: Overcoming the Limitations of Learning With Rewards,

    M. Calvo-Fullana, S. Paternain, L. F. O. Chamon, and A. Ribeiro, “State Augmented Constrained Reinforcement Learning: Overcoming the Limitations of Learning With Rewards,” IEEE Transactions on Automatic Control, vol. 69, no. 7, pp. 4275–4290, Jul. 2024, conference Name: IEEE T...

  180. [187]

    Saute RL: Almost Surely Safe Reinforcement Learning Using State Augmentation,

    A. Sootla, A. I. Cowen-Rivers, T. Jafferjee, Z. Wang, D. H. Mguni, J. Wang, and H. Ammar, “Saute RL: Almost Surely Safe Reinforcement Learning Using State Augmentation,” in Proceedings of the 39th International Conference on Machine Learning . PMLR, Jun. 2022, pp. 20 423–20 44...

  181. [188]

    Constrained Reinforcement Learning with Smoothed Log Barrier Function,

    B. Zhang, Y . Zhang, L. Frison, T. Brox, and J. B ¨odecker, “Constrained Reinforcement Learning with Smoothed Log Barrier Function,” Mar. 2024, arXiv:2403.14508 [cs]

  182. [189]

    Computational complexity certification for real-time MPC with input constraints based on the fast gradient method,

    S. Richter, C. N. Jones, and M. Morari, “Computational complexity certification for real-time MPC with input constraints based on the fast gradient method,” IEEE Transactions on Automatic Control , vol. 57, no. 6, pp. 1391–1403, 2011, publisher: IEEE

  183. [190]

    Real-Time Certified MPC : Reliable Active-Set QP Solvers,

    D. Arnstr ¨om, “Real-Time Certified MPC : Reliable Active-Set QP Solvers,” 2023, publisher: Link ¨oping University Electronic Press

  184. [191]

    On the importance of hyperparameter optimization for model-based reinforcement learning,

    B. Zhang, R. Rajan, L. Pineda, N. Lambert, A. Biedenkapp, K. Chua, F. Hutter, and R. Calandra, “On the importance of hyperparameter optimization for model-based reinforcement learning,” in International Conference on Artificial Intelligence and Statistics . PMLR, 2021, pp. 4015–4023

  185. [192]

    Rethinking Out-of-Distribution Detection for Reinforcement Learning: Advancing Methods for Evaluation and Detection,

    L. Nasvytis, K. Sandbrink, J. Foerster, T. Franzmeyer, and C. Schroeder de Witt, “Rethinking Out-of-Distribution Detection for Reinforcement Learning: Advancing Methods for Evaluation and Detection,” in Proceedings of the 23rd International Conference on Autonomous Agents and ...

  186. [193]

    A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,

    S. Ross, G. Gordon, and D. Bagnell, “A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning,” inProceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, ser. Proceedings of Machine Learning Research, G....

  187. [194]

    Rapidly adaptable legged robots via evolutionary meta- learning,

    X. Song, Y . Yang, K. Choromanski, K. Caluwaerts, W. Gao, C. Finn, and J. Tan, “Rapidly adaptable legged robots via evolutionary meta- learning,” in 2020 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) . IEEE, 2020, pp. 3769–3776

  188. [195]

    Gurobi Optimizer Reference Manual,

    Gurobi Optimization, LLC, “Gurobi Optimizer Reference Manual,”

  189. [196]

    Quantitative comparison of reinforcement learning and data- driven model predictive control for chemical and biological processes,

    T. H. Oh, “Quantitative comparison of reinforcement learning and data- driven model predictive control for chemical and biological processes,” Computers & Chemical Engineering , vol. 181, p. 108558, Feb. 2024

  190. [197]

    Reinforcement Learning Versus Model Predictive Control: A Comparison on a Power System Problem,

    D. Ernst, M. Glavic, F. Capitanescu, and L. Wehenkel, “Reinforcement Learning Versus Model Predictive Control: A Comparison on a Power System Problem,” IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), vol. 39, no. 2, pp. 517–529, Apr. 2009, conference ...

  191. [198]

    Comparison of reinforcement learning and model predictive control for building energy system optimization,

    D. Wang, W. Zheng, Z. Wang, Y . Wang, X. Pang, and W. Wang, “Comparison of reinforcement learning and model predictive control for building energy system optimization,” Applied Thermal Engineering , vol. 228, p. 120430, Jun. 2023

  192. [199]

    Comparison of Deep Reinforcement Learning and Model Predictive Control for Real- Time Depth Optimization of a Lifting Surface Controlled Ocean Current Turbine,

    A. Hasankhani, Y . Tang, J. VanZwieten, and C. Sultan, “Comparison of Deep Reinforcement Learning and Model Predictive Control for Real- Time Depth Optimization of a Lifting Surface Controlled Ocean Current Turbine,” in IEEE Conference on Control Technology and Applications (C...

  193. [200]

    Iteratively Extending Time Horizon Reinforcement Learning,

    D. Ernst, P. Geurts, and L. Wehenkel, “Iteratively Extending Time Horizon Reinforcement Learning,” in Machine Learning: ECML 2003 , ser. Lecture Notes in Computer Science, N. Lavra ˇc, D. Gamberger, H. Blockeel, and L. Todorovski, Eds. Berlin, Heidelberg: Springer, 2003, pp. 96–107

  194. [201]

    V12. 1: User’s Manual for CPLEX,

    I. I. Cplex, “V12. 1: User’s Manual for CPLEX,” International Business Machines Corporation, vol. 46, no. 53, p. 157, 2009

  195. [202]

    Comparison of Deep Reinforcement Learning and Model Predictive Control for Adaptive Cruise Control,

    Y . Lin, J. McPhee, and N. L. Azad, “Comparison of Deep Reinforcement Learning and Model Predictive Control for Adaptive Cruise Control,” IEEE Transactions on Intelligent Vehicles , vol. 6, no. 2, pp. 221–231, Jun. 2021, conference Name: IEEE Transactions on Intelligent Vehicles

  196. [203]

    On the implementation of an interior- point filter line-search algorithm for large-scale nonlinear programming,

    A. W ¨achter and L. T. Biegler, “On the implementation of an interior- point filter line-search algorithm for large-scale nonlinear programming,” Mathematical Programming, vol. 106, no. 1, pp. 25–57, Mar. 2006

  197. [204]

    Lessons Learned from Data-Driven Building Control Experiments: Contrasting Gaussian Process-based MPC, Bilevel DeePC, and Deep Reinforcement Learning,

    L. Di Natale, Y . Lian, E. T. Maddalena, J. Shi, and C. N. Jones, “Lessons Learned from Data-Driven Building Control Experiments: Contrasting Gaussian Process-based MPC, Bilevel DeePC, and Deep Reinforcement Learning,” in IEEE 61st Conference on Decision and Control (CDC) , De...

  198. [205]

    Data-Enabled Predictive Control: In the Shallows of the DeePC,

    J. Coulson, J. Lygeros, and F. D ¨orfler, “Data-Enabled Predictive Control: In the Shallows of the DeePC,” in 18th European Control Conference (ECC), Jun. 2019, pp. 307–312

  199. [206]

    An experimental study of two predictive reinforcement learning methods and comparison with model-predictive control,

    D. Dobriborsci, P. Osinenko, and W. Aumer, “An experimental study of two predictive reinforcement learning methods and comparison with model-predictive control,” IFAC-PapersOnLine, vol. 55, no. 10, pp. 1545–1550, Jan. 2022

  200. [207]

    SciPy v1.15.0 Manual

    “SciPy v1.15.0 Manual.” [Online]. Available: https://docs.scipy.org/doc/ scipy/reference/optimize.minimize-slsqp.html

  201. [208]

    Evaluating Model-Based Planning and Planner Amortization for Continuous Control,

    A. Byravan, L. Hasenclever, P. Trochim, M. Mirza, A. D. Ialongo, J. T. Tassa, Y . andSpringenberg, A. Abdolmaleki, N. Heess, J. Merel, and M. Riedmiller, “Evaluating Model-Based Planning and Planner Amortization for Continuous Control,” in 10th International Conference on Lear...

  202. [209]

    Probabilistic Planning with Sequential Monte Carlo methods,

    A. Piche, V . Thomas, C. Ibrahim, Y . Bengio, and C. Pal, “Probabilistic Planning with Sequential Monte Carlo methods,” Sep. 2018

  203. [210]

    Maximum a Posteriori Policy Optimisation,

    A. Abdolmaleki, J. T. Springenberg, Y . Tassa, R. Munos, N. Heess, and M. Riedmiller, “Maximum a Posteriori Policy Optimisation,” in International Conference on Learning Representations , 2018

  204. [211]

    Reaching the limit in autonomous racing: Optimal control versus reinforcement learning,

    Y . Song, A. Romero, M. M ¨uller, V . Koltun, and D. Scaramuzza, “Reaching the limit in autonomous racing: Optimal control versus reinforcement learning,” Science Robotics, vol. 8, no. 82, p. eadg1462, Sep. 2023, publisher: American Association for the Advancement of Science

  205. [212]

    Model-Based Predictive Control and Reinforcement Learning for Planning Vehicle-Parking Trajectories for Vertical Parking Spaces,

    J. Shi, K. Li, C. Piao, J. Gao, and L. Chen, “Model-Based Predictive Control and Reinforcement Learning for Planning Vehicle-Parking Trajectories for Vertical Parking Spaces,” Sensors, vol. 23, no. 16, p. 7124, Jan. 2023, number: 16 Publisher: Multidisciplinary Digital Publish...

  206. [213]

    Comparison of Traffic Control with Model Predictive Control and Deep Reinforcement Learning,

    M. Imran, R. Izzo, A. Tortorelli, and F. Liberati, “Comparison of Traffic Control with Model Predictive Control and Deep Reinforcement Learning,” in 9th International Conference on Control, Decision and Information Technologies (CoDIT), Jul. 2023, pp. 989–994, iSSN: 2576- 3555

  207. [214]

    A Hierarchical Approach for Strategic Motion Planning in Autonomous Racing,

    R. Reiter, J. Hoffmann, J. Boedecker, and M. Diehl, “A Hierarchical Approach for Strategic Motion Planning in Autonomous Racing,” in European Control Conference (ECC) , Jun. 2023, pp. 1–8

  208. [215]

    Reinforcement Learning versus Model Predictive Control on greenhouse 33 climate control,

    B. Morcego, W. Yin, S. Boersma, E. van Henten, V . Puig, and C. Sun, “Reinforcement Learning versus Model Predictive Control on greenhouse 33 climate control,” Computers and Electronics in Agriculture , vol. 215, p. 108372, Dec. 2023

  209. [216]

    Comparison of Reinforcement Learning and Model Predictive Control for Automated Generation of Optimal Control for Dynamic Systems within a Design Space Explo- ration Framework,

    P. Hoffmann, K. Gorelik, and V . Ivanov, “Comparison of Reinforcement Learning and Model Predictive Control for Automated Generation of Optimal Control for Dynamic Systems within a Design Space Explo- ration Framework,” International Journal of Automotive Engineering , vol. 15...

  210. [217]

    Guided Policy Search,

    S. Levine and V . Koltun, “Guided Policy Search,” in Proceedings of the 30th International Conference on Machine Learning . PMLR, May 2013, pp. 1–9, iSSN: 1938-7228

  211. [218]

    DTC: Deep Tracking Control,

    F. Jenelten, J. He, F. Farshidian, and M. Hutter, “DTC: Deep Tracking Control,” Science Robotics , vol. 9, no. 86, p. eadh5401, Jan. 2024, publisher: American Association for the Advancement of Science

  212. [219]

    Imitation Learning from Nonlinear MPC via the Exact Q-Loss and its Gauss- Newton Approximation,

    A. Ghezzi, J. Hoffman, J. Frey, J. Boedecker, and M. Diehl, “Imitation Learning from Nonlinear MPC via the Exact Q-Loss and its Gauss- Newton Approximation,” in 62nd IEEE Conference on Decision and Control (CDC), Dec. 2023, pp. 4766–4771, iSSN: 2576-2370

  213. [220]

    Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control,

    K. Lowrey, A. Rajeswaran, S. Kakade, E. Todorov, and I. Mordatch, “Plan Online, Learn Offline: Efficient Learning and Exploration via Model-Based Control,” in 7th International Conference on Learning Representations, 2019

  214. [221]

    Where to go Next: Learning a Subgoal Recommendation Policy for Navigation in Dynamic Environments,

    B. Brito, M. Everett, J. P. How, and J. Alonso-Mora, “Where to go Next: Learning a Subgoal Recommendation Policy for Navigation in Dynamic Environments,” IEEE Robotics and Automation Letters , vol. 6, no. 3, pp. 4616–4623, Jul. 2021

  215. [222]

    N-MPC for Deep Neural Network- Based Collision Avoidance exploiting Depth Images,

    M. Jacquet and K. Alexis, “N-MPC for Deep Neural Network- Based Collision Avoidance exploiting Depth Images,” Feb. 2024, arXiv:2402.13038 [cs]

  216. [223]

    Blending MPC & Value Function Approximation for Efficient Reinforcement Learning,

    M. Bhardwaj, S. Choudhury, and B. Boots, “Blending MPC & Value Function Approximation for Efficient Reinforcement Learning,” in 9th International Conference on Learning Representations, ICLR , 2021

  217. [224]

    Data-Driven Economic NMPC Using Rein- forcement Learning,

    S. Gros and M. Zanon, “Data-Driven Economic NMPC Using Rein- forcement Learning,” IEEE Transactions on Automatic Control , vol. 65, no. 2, pp. 636–648, Feb. 2020

  218. [225]

    Model Predictive Control-Based Reinforcement Learning Using Expected Sarsa,

    H. Moradimaryamnegari, M. Frego, and A. Peer, “Model Predictive Control-Based Reinforcement Learning Using Expected Sarsa,” IEEE Access, vol. 10, pp. 81 177–81 191, 2022, conference Name: IEEE Access

  219. [226]

    Learning to Play Trajectory Games Against Opponents With Unknown Objectives,

    X. Liu, L. Peters, and J. Alonso-Mora, “Learning to Play Trajectory Games Against Opponents With Unknown Objectives,” IEEE Robotics and Automation Letters , vol. 8, no. 7, pp. 4139–4146, Jul. 2023, conference Name: IEEE Robotics and Automation Letters

  220. [227]

    RL + Model-Based Control: Using On-Demand Optimal Control to Learn Versatile Legged Locomotion,

    D. Kang, J. Cheng, M. Zamora, F. Zargarbashi, and S. Coros, “RL + Model-Based Control: Using On-Demand Optimal Control to Learn Versatile Legged Locomotion,” IEEE Robotics and Automation Letters , vol. 8, no. 10, pp. 6619–6626, Oct. 2023

  221. [228]

    Variational Policy Search via Trajectory Optimization,

    S. Levine and V . Koltun, “Variational Policy Search via Trajectory Optimization,” in Advances in Neural Information Processing Systems , vol. 26. Curran Associates, Inc., 2013

  222. [229]

    Variational inference for policy search in changing situations,

    G. Neumann, “Variational inference for policy search in changing situations,” in Proceedings of the 28th International Conference on International Conference on Machine Learning . Bellevue, Washington, USA: Omnipress, 2011, pp. 817–824

  223. [230]

    Combining the benefits of function approximation and trajectory optimization,

    I. Mordatch and E. Todorov, “Combining the benefits of function approximation and trajectory optimization,” vol. 10, Jul. 2014

  224. [231]

    End-to-end training of deep visuomotor policies,

    S. Levine, C. Finn, T. Darrell, and P. Abbeel, “End-to-end training of deep visuomotor policies,” Journal of Machine Learning Research , vol. 17, no. 39, pp. 1–40, 2016, iSBN: 1533-7928

  225. [232]

    Bregman Alternating Direction Method of Multipliers,

    H. Wang and A. Banerjee, “Bregman Alternating Direction Method of Multipliers,” in Advances in Neural Information Processing Systems , vol. 27. Curran Associates, Inc., 2014

  226. [233]

    A Fast Integrated Planning and Control Framework for Autonomous Driving via Imitation Learning

    L. Sun, C. Peng, W. Zhan, and M. Tomizuka, “A Fast Integrated Planning and Control Framework for Autonomous Driving via Imitation Learning.” American Society of Mechanical Engineers Digital Collection, Nov. 2018

  227. [234]

    Exploring Model-based Planning with Policy Networks,

    T. Wang and J. Ba, “Exploring Model-based Planning with Policy Networks,” Sep. 2019

  228. [235]

    Extracting Strong Policies for Robotics Tasks from Zero-Order Trajectory Optimizers,

    C. Pinneri, S. Sawant, S. Blaes, and G. Martius, “Extracting Strong Policies for Robotics Tasks from Zero-Order Trajectory Optimizers,” in International Conference on Learning Representations , 2021

  229. [236]

    Learning to Optimize in Model Predictive Control,

    J. Sacks and B. Boots, “Learning to Optimize in Model Predictive Control,” in International Conference on Robotics and Automation (ICRA), May 2022, pp. 10 549–10 556

  230. [237]

    Handling Sparse Rewards in Reinforcement Learning Using Model Predictive Control,

    M. Dawood, N. Dengler, J. de Heuvel, and M. Bennewitz, “Handling Sparse Rewards in Reinforcement Learning Using Model Predictive Control,” in IEEE International Conference on Robotics and Automation (ICRA), May 2023, pp. 879–885

  231. [238]

    Model Predictive Control via On-Policy Imitation Learning,

    K. Ahn, Z. Mhammedi, H. Mania, Z.-W. Hong, and A. Jadbabaie, “Model Predictive Control via On-Policy Imitation Learning,” in Proceedings of The 5th Annual Learning for Dynamics and Control Conference. PMLR, Jun. 2023, pp. 1493–1505, iSSN: 2640-3498

  232. [239]

    Enforcing the consensus between Trajectory Optimization and Policy Learning for precise robot control,

    Q. L. Lidec, W. Jallet, I. Laptev, C. Schmid, and J. Carpentier, “Enforcing the consensus between Trajectory Optimization and Policy Learning for precise robot control,” in IEEE International Conference on Robotics and Automation (ICRA) , May 2023, pp. 946–952

  233. [240]

    Crocoddyl: An Efficient and Versatile Framework for Multi-Contact Optimal Control,

    C. Mastalli, R. Budhiraja, W. Merkt, G. Saurel, B. Hammoud, M. Naveau, J. Carpentier, S. Vijayakumar, and N. Mansard, “Crocoddyl: An Efficient and Versatile Framework for Multi-Contact Optimal Control,” in ICRA 2020 IEEE International Conference on Robotics and Automation, Par...

  234. [241]

    Learning Robust Rewards with Adversarial Inverse Reinforcement Learning,

    J. Fu, K. Luo, and S. Levine, “Learning Robust Rewards with Adversarial Inverse Reinforcement Learning,” Aug. 2018

  235. [242]

    Generative Adversarial Imitation Learning,

    J. Ho and S. Ermon, “Generative Adversarial Imitation Learning,” in Advances in Neural Information Processing Systems , vol. 29. Curran Associates, Inc., 2016

  236. [243]

    PlanNetX: Learning an efficient neural network planner from MPC for longitudinal control,

    J. Hoffmann, D. F. Clausen, J. Brosseit, J. Bernhard, K. Esterle, M. Werling, M. Karg, and J. J. B ¨odecker, “PlanNetX: Learning an efficient neural network planner from MPC for longitudinal control,” in Proceedings of the 6th Annual Learning for Dynamics & Control Conference....

  237. [244]

    Learning When to Trust the Expert for Guided Exploration in RL,

    F. Schulz, J. Hoffmann, Y . Zhang, and J. Boedecker, “Learning When to Trust the Expert for Guided Exploration in RL,” in ICML 2024 Workshop: Foundations of Reinforcement Learning and Control – Connections and Perspectives, 2024

  238. [245]

    The explicit solution of constrained LP-Based receding horizon control,

    A. Bemporad, F. Borrelli, and M. Morari, “The explicit solution of constrained LP-Based receding horizon control,” in Proceedings of the IEEE conference on decision and control (CDC) , Sydney, Australia, 1999

  239. [246]

    Model predictive control based on linear programming - the explicit solution,

    ——, “Model predictive control based on linear programming - the explicit solution,” IEEE Transactions on Automatic Control , vol. 47, no. 12, pp. 1974–1985, Dec. 2002, conference Name: IEEE Transactions on Automatic Control

  240. [247]

    Approximate explicit constrained linear model predictive control via orthogonal search tree,

    T. Johansen and A. Grancharova, “Approximate explicit constrained linear model predictive control via orthogonal search tree,” IEEE Trans. Automatic Control, vol. 48, pp. 810–815, 2003

  241. [248]

    Accelerating nonlinear model predictive control through machine learning,

    Y . Vaupel, N. C. Hamacher, A. Caspari, A. Mhamdi, I. G. Kevrekidis, and A. Mitsos, “Accelerating nonlinear model predictive control through machine learning,” Journal of Process Control , vol. 92, pp. 261–270, Aug. 2020

  242. [249]

    Learning an Approximate Model Predictive Controller With Guarantees,

    M. Hertneck, J. K ¨ohler, S. Trimpe, and F. Allg ¨ower, “Learning an Approximate Model Predictive Controller With Guarantees,” IEEE Control Systems Letters, vol. 2, no. 3, pp. 543–548, Jul. 2018, conference Name: IEEE Control Systems Letters

  243. [250]

    A neural network model predictive controller,

    B. M. ˚Akesson and H. T. Toivonen, “A neural network model predictive controller,” Journal of Process Control , vol. 16, no. 9, pp. 937–946, Oct. 2006

  244. [251]

    Deep Neural Network Approximation of Nonlinear Model Predictive Control,

    Y . Cao and R. B. Gopaluni, “Deep Neural Network Approximation of Nonlinear Model Predictive Control,” IFAC-PapersOnLine, vol. 53, no. 2, pp. 11 319–11 324, Jan. 2020

  245. [252]

    A Neural Network Architecture to Learn Explicit MPC Controllers from Data,

    “A Neural Network Architecture to Learn Explicit MPC Controllers from Data,” Ifac Papersonline, 2020

  246. [253]

    Approximating Explicit Model Predictive Control Using Constrained Neural Networks,

    S. Chen, K. Saulnier, N. Atanasov, D. D. Lee, V . Kumar, G. J. Pappas, and M. Morari, “Approximating Explicit Model Predictive Control Using Constrained Neural Networks,” in 2018 Annual American Control Conference (ACC), Jun. 2018, pp. 1520–1527

  247. [254]

    Efficient Representation and Approximation of Model Predictive Control Laws via Deep Learning,

    B. Karg and S. Lucia, “Efficient Representation and Approximation of Model Predictive Control Laws via Deep Learning,” IEEE Transactions on Cybernetics, vol. 50, no. 9, pp. 3866–3878, Sep. 2020

  248. [255]

    End-to-End Imitation Learning with Safety Guarantees using Control Barrier Functions,

    R. K. Cosner, Y . Yue, and A. D. Ames, “End-to-End Imitation Learning with Safety Guarantees using Control Barrier Functions,” in 2022 IEEE 61st Conference on Decision and Control (CDC) . Cancun, Mexico: IEEE, Dec. 2022, pp. 5316–5322

  249. [256]

    Stability Verification of Neural Network Controllers Using Mixed-Integer Programming,

    R. Schwan, C. N. Jones, and D. Kuhn, “Stability Verification of Neural Network Controllers Using Mixed-Integer Programming,” IEEE Transactions on Automatic Control , pp. 1–16, 2023, conference Name: IEEE Transactions on Automatic Control

  250. [257]

    Learning Constrained Parametric Differentiable Predictive Control Policies With Guarantees,

    J. Drgo ˇna, A. Tuor, and D. Vrabie, “Learning Constrained Parametric Differentiable Predictive Control Policies With Guarantees,” IEEE Transactions on Systems, Man, and Cybernetics: Systems , pp. 1–12, 2024

  251. [258]

    Neural Network Verification in Control,

    M. Everett, “Neural Network Verification in Control,” in 2021 60th IEEE Conference on Decision and Control (CDC) . Austin, TX, USA: IEEE, Dec. 2021, pp. 6326–6340. 34

  252. [259]

    MPC-Net: A First Principles Guided Policy Search,

    J. Carius, F. Farshidian, and M. Hutter, “MPC-Net: A First Principles Guided Policy Search,” IEEE Robotics and Automation Letters , vol. 5, no. 2, pp. 2897–2904, Apr. 2020

  253. [260]

    Imitation Learning from MPC for Quadrupedal Multi-Gait Control,

    A. Reske, J. Carius, Y . Ma, F. Farshidian, and M. Hutter, “Imitation Learning from MPC for Quadrupedal Multi-Gait Control,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . Xi’an, China: IEEE, May 2021, pp. 5014–5020

  254. [261]

    Between MDPs and semi- MDPs: A framework for temporal abstraction in reinforcement learning,

    R. S. Sutton, D. Precup, and S. Singh, “Between MDPs and semi- MDPs: A framework for temporal abstraction in reinforcement learning,” Artificial intelligence, vol. 112, no. 1-2, pp. 181–211, 1999, publisher: Elsevier

  255. [262]

    Equivalence of Optimality Criteria for Markov Decision Process and Model Predictive Control,

    A. B. Kordabad, M. Zanon, and S. Gros, “Equivalence of Optimality Criteria for Markov Decision Process and Model Predictive Control,” IEEE Transactions on Automatic Control, vol. 69, no. 2, pp. 1149–1156, Feb. 2024, conference Name: IEEE Transactions on Automatic Control

  256. [264]

    Value function approximation and model predictive control,

    M. Zhong, M. Johnson, Y . Tassa, T. Erez, and E. Todorov, “Value function approximation and model predictive control,” in IEEE Sympo- sium on Adaptive Dynamic Programming and Reinforcement Learning (ADPRL), Apr. 2013, pp. 100–107, iSSN: 2325-1867

  257. [265]

    Provably safe and robust learning-based model predictive control,

    A. Aswani, H. Gonzalez, S. S. Sastry, and C. Tomlin, “Provably safe and robust learning-based model predictive control,” Automatica, vol. 49, no. 5, pp. 1216–1226, 2013

  258. [266]

    SNOPT: An SQP Algorithm for Large-Scale Constrained Optimization,

    P. E. Gill, W. Murray, and M. A. Saunders, “SNOPT: An SQP Algorithm for Large-Scale Constrained Optimization,” SIAM Journal on Optimization, vol. 12, no. 4, pp. 979–1006, 2002

  259. [267]

    R. Y . Rubinstein and D. P. Kroese, The Cross-Entropy Method, ser. In- formation Science and Statistics, M. Jordan, J. Kleinberg, B. Sch ¨olkopf, F. P. Kelly, and I. Witten, Eds. New York, NY: Springer, 2004

  260. [268]

    LVIS: Learning from Value Function Intervals for Contact-Aware Robot Controllers,

    R. Deits, T. Koolen, and R. Tedrake, “LVIS: Learning from Value Function Intervals for Contact-Aware Robot Controllers,” in Interna- tional Conference on Robotics and Automation (ICRA) . Montreal, QC, Canada: IEEE Press, 2019, pp. 7762–7768

  261. [269]

    Data Efficient Reinforcement Learning for Legged Robots,

    Y . Yang, K. Caluwaerts, A. Iscen, T. Zhang, J. Tan, and V . Sindhwani, “Data Efficient Reinforcement Learning for Legged Robots,” in Pro- ceedings of the Conference on Robot Learning . PMLR, May 2020, pp. 1–10, iSSN: 2640-3498

  262. [270]

    Design Principles for a Family of Direct-Drive Legged Robots,

    G. Kenneally, A. De, and D. E. Koditschek, “Design Principles for a Family of Direct-Drive Legged Robots,” IEEE Robotics and Automation Letters, vol. 1, no. 2, pp. 900–907, Jul. 2016

  263. [271]

    Low-Level Control of a Quadrotor With Deep Model- Based Reinforcement Learning,

    N. Lambert, D. S. Drew, J. Yaconelli, S. Levine, R. Calandra, and K. S. J. Pister, “Low-Level Control of a Quadrotor With Deep Model- Based Reinforcement Learning,” IEEE Robotics and Automation Letters, vol. 4, pp. 4224–4230, 2019

  264. [272]

    Practical Reinforcement Learning For MPC: Learning from sparse objectives in under an hour on a real robot,

    N. Karnchanachari, M. I. Valls, D. Hoeller, and M. Hutter, “Practical Reinforcement Learning For MPC: Learning from sparse objectives in under an hour on a real robot,” in Proceedings of the 2nd Conference on Learning for Dynamics and Control . PMLR, Jul. 2020, pp. 211–224, iS...

  265. [273]

    A Q-learning predictive control scheme with guaranteed stability,

    L. Beckenbach, P. Osinenko, and S. Streif, “A Q-learning predictive control scheme with guaranteed stability,” European Journal of Control, vol. 56, pp. 167–178, Nov. 2020

  266. [274]

    Deep Value Model Predictive Control,

    D. Hoeller, F. Farshidian, and M. Hutter, “Deep Value Model Predictive Control,” in Proceedings of the Conference on Robot Learning . PMLR, May 2020, pp. 990–1004, iSSN: 2640-3498

  267. [275]

    The Value of Planning for Infinite-Horizon Model Predictive Control,

    N. Hatch and B. Boots, “The Value of Planning for Infinite-Horizon Model Predictive Control,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) , May 2021, pp. 7372–7378, iSSN: 2577-087X

  268. [276]

    Model Predictive Actor-Critic: Accelerating Robot Skill Acquisition with Deep Reinforcement Learning,

    A. S. Morgan, D. Nandha, G. Chalvatzaki, C. D’Eramo, A. M. Dollar, and J. Peters, “Model Predictive Actor-Critic: Accelerating Robot Skill Acquisition with Deep Reinforcement Learning,” in 2021 IEEE International Conference on Robotics and Automation (ICRA) . IEEE Press, 2021,...

  269. [277]

    Approximate infinite-horizon predictive control,

    L. Beckenbach and S. Streif, “Approximate infinite-horizon predictive control,” in IEEE 61st Conference on Decision and Control (CDC) , Dec. 2022, pp. 3711–3717, iSSN: 2576-2370

  270. [278]

    Predictive Control with Learning-Based Terminal Costs Using Approximate Value Iteration,

    F. Moreno-Mora, L. Beckenbach, and S. Streif, “Predictive Control with Learning-Based Terminal Costs Using Approximate Value Iteration,” IFAC-PapersOnLine, vol. 56, no. 2, pp. 3874–3879, Jan. 2023

  271. [279]

    MATLAB version: 9.13.0 (R2022b),

    T. M. Inc, “MATLAB version: 9.13.0 (R2022b),” Natick, Massachusetts, United States, 2022. [Online]. Available: https://www.mathworks.com

  272. [280]

    Reinforcement Learning- Based Model Predictive Control for Discrete-Time Systems,

    M. Lin, Z. Sun, Y . Xia, and J. Zhang, “Reinforcement Learning- Based Model Predictive Control for Discrete-Time Systems,” IEEE Transactions on Neural Networks and Learning Systems , vol. 35, no. 3, pp. 3312–3324, Mar. 2024, conference Name: IEEE Transactions on Neural Network...

  273. [281]

    RL-Driven MPPI: Accelerating Online Control Laws Calculation With Offline Policy,

    Y . Qu, H. Chu, S. Gao, J. Guan, H. Yan, L. Xiao, S. E. Li, and J. Duan, “RL-Driven MPPI: Accelerating Online Control Laws Calculation With Offline Policy,” IEEE Transactions on Intelligent Vehicles, vol. 9, no. 2, pp. 3605–3616, Feb. 2024

  274. [282]

    A Learning-Based Model Predictive Control Strategy for Home Energy Management Systems,

    W. Cai, S. Sawant, D. Reinhardt, S. Rastegarpour, and S. Gros, “A Learning-Based Model Predictive Control Strategy for Home Energy Management Systems,” IEEE Access , vol. 11, pp. 145 264–145 280, 2023

  275. [283]

    Differentiable MPC for End-to-end Planning and Control,

    B. Amos, I. Jimenez, J. Sacks, B. Boots, and J. Z. Kolter, “Differentiable MPC for End-to-end Planning and Control,” in Advances in Neural Information Processing Systems , vol. 31. Curran Associates, Inc., 2018

  276. [284]

    Safe Reinforcement Learning Using Robust MPC,

    M. Zanon and S. Gros, “Safe Reinforcement Learning Using Robust MPC,” IEEE Transactions on Automatic Control, vol. 66, no. 8, pp. 3638– 3652, Aug. 2021, conference Name: IEEE Transactions on Automatic Control

  277. [285]

    Actor-Critic Model Predictive Control,

    A. Romero, Y . Song, and D. Scaramuzza, “Actor-Critic Model Predictive Control,” in 2024 IEEE International Conference on Robotics and Automation (ICRA). Yokohama, Japan: IEEE, May 2024, pp. 14 777– 14 784

  278. [286]

    Reinforcement Learning Based on Real-Time Iteration NMPC,

    M. Zanon, V . Kungurtsev, and S. Gros, “Reinforcement Learning Based on Real-Time Iteration NMPC,” IFAC-PapersOnLine, vol. 53, no. 2, pp. 5213–5218, Jan. 2020

  279. [287]

    Adaptive Stochastic Nonlinear Model Predictive Control with Look-ahead Deep Reinforcement Learn- ing for Autonomous Vehicle Motion Control,

    B. Zarrouki, C. Wang, and J. Betz, “Adaptive Stochastic Nonlinear Model Predictive Control with Look-ahead Deep Reinforcement Learn- ing for Autonomous Vehicle Motion Control,” in 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , Oct. 2024, pp. ...

  280. [288]

    Bias Correction of Discounted Optimal Steady-State using Cost Modification,

    A. B. Kordabad and S. Gros, “Bias Correction of Discounted Optimal Steady-State using Cost Modification,” in 2023 European Control Conference (ECC), Jun. 2023, pp. 1–6

  281. [289]

    Towards Safe Reinforcement Learning Using NMPC and Policy Gradients: Part I - Stochastic case,

    S. Gros and M. Zanon, “Towards Safe Reinforcement Learning Using NMPC and Policy Gradients: Part I - Stochastic case,” Jun. 2019, arXiv:1906.04057 [cs]

  282. [290]

    Optimistic Exploration even with a Pessimistic Initialisation,

    T. Rashid, B. Peng, W. B ¨ohmer, and S. Whiteson, “Optimistic Exploration even with a Pessimistic Initialisation,” in 8th International Conference on Learning Representations, ICLR , 2020

  283. [291]

    Reinforcement learning and model predictive control for robust embedded quadrotor guidance and control,

    C. Greatwood and A. G. Richards, “Reinforcement learning and model predictive control for robust embedded quadrotor guidance and control,” Autonomous Robots, vol. 43, no. 7, pp. 1681–1693, Oct. 2019

  284. [292]

    Learning When to Drive in Intersections by Combining Reinforcement Learning and Model Predictive Control,

    T. Tram, I. Batkovic, M. Ali, and J. Sj ¨oberg, “Learning When to Drive in Intersections by Combining Reinforcement Learning and Model Predictive Control,” in 2019 IEEE Intelligent Transportation Systems Conference (ITSC), Oct. 2019, pp. 3263–3268

  285. [293]

    FORCES Professional,

    A. Domahidi and J. Jerez, “FORCES Professional,” 2014, published: Embotech AG, url=https://embotech.com/FORCES-Pro

  286. [294]

    Learning Interaction-Aware Guidance for Trajectory Optimization in Dense Traffic Scenarios,

    B. Brito, A. Agarwal, and J. Alonso-Mora, “Learning Interaction-Aware Guidance for Trajectory Optimization in Dense Traffic Scenarios,” IEEE Transactions on Intelligent Transportation Systems , vol. 23, no. 10, pp. 18 808–18 821, Oct. 2022, conference Name: IEEE Transactions o...

  287. [295]

    Learning-Based Model Predictive Control for Quadruped Locomotion on Slippery Ground,

    Z. Zhang, H. An, Q. Wei, and H. Ma, “Learning-Based Model Predictive Control for Quadruped Locomotion on Slippery Ground,” in 4th International Conference on Control and Robotics (ICCR) , Dec. 2022, pp. 47–52

  288. [296]

    Safe Reinforcement Learning with Chance-constrained Model Predictive Control,

    S. Pfrommer, T. Gautam, A. Zhou, and S. Sojoudi, “Safe Reinforcement Learning with Chance-constrained Model Predictive Control,” in Proceedings of The 4th Annual Learning for Dynamics and Control Conference. PMLR, May 2022, pp. 291–303, iSSN: 2640-3498

  289. [297]

    ApS, The MOSEK optimization toolbox for MATLAB manual

    M. ApS, The MOSEK optimization toolbox for MATLAB manual. Version 10.1., 2024. [Online]. Available: http://docs.mosek.com/latest/ toolbox/index.html

  290. [298]

    DiffTune- MPC: Closed-Loop Learning for Model Predictive Control,

    R. Tao, S. Cheng, X. Wang, S. Wang, and N. Hovakimyan, “DiffTune- MPC: Closed-Loop Learning for Model Predictive Control,” IEEE Robotics and Automation Letters , 2024, publisher: IEEE

  291. [299]

    A Safe Reinforcement Learn- ing driven Weights-varying Model Predictive Control for Autonomous Vehicle Motion Control: 35th IEEE Intelligent Vehicles Symposium, IV 2024,

    B. Zarrouki, M. Spanakakis, and J. Betz, “A Safe Reinforcement Learn- ing driven Weights-varying Model Predictive Control for Autonomous Vehicle Motion Control: 35th IEEE Intelligent Vehicles Symposium, IV 2024,” 35th IEEE Intelligent Vehicles Symposium, IV 2024 , pp. 1401–1408, 2024

  292. [2023]

    Available: https://www.gurobi.com

    [Online]. Available: https://www.gurobi.com

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.