Pith. sign in

REVIEW 3 major objections 5 minor 300 references

Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review

T0 review · 3 major / 5 minor · reviewed 2026-08-15 · deepseek-v4-flash

Pith's one-line read This review argues that combining Bayesian inference with reinforcement learning yields agents that are more data-efficient, generalizable, interpretable, and safe, and it offers the first systematic meta-perspective on those combinations.

desk verdict A useful survey of Bayesian+RL combinations, but its 'first meta-perspective' claim and Table I ratings rest on an undocumented selection process. read the letter →

arxiv 2505.07911 v1 pith:DI4CJWDG submitted 2025-05-12 cs.LG cs.AI

classification cs.LGcs.AI
keywords Bayesianinferencereinforcementlearningagentdecisionmakinguncertaintyquantificationmeta-learninglifelongvariationaloptimization
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Bayesian inference brings uncertainty quantification, while reinforcement learning brings a decision-making loop; the paper argues that combining them yields agents that need less data, generalize more easily, behave more interpretably, and act more safely. To organize that claim, the review groups Bayesian methods into seven potential families and then traces how each family has been, or could be, combined with model-based RL, model-free RL, and inverse RL. It compares these combinations on four indicators — data efficiency, generalization, interpretability, and safety — and analyzes how Bayesian methods intervene at the data collection, data processing, and policy learning stages of RL. The paper asserts this is the first systematic meta-perspective on these combinations, and it concludes that the shared obstacle is multi-layer or multi-dimensional optimization, which hierarchical Bayesian structures are positioned to address.

What carries the argument

The load-bearing machinery is a classification scheme rather than a theorem. It is a three-part grid: seven potential Bayesian method families as rows, the RL pipeline (data collection, data processing, policy learning) as one axis, and four evaluation indicators (data efficiency, generalization, interpretability, safety) as the comparison axis. The paper uses this grid to assign each combination a qualitative rating, from poor to excellent, and to express the field's shared obstacle as a two-layer optimization problem in which Bayesian methods parameterize unknown components as local models inside a global RL objective.

What would settle it

A systematic re-derivation of Table I from the cited papers, with explicit inclusion criteria and independent raters, would settle whether the map is complete: disagreement about which family fits a given method, or low inter-rater agreement on the indicator levels, would falsify the claim of a systematic meta-perspective.

Watch

Extended reading notes

Core claim

The central claim is that the design space of combining Bayesian inference with RL is orderly and analysable, not a scattered collection of tricks. The paper's contribution is a map: seven potential Bayesian methods — variational inference, Bayesian optimization, Bayesian neural networks, Bayesian active learning, Bayesian generative models, Bayesian meta-learning, and lifelong Bayesian learning — each paired with RL in classical and recent forms, then evaluated by four indicators. Within that map, model-based Bayesian RL, model-free Bayesian RL, and Bayesian inverse RL are treated as classical combinations, while the seven families are treated as current frontiers. The paper also identifies a common diagnosis across six hard RL variants (unknown reward, partial observability, multi-agent, multi-task, nonlinear non-Gaussian, and hierarchical RL): the hard part is two-layer or multi-dimensional optimization, and Bayesian methods help by parameterizing unknown components as local models within a global RL objective.

Load-bearing premise

The load-bearing premise is that the review's selection of seven Bayesian method families and four evaluation indicators, together with the qualitative ratings assigned in Table I, is a fair and complete representation of the field; if the selection is unrepresentative or the ratings are not reproducible, the map loses its authority.

Editorial extensions

If this is right

  • A researcher choosing an approach can read Table I as a placement map: each Bayesian-plus-RL combination comes with a stated task scope and a level for data efficiency, generalization, interpretability, and safety.
  • Bayesian meta-learning and lifelong Bayesian learning are predicted to give the largest data-efficiency and generalization gains because they reuse policies across tasks.
  • Bayesian optimization and variational inference are stage-agnostic tools that can appear anywhere in the RL pipeline, serving to find informative samples, reduce dimensions, and approximate intractable posteriors.
  • The paper's ten open questions imply that future progress depends less on bigger networks and datasets and more on designing hierarchical solutions that split each problem into local and global optimization levels.
  • Diffusion models are singled out as a promising vehicle for safe RL policies, because safety constraints can enter through the reward model that conditions the denoising process.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • If the qualitative ratings in Table I were replaced by quantitative benchmarks across the same four indicators, the table could be turned into a reusable selection guide for Bayesian-plus-RL method choice; the paper does not do that measurement.
  • The two-layer optimization diagnosis suggests a testable design rule: algorithms that explicitly separate inner and outer loops, as Bayesian meta-learning and lifelong Bayesian nonparametric models do, should outperform flat methods on multi-task adaptation benchmarks even when network capacity is held fixed.
  • The review's own account indicates a gap: Bayesian active learning mostly improves data quality rather than RL convergence directly, so a natural experiment is to treat episode selection for training as a bandit or RL problem and measure whether the resulting query policy improves sample efficiency.
  • The paper claims but does not independently verify the completeness of its map; a community-curated taxonomy test could check whether newly published Bayesian-plus-RL methods fit into the seven families or require an eighth family.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 5 minor

Summary. This paper is a survey of methods that combine Bayesian inference with reinforcement learning (RL) for agent decision making. It introduces seven 'potential Bayesian methods' (variational inference, Bayesian optimization, Bayesian neural networks, Bayesian active learning, Bayesian generative models, Bayesian meta-learning, and lifelong Bayesian learning), reviews their classical and recent combinations with model-based RL, model-free RL, and inverse RL, and then analytically compares these combinations on four indicators: data efficiency, generalization, interpretability, and safety. The paper also discusses Bayesian approaches to safe decision making and analyzes six complex RL variants (unknown reward, partial observability, multi-agent, multi-task, nonlinear non-Gaussian, and hierarchical RL), ending with ten open questions. The central claim is that combining Bayesian inference with RL yields important advantages in agent decision making and that the paper provides the first systematic meta-perspective on these combinations.

Significance. If the taxonomy and comparative conclusions are accepted, the survey provides a useful structured entry point to a fragmented literature, connecting classical topics (BAMDP, GP-based Bayesian RL, Bayesian IRL) with recent developments (diffusion planners, Bayesian meta-RL, lifelong Bayesian learning). The mathematical descriptions are mostly standard and the breadth of coverage is substantial; the ten open questions in Section VII are concrete and could guide future research. The paper also explicitly discusses safety, which is often underserved in surveys. However, the significance is moderated by the fact that the claimed novelty—the meta-perspective and the four-indicator comparison—rests on a selection of methods and on Table I ratings that are not derived from a documented, reproducible methodology.

major comments (3)
  1. [Section I] The statement 'RL is a subset of Bayesian inference because not all Bayesian inference problems can be formulated as RL within the MDP framework' is logically imprecise and not established. The 'RL as inference' literature (e.g., Levine, 2018) shows an equivalence between certain RL objectives and probabilistic inference under specific model assumptions, but this does not amount to set-theoretic containment of RL within Bayesian inference; both frameworks are general modeling paradigms and the direction of inclusion depends on the formalization. This claim is presented as a premise for the review's framing and should be either corrected to a precise statement (e.g., 'some RL problems can be formulated as Bayesian inference problems') or supported with a formal argument.
  2. [Section VI, Table I] The central comparative contribution is Table I, which rates algorithms as 'Poor,' 'Acc.' (acceptable), 'Good,' or 'Exc.' (excellent) on data efficiency, interpretability, generalization, and safety. Section VI states only that the authors 'make general analysis and comparisons given the utility/strength of Bayesian methods'; no systematic evaluation protocol, inclusion/exclusion criteria, evidence anchors, or inter-rater criteria are provided. Because the paper's claimed novelty is the meta-perspective itself, these ratings are load-bearing; as they stand, they are non-reproducible qualitative judgments. The authors should either (a) document a systematic search and screening protocol for the methods and papers underlying each row, and define what evidence would justify a rating, or (b) explicitly present Table I as an opinionated synthesis rather than a systematic comparison, and remove or qualify the 'first systematic' claim accordingly.
  3. [Section I] The claim 'Given our knowledge, this is the first paper to systematically investigate the combinations of Bayesian inference and RL for agent decision making from a meta perspective' is an unsupported novelty assertion. The authors do not report a literature search protocol, databases consulted, or comparison against other surveys beyond a brief list of prior reviews in Section I. While 'given our knowledge' is a hedge, the phrase 'systematically investigate' implies a methodology that is not described. The paper should describe its selection process for the seven Bayesian methods and for the papers cited in Sections IV and V, or soften the claim to match the actual narrative scope. This is important because the novelty claim is part of the central contribution.
minor comments (5)
  1. [Section IV.E (heading)] The subsection heading 'Combing Bayesian generative models with RL' contains a typo; it should read 'Combining Bayesian generative models with RL'.
  2. [Sections II.C and IV.G] The acronym 'DMPP' appears several times (e.g., Section II.C, Section IV.G) where the intended term is 'DPMM' (Dirichlet process mixture model), as used elsewhere in the paper. Please make the usage consistent.
  3. [Section II.C, Diffusion models paragraph] The sentence 'A suitable noise schedule results in balanced exploration and exploration' should read 'balanced exploration and exploitation'.
  4. [Table I] The table uses '---' in several cells (e.g., Generalization column for GPR-based Bayesian learning and model-free Bayesian RL) without explaining its meaning. The caption should define '---' as 'not assessed' or 'insufficient evidence'.
  5. [References] Several references are incomplete or inconsistently formatted, e.g., reference [16] is given as 'R. Learning' with the title 'Model-based and Model-free RL for Robot Control' and lacks venue details, and reference [41] cites 'Artificial Intelligence: A Modern Approach' without full book information. A thorough copyedit of the reference list is needed.

Circularity Check

0 steps flagged · score 2.0 of 10

No load-bearing circularity; one minor self-citation ([90]) is not used to derive any central claim, and the survey's qualitative comparisons do not reduce to their inputs.

full rationale

This is a survey rather than a derivation paper, so most circularity patterns do not apply. The central claims are (i) that Bayesian inference confers data-efficiency, generalization, interpretability and safety advantages when combined with RL, and (ii) that this is the first systematic meta-perspective on such combinations. Neither claim is obtained by fitting, renaming, or definitional construction: the four advantages are supported by cited external work, and the novelty claim is an assertion about the literature ('Given our knowledge, this is the first paper to systematically investigate...'), not a result derived from the paper's own definitions. The comparisons in Section VI are explicitly qualitative ('We here make general analysis and comparisons given the utility/strength of Bayesian methods'), and Table I entries such as 'Exc.'/'Acc.'/'Good' are categorical judgments; this is a reproducibility limitation, not circularity. The only self-citation found is [90] (Verdoja and Kyrki, ICML 2021 Workshop), used to note a 'latest flaw in providing consistent uncertainty estimations' of MC dropout in Section II-C. That claim is not load-bearing: it does not support the paper's central conclusions, and removing it would not change any rating or recommendation. No equation is shown to be identical to an input by construction, no fitted parameter is renamed as a prediction, and no uniqueness theorem from the authors' prior work is invoked. Accordingly, the paper is self-contained as a review and receives a low score reflecting only the presence of the minor, non-load-bearing self-citation.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The review introduces no fitted numbers or new entities. It relies on standard Bayesian mathematics and on domain assumptions about the adequacy of its chosen methods and indicators.

assumptions (4)
  • standard math Bayes' theorem and Gaussian process regression provide a valid basis for uncertainty quantification in RL.
    Used throughout Sections II and III; these are established mathematical tools, not original.
  • domain assumption The four indicators (data efficiency, generalization, interpretability, safety) are sufficient and meaningful for comparing Bayesian-RL combinations.
    Section VI; the review rates methods in Table I on these axes without defining metrics or empirical validation.
  • domain assumption The seven selected 'potential Bayesian methods' are representative of Bayesian methods relevant to RL.
    Sections I and IV; inclusion and exclusion criteria are not given.
  • ad hoc to paper RL can be viewed as a subset of Bayesian inference under the MDP framework.
    Section I; this is a contested framing used to justify the review's scope, not a proven theorem.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review." pith.science (2026). https://pith.science/paper/DI4CJWDG

@misc{pith2026250507911,
  author       = {Pith},
  title        = {Pith review of: Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DI4CJWDG}},
  note         = {Machine review of arXiv:2505.07911}
}
read the original abstract

Bayesian inference has many advantages in decision making of agents (e.g. robotics/simulative agent) over a regular data-driven black-box neural network: Data-efficiency, generalization, interpretability, and safety where these advantages benefit directly/indirectly from the uncertainty quantification of Bayesian inference. However, there are few comprehensive reviews to summarize the progress of Bayesian inference on reinforcement learning (RL) for decision making to give researchers a systematic understanding. This paper focuses on combining Bayesian inference with RL that nowadays is an important approach in agent decision making. To be exact, this paper discusses the following five topics: 1) Bayesian methods that have potential for agent decision making. First basic Bayesian methods and models (Bayesian rule, Bayesian learning, and Bayesian conjugate models) are discussed followed by variational inference, Bayesian optimization, Bayesian deep learning, Bayesian active learning, Bayesian generative models, Bayesian meta-learning, and lifelong Bayesian learning. 2) Classical combinations of Bayesian methods with model-based RL (with approximation methods), model-free RL, and inverse RL. 3) Latest combinations of potential Bayesian methods with RL. 4) Analytical comparisons of methods that combine Bayesian methods with RL with respect to data-efficiency, generalization, interpretability, and safety. 5) In-depth discussions in six complex problem variants of RL, including unknown reward, partial-observability, multi-agent, multi-task, non-linear non-Gaussian, and hierarchical RL problems and the summary of how Bayesian methods work in the data collection, data processing and policy learning stages of RL to pave the way for better agent decision-making strategies.

Figures

Figures reproduced from arXiv: 2505.07911 by the authors.

Figure 1
Figure 1. An overview of the paper structure. To achieve the above goal, this review ( [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 5
Figure 5. Bayesian methods on different processes of RL [PITH_FULL_IMAGE:figures/full_fig_p014_5.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

300 extracted references · 37 canonical work pages

  1. [1]

    Attention is all you need,

    A. Vaswani et al. , “Attention is all you need,” Adv. Neural Inf. Process. Syst., pp. 5999–6009, 2017

  2. [2]

    Attention in Psychology , Neuroscience , and Machine Learning,

    G. W. Lindsay, “Attention in Psychology , Neuroscience , and Machine Learning,” vol. 14, no. 29, pp. 1–21, 2020

  3. [3]

    DeepSeek -V3 Technical Report,

    DeepSeek-AI et al. , “DeepSeek -V3 Technical Report,” 18 Zhou, C. et al. Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review arXiv:2412.19437v1, vol. 2024, pp. 1–53, 2024

  4. [4]

    Mathematical Capabilities of ChatGPT,

    S. Frieder et al. , “Mathematical Capabilities of ChatGPT,” Thirty- seventh Annu. Conf. Neural Inf. Process. Syst. (NeurIPS 2023), pp. 1–46, 2023

  5. [5]

    Modulationof neuronal activity by target uncertainty,

    M. A. Basso and R. H. Wurtz, “Modulationof neuronal activity by target uncertainty,” Nature, vol. 389, pp. 66–69, 1997

  6. [6]

    Neural implementations of Bayesian inference,

    H. Sohn and D. Narain, “Neural implementations of Bayesian inference,” Curr. Opin. Neurobiol., vol. 70, pp. 121–129, 2021

  7. [7]

    Neural substrate of dynamic Bayesian inference in the cerebral cortex,

    A. Funamizu, B. Kuhn, and K. Doya, “Neural substrate of dynamic Bayesian inference in the cerebral cortex,” Nat. Neurosci., vol. 19, pp. 1682–1689, 2016

  8. [8]

    Neural implementation of Bayesian inference in a sensorimotor behavior,

    T. R. Darlington, J. M. Beck, and S. G. Lisberger, “Neural implementation of Bayesian inference in a sensorimotor behavior,” Nat. Neurosci., vol. 21, pp. 1442–1451, 2018

Show all 300 references
  1. [9]

    Bayesian Computation through Cortical Latent Dynamics,

    H. Sohn and D. Narain, “Bayesian Computation through Cortical Latent Dynamics,” Neuron, vol. 103, no. 5, pp. 934-947.e5, 2019

  2. [10]

    Neural Correlates of Optimal Multisensory Decision Making under Time -Varying Reliabilities with an Invariant Linear Probabilistic Population Code,

    H. Hou et al., “Neural Correlates of Optimal Multisensory Decision Making under Time -Varying Reliabilities with an Invariant Linear Probabilistic Population Code,” Neuron, vol. 104, pp. 1010-1021.e10, 2019

  3. [11]

    A neural basis of probabilistic computation in visual cortex,

    E. Y. Walker, R. J. Cotton, W. J. Ma, and A. S. Tolias, “A neural basis of probabilistic computation in visual cortex,” Nat. Neurosci. , vol. 23, pp. 122–129, 2020

  4. [12]

    Reinforcement learning and its connections with neuroscience and psychology,

    A. Subramanian, S. Chitlangia, and V. Baths, “Reinforcement learning and its connections with neuroscience and psychology,” Neural Networks, vol. 145, pp. 271–287, 2022

  5. [13]

    An Investigation of Model -Free Planning,

    A. Guez et al. , “An Investigation of Model -Free Planning,” arXiv:1901.03559, pp. 1–21, 2019

  6. [14]

    Meta - learning, social cognition and consciousness in brains and machines,

    A. Langdon, M. Botvinick, H. Nakahara, and K. Tanaka, “Meta - learning, social cognition and consciousness in brains and machines,” Neural Networks, vol. 145, pp. 80–89, 2022

  7. [15]

    Deep Reinforcement Learning and its Neuroscientific Implications,

    M. Botvinick, J. X. Wang, W. Dabney, K. J. Miller, and Z. Kurth - nelson, “Deep Reinforcement Learning and its Neuroscientific Implications,” arXiv:2007.03750v1, pp. 1–22, 2020

  8. [16]

    Model -based and Model -free RL for Robot Control,

    R. Learning, “Model -based and Model -free RL for Robot Control,” Lect. Notes Stanford Univ. (Chapter 3), pp. 1–13, 2019

  9. [17]

    Reinforcement Learning and Control as Probabilistic Inference : Tutorial and Review,

    S. Levine, “Reinforcement Learning and Control as Probabilistic Inference : Tutorial and Review,” arXiv:1805.00909, pp. 1–22, 2018

  10. [18]

    Variational Bayesian Reinforcement Learning with Regret Bounds,

    B. O. Donoghue, “Variational Bayesian Reinforcement Learning with Regret Bounds,” arXiv:1807.09647, pp. 1–22, 2022

  11. [19]

    Interpretable End -to- End Urban Autonomous Driving With Latent Deep Reinforcement Learning,

    J. Chen, S. E. Li, M. Tomizuka, and L. Fellow, “Interpretable End -to- End Urban Autonomous Driving With Latent Deep Reinforcement Learning,” IEEE Trans. Intell. Transp. Syst., vol. 23, no. 6, pp. 5068 – 5078, 2022

  12. [20]

    Interpretable Decision -Making for Autonomous Vehicles at Highway On -Ramps With Latent Space Reinforcement Learning,

    H. Wang et al. , “Interpretable Decision -Making for Autonomous Vehicles at Highway On -Ramps With Latent Space Reinforcement Learning,” IEEE Trans. Veh. Technol., vol. 70, no. 9, pp. 8707 –8719, 2021

  13. [21]

    Making Sense of Reinforcement Learning and Probabilistic Inference,

    I. Osband and C. Ionescu, “Making Sense of Reinforcement Learning and Probabilistic Inference,” Eighth Int. Conf. Learn. Represent. (ICLR 2020), Apr 26th through May 1st Virtual Only Conf., pp. 1–16, 2020

  14. [22]

    Bayesian Reinforcement Learning : A Survey,

    M. Ghavamzadeh and S. M. Technion, “Bayesian Reinforcement Learning : A Survey,” arXiv:1609.04436, pp. 1–147, 2016

  15. [23]

    State estimation for robotics,

    T. D. Barfoot, “State estimation for robotics,” Cambridge Univ. Press. Ca, pp. 1–399, 2022

  16. [24]

    Variational Inference: A Review for Statisticians,

    D. M. Blei, A. Kucukelbir, and J. D. McAuliffe, “Variational Inference: A Review for Statisticians,” J. Am. Stat. Assoc. , vol. 112, no. 518, pp. 859–877, 2017

  17. [25]

    Taking the Human Out of the Loop: A Review of Bayesian Optimization,

    B. Shahriari, K. Swersky, Z. Wang, R. P. Adams, and N. De Freitas, “Taking the Human Out of the Loop: A Review of Bayesian Optimization,” Proc. IEEE, vol. 104, no. 1, pp. 148–175, 2016

  18. [26]

    Bayesian Neural Networks: An Introduction and Survey,

    E. Goan, C. Fookes, and M. L. Jun, “Bayesian Neural Networks: An Introduction and Survey,” arXiv:2006.12024, pp. 1–44, 2020

  19. [27]

    A Survey of Deep Active Learning,

    P. Ren et al. , “A Survey of Deep Active Learning,” arXiv:2009.00236v2, pp. 1–40, 2021

  20. [28]

    Model -based Multi -agent Reinforcement Learning: Recent Progress and Prospects,

    X. Wang, Z. Zhang, and W. Zhang, “Model -based Multi -agent Reinforcement Learning: Recent Progress and Prospects,” arXiv:2203.1060, pp. 1–8, 2022

  21. [29]

    A Survey of Multi -Task Deep Reinforcement Learning,

    N. V. Varghese and Q. H. Mahmoud, “A Survey of Multi -Task Deep Reinforcement Learning,” Electronics, vol. 9, no. 9, p. 1363, 2020

  22. [30]

    A Survey of Meta -Reinforcement Learning,

    J. Beck et al. , “A Survey of Meta -Reinforcement Learning,” arXiv:2301.08028, pp. 1–53, 2023

  23. [31]

    Practical Realization of Bessel’s Correction for a Bias - Free Estimation of the Auto -Covariance and the Cross -Covariance Functions,

    H. Nobach, “Practical Realization of Bessel’s Correction for a Bias - Free Estimation of the Auto -Covariance and the Cross -Covariance Functions,” arXiv:2303.11047, pp. 1–17, 2023

  24. [32]

    Bayes’ Theorem,

    M. F. Triola, “Bayes’ Theorem,” Metaphys. Res. Lab, Stanford Univ., pp. 1–9, 2021

  25. [33]

    Predicting human navigation goals based on Bayesian inference and activity regions,

    L. Bruckschen, K. Bungert, N. Dengler, and M. Bennewitz, “Predicting human navigation goals based on Bayesian inference and activity regions,” Rob. Auton. Syst., vol. 134, p. 103664, 2020

  26. [34]

    Adaptive Bayesian inference system for recognition of walking activities and prediction of gait events using wearable sensors,

    U. Martinez -hernandez and A. A. Dehghani -sanij, “Adaptive Bayesian inference system for recognition of walking activities and prediction of gait events using wearable sensors,” Neural Networks, vol. 102, pp. 107–119, 2018

  27. [35]

    Safety Assurances for Human -Robot Interaction via Confidence -aware Game-theoretic Human Models,

    R. Tian, L. Sun, A. Bajcsy, M. Tomizuka, and A. D. Dragan, “Safety Assurances for Human -Robot Interaction via Confidence -aware Game-theoretic Human Models,” pp. 11229–11235, 2022

  28. [36]

    Bayesian generalized kernel inference for occupancy map prediction Bayesian Generalized Kernel Inference for Occupancy Map Prediction,

    K. Doherty, J. Wang, and B. Englot, “Bayesian generalized kernel inference for occupancy map prediction Bayesian Generalized Kernel Inference for Occupancy Map Prediction,” 2017 IEEE Int. Conf. Robot. Autom. (ICRA),29 May 2017 - 03 June 2017,Singapore, 2017

  29. [37]

    Gaussian Processes for Machine Learning,

    C. E. Rasmussen, C. K. I. Williams, G. Processes, M. I. T. Press, and M. I. Jordan, “Gaussian Processes for Machine Learning,” Cambridge, MIT Press, 2006

  30. [38]

    The pitfalls of using Gaussian Process Regression for normative modeling,

    B. Xu, R. Kuplicki, S. Sen, and M. P. Paulus, “The pitfalls of using Gaussian Process Regression for normative modeling,” PLoS One, vol. 16, pp. 1–14, 2021

  31. [39]

    Non -Gaussian Process Regression,

    Y. Kındap and S. Godsill, “Non -Gaussian Process Regression,” arXiv:2209.03117, pp. 1–16, 2022

  32. [40]

    Bayesian nonparametric kernel -learning,

    J. B. Oliva, A. Dubey, A. G. Wilson, B. Póczos, J. Schneider, and E. P. Xing, “Bayesian nonparametric kernel -learning,” Proc. 19th Int. Conf. Artif. Intell. Stat. AISTATS 2016, Cadiz, Spain , vol. 41, pp. 1078–1086, 2016

  33. [41]

    Artificial intelligence: A Modern Approach (Third Edition),

    E. Davis et al. , “Artificial intelligence: A Modern Approach (Third Edition),” Pearson Educ. Inc, Up. Saddle River, New Jersey, United States Am., 2010

  34. [42]

    Gaussian process dynamical models for human motion,

    J. M. Wang, D. J. Fleet, and A. Hertzmann, “Gaussian process dynamical models for human motion,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 30, no. 2, pp. 283–298, Feb. 2008

  35. [43]

    Meta -Learning Priors for Efficient Online Bayesian Regression,

    J. Harrison, A. Sharma, and M. Pavone, “Meta -Learning Priors for Efficient Online Bayesian Regression,” arXiv:1807.08912, pp. 1 –28, 2018

  36. [44]

    On the sub -Gaussianity of the Beta and Dirichlet distributions,

    O. Marchal, J. Arbel, J. Monnet, I. C. Jordan, and L. J. Kuntzmann, “On the sub -Gaussianity of the Beta and Dirichlet distributions,” arXiv:1705.00048, pp. 1–13, 2017

  37. [45]

    Bayesian approaches to Gaussian mixture modeling,

    S. J. Roberts, D. Husmeier, I. Rezek, and W. Penny, “Bayesian approaches to Gaussian mixture modeling,” IEEE Trans. Pattern Anal. Mach. Intell., vol. 20, no. 11, pp. 1133–1142, 1998

  38. [46]

    Hierarchical Gaussian Mixture Model,

    F. N. Vincent Garcia and R. Nock, “Hierarchical Gaussian Mixture Model,” Proc. IEEE Int. Conf. Acoust. Speech, Signal Process. ICASSP 2010, 14 -19 March 2010, Sherat. Dallas Hotel. Dallas, Texas, USA, pp. 1–4, 2010

  39. [47]

    Hierarchical Clustering of a Mixture Model,

    J. Goldberger and S. Roweis, “Hierarchical Clustering of a Mixture Model,” Proc. 17th Int. Conf. Neural Inf. Process. Syst. December 2004, pp. 505–512, 2014

  40. [48]

    Hierarchical Gaussian Mixture Model with Objects Attached to Terminal and Non -terminal Dendrogram Nodes,

    L. P. Olech and M. Paradowski, “Hierarchical Gaussian Mixture Model with Objects Attached to Terminal and Non -terminal Dendrogram Nodes,” arXiv:1603.08342, pp. 1–10, 2016

  41. [49]

    HGMR : Hierarchical Gaussian Mixtures for Adaptive 3D Registration,

    B. Eckart, K. Kim, and J. Kautz, “HGMR : Hierarchical Gaussian Mixtures for Adaptive 3D Registration,” 15th Eur. Conf. Munich, Ger. Sept. 8-14, 2018, pp. 730–746, 2018

  42. [50]

    Flexible Hierarchical Gaussian Mixture Model for High -Resolution Remote Sensing Image Segmentation,

    H. R. Sensing and I. Segmentation, “Flexible Hierarchical Gaussian Mixture Model for High -Resolution Remote Sensing Image Segmentation,” Remote Sens., vol. 12, no. 7, p. 1219, 2020

  43. [51]

    A BAYESIAN HIERARCHICAL MIXTURE OF GAUSSIAN MODEL FOR MULTI-SPEAKER DOA ESTIMATION AND SEPARATION,

    Y. Laufer and S. Gannot, “A BAYESIAN HIERARCHICAL MIXTURE OF GAUSSIAN MODEL FOR MULTI-SPEAKER DOA ESTIMATION AND SEPARATION,” 2020 IEEE Int. Work. Mach. Learn. SIGNAL Process. SEPT. 21–24, 2020, ESPOO, Finl., 2020

  44. [52]

    Bayesian hierarchical mixture models for detecting non -normal clusters applied to noisy genomic and environmental datasets,

    H. Zhang, B. Swallow, and M. Gupta, “Bayesian hierarchical mixture models for detecting non -normal clusters applied to noisy genomic and environmental datasets,” Aust. N. Z. J. Stat. , vol. 64, no. 2, pp. 313–337, 2022

  45. [53]

    A tutorial on Dirichlet process mixture modeling,

    Y. Li, E. Schofield, and M. Gönen, “A tutorial on Dirichlet process mixture modeling,” J. Math. Psychol., vol. 91, pp. 128–144, 2019

  46. [54]

    A tutorial on Bayesian nonparametric models,

    S. J. Gershman and D. M. Blei, “A tutorial on Bayesian nonparametric models,” J. Math. Psychol. , vol. 56, no. 1, pp. 1 –12, 2012

  47. [55]

    Context -Based Meta - Reinforcement Learning with Bayesian Nonparametric Models,

    Z. Bing, Y. Yun, K. Huang, and A. Knoll, “Context -Based Meta - Reinforcement Learning with Bayesian Nonparametric Models,” IEEE Trans. Pattern Anal. Mach. Intell., 2024

  48. [56]

    Model Selection for Mixture Models -Perspectives and Strategies,

    G. Celeux et al., “Model Selection for Mixture Models -Perspectives and Strategies,” Handb. Mix. Anal. CRC Press, 2018

  49. [57]

    Priors in Bayesian Deep Learning : A Review,

    V. Fortuin, “Priors in Bayesian Deep Learning : A Review,” Int. Stat. Rev., vol. 90, no. 3, pp. 563–591, 2022

  50. [58]

    Bayesian Model-Agnostic Meta -Learning,

    J. Yoon, T. Kim, O. Dia, S. Kim, Y. Bengio, and S. Ahn, “Bayesian Model-Agnostic Meta -Learning,” 32nd Conf. Neural Inf. Process. Syst. (NeurIPS 2018), Montréal, Canada, pp. 1–11, 2018

  51. [59]

    Finite Mixture Models,

    G. J. Mclachlan, S. X. Lee, and S. I. Rathnayake, “Finite Mixture Models,” Annu. Rev., vol. 6, pp. 355–378, 2019

  52. [60]

    On the convergence of coordinate ascent variational inference,

    A. Bhattacharya, D. Pati, Y. Yang, and M. L. Jun, “On the convergence of coordinate ascent variational inference,” arXiv:2306.01122, pp. 1–47

  53. [61]

    Monte Carlo co -ordinate 19 Zhou, C. et al. Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review ascent variational inference,

    L. Ye, A. Beskos, M. De Iorio, and J. Hao, “Monte Carlo co -ordinate 19 Zhou, C. et al. Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review ascent variational inference,” Stat. Comput., vol. 30, no. 4, pp. 887 – 905, 2020

  54. [62]

    Stochastic Optimization: A Review,

    D. Fouskakis and D. Draper, “Stochastic Optimization: A Review,” Int. Stat. Rev. / Rev. Int. Stat., vol. 70, no. 3, pp. 315–349, Jul. 2002

  55. [63]

    Stochastic Variational Inference,

    M. D. Hoffman, D. M. Blei, C. Wang, and J. Paisley, “Stochastic Variational Inference,” J. Mach. Learn. Res., vol. 14, pp. 1303 –1347, 2013

  56. [64]

    Practical Bayesian optimization of machine learning algorithms,

    B. J. Snoek, H. Larochelle, and R. P. Adams, “Practical Bayesian optimization of machine learning algorithms,” arXiv:1206.2944, pp. 1–12, 2012

  57. [65]

    Computing the racing line using Bayesian optimization,

    A. Jain and M. Morari, “Computing the racing line using Bayesian optimization,” 2020 59th IEEE Conf. Decis. Control (CDC),14 -18 December 2020, Jeju, Korea, 2020

  58. [66]

    A New Method of Locating the Maximum Point of an Arbitrary Multipeak Curve in the Presence of Noise,

    H. J. Kushner, “A New Method of Locating the Maximum Point of an Arbitrary Multipeak Curve in the Presence of Noise,” J. Fluids Eng., vol. 86, no. 1, pp. 97–106, 1964

  59. [67]

    The application of Bayesian methods for seeking the extremum,

    J. Mockus, “The application of Bayesian methods for seeking the extremum,” Towar. Glob. Optim., vol. 2, p. 117, 1998

  60. [68]

    Generating Adversarial Driving Scenarios in High-Fidelity Simulators,

    Y. Abeysirigoonawardena, F. Shkurti, and G. Dudek, “Generating Adversarial Driving Scenarios in High-Fidelity Simulators,” 2019 Int. Conf. Robot. Autom. (ICRA),20-24 May 2019,Montreal, QC, Canada, 2019

  61. [69]

    Greed is Good: Exploration and Exploitation Trade -offs in Bayesian Optimisation,

    G. D. E. Ath, R. M. Everson, A. A. M. Rahat, and J. E. Fieldsend, “Greed is Good: Exploration and Exploitation Trade -offs in Bayesian Optimisation,” arXiv:1911.12809v2, pp. 1–49, 2021

  62. [70]

    Upper confidence bound and pure exploration,

    E. Contal, D. Buffoni, A. Robicquet, and N. Vayatis, “Upper confidence bound and pure exploration,” arXiv:1304.5350, pp. 1 –16, 2013

  63. [71]

    Entropy Search for Information - Efficient Global Optimization,

    P. Hennig and C. J. Schuler, “Entropy Search for Information - Efficient Global Optimization,” J. Mach. Learn. Res. , vol. 13, pp. 1809–1837, 2012

  64. [72]

    Analysis of Thompson Sampling for the Multi -armed Bandit Problem,

    S. Agrawal, “Analysis of Thompson Sampling for the Multi -armed Bandit Problem,” JMLR Work. Conf. Proc., vol. 23, no. 39, pp. 1 –26, 2012

  65. [73]

    Sparse Spectrum Gaussian Process Regression,

    M. Lázaro-gredilla, J. Quiñonero-candela, C. E. Rasmussen, and A. R. Figueiras-Vidal, “Sparse Spectrum Gaussian Process Regression,” J. Mach. Learn. Res., vol. 11, pp. 1865–1881, 2010

  66. [74]

    Predictive Entropy Search for Efficient Global Optimization of Black-box Functions,

    M. W. Hoffman, “Predictive Entropy Search for Efficient Global Optimization of Black-box Functions,” arXiv:1406.2541v1, pp. 1–12, 2014

  67. [75]

    Max -value Entropy Search for Efficient Bayesian Optimization,

    Z. Wang and S. Jegelka, “Max -value Entropy Search for Efficient Bayesian Optimization,” arXiv:1703.01968v3, pp. 1–12, 2018

  68. [76]

    Portfolio Allocation for Bayesian Optimization,

    E. Brochu, M. Hoffman, and N. De Freitas, “Portfolio Allocation for Bayesian Optimization,” arXiv:1009.5419, pp. 1–20, 2011

  69. [77]

    An Entropy Search Portfolio for Bayesian Optimization,

    B. Shahriari, Z. Wang, M. W. Hoffman, A. Bouchard-Côté, and N. de Freitas, “An Entropy Search Portfolio for Bayesian Optimization,” arXiv:1406.4625v4, pp. 1–10, 2014

  70. [78]

    Deep Gaussian processes,

    A. C. Damianou and N. D. Lawrence, “Deep Gaussian processes,” Proc. 16th Int. Conf. Artif. Intell. Stat. 2013, Scottsdale, AZ, USA , 2013

  71. [79]

    Transforming Neural -Net Output Levels to Probability Distributions,

    J. S. Denker and Y. LeCun, “Transforming Neural -Net Output Levels to Probability Distributions,” Adv. Neural Inf. Process. Syst. 3 , pp. 853–859, 1991

  72. [80]

    Understanding the Metropolis -Hastings algorithm,

    S. Chib and E. Greenberg, “Understanding the Metropolis -Hastings algorithm,” Am. Stat., vol. 49, no. 4, pp. 327–336, 1995

  73. [81]

    Explaining the Gibbs Sampler,

    G. Casella and E. I. George, “Explaining the Gibbs Sampler,” Am. Stat., vol. 46, no. 3, pp. 167–174, Jul. 1992

  74. [82]

    Mean Field Variational Bayes for Elaborate Distributions,

    M. P. Wand, J. T. Ormerod, S. A. Padoan, and R. Fr, “Mean Field Variational Bayes for Elaborate Distributions,” Bayesian Anal., vol. 6, no. 4, pp. 847–900, 2011

  75. [83]

    The Variational Gaussian Approximation Revisited,

    M. Opper and C. Archambeau, “The Variational Gaussian Approximation Revisited,” Neural Comput., vol. 21, no. 3, pp. 786 – 792, 2009

  76. [84]

    The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables,

    C. J. Maddison, A. Mnih, Y. W. Teh, U. Kingdom, and U. Kingdom, “The Concrete Distribution: A Continuous Relaxation of Discrete Random Variables,” arXiv:1611.00712, pp. 1–20, 2017

  77. [85]

    Do Bayesian Neural Networks Need To Be Fully Stochastic?,

    M. Sharma, S. Farquhar, E. Nalisnick, and T. Rainforth, “Do Bayesian Neural Networks Need To Be Fully Stochastic?,” arXiv:2211.06291v2, vol. 206, pp. 1–29, 2023

  78. [86]

    Dropout: A simple way to prevent neural networks from overfitting,

    N. Srivastava, G. Hinton, A. Krizhevsky, I. Sutskever, and R. Salakhutdinov, “Dropout: A simple way to prevent neural networks from overfitting,” J. Mach. Learn. Res., vol. 15, pp. 1929–1958, 2014

  79. [87]

    Computing with infinite networks,

    C. K. I. Williams, “Computing with infinite networks,” Adv. Neural Inf. Process. Syst., pp. 295–301, 1997

  80. [88]

    Dropout as a Bayesian approximation: Representing model uncertainty in deep learning,

    Y. Gal and Z. Ghahramani, “Dropout as a Bayesian approximation: Representing model uncertainty in deep learning,” 33rd Int. Conf. Mach. Learn. ICML 2016, vol. 3, pp. 1651–1660, 2016

  81. [89]

    A Decentralized Bayesian Approach for Snake Robot Control,

    Y. Jia and S. Ma, “A Decentralized Bayesian Approach for Snake Robot Control,” IEEE Robot. Autom. Lett. , vol. 6, no. 4, pp. 6955 – 6960, 2021

  82. [90]

    Notes on the Behavior of MC Dropout,

    F. Verdoja and V. Kyrki, “Notes on the Behavior of MC Dropout,” ICML 2021 Work. Uncertain. Robustness Deep Learn., 2021

  83. [91]

    Bayesian deep convolutional networks with many channels are Gaussian processes,

    R. Novak et al., “Bayesian deep convolutional networks with many channels are Gaussian processes,” arXiv:1810.05148, pp. 1–35, 2019

  84. [92]

    Deep Predictive Models for Collision Risk Assessment in Autonomous Driving,

    M. Strickland, G. Fainekos, and H. B. Amor, “Deep Predictive Models for Collision Risk Assessment in Autonomous Driving,” in 2018 IEEE International Conference on Robotics and Automation (ICRA), 2018, pp. 4685–4692

  85. [93]

    Fully Bayesian Recurrent Neural Networks for Safe Reinforcement Learning,

    M. Benatan and E. O. Pyzer-knapp, “Fully Bayesian Recurrent Neural Networks for Safe Reinforcement Learning,” arXiv:1911.03308, pp. 1–10, 2019

  86. [94]

    Deep similarity-based batch mode active learning with exploration-exploitation,

    C. Yin et al., “Deep similarity-based batch mode active learning with exploration-exploitation,” Proc. - IEEE Int. Conf. Data Mining, ICDM, pp. 575–584, 2017

  87. [95]

    Active Learning Literature Survey,

    B. Settles, “Active Learning Literature Survey,” Tech. Report. Univ. Wisconsin-Madison Dep. Comput. Sci., pp. 1–47, 2009

  88. [96]

    Active Learning for Convolutional Neural Networks: A Core-Set Approach,

    N. E. A. C. Ore, E. T. A. Pproach, O. Sener, and S. Savarese, “Active Learning for Convolutional Neural Networks: A Core-Set Approach,” arXiv:1708.00489, pp. 1–13, 2018

  89. [97]

    Bayesian Active Learning for Classification and Preference Learning,

    N. Houlsby, F. Huszár, Z. Ghahramani, and M. Lengyel, “Bayesian Active Learning for Classification and Preference Learning,” arXiv:1112.5745, pp. 1–17, 2011

  90. [98]

    BatchBALD: Efficient and diverse batch acquisition for deep Bayesian active learning,

    A. Kirsch, J. van Amersfoort, and Y. Gal, “BatchBALD: Efficient and diverse batch acquisition for deep Bayesian active learning,” 33rd Conf. Neural Inf. Process. Syst. (NeurIPS 2019), Vancouver, Canada., 2019

  91. [99]

    Deep Bayesian Active Learning with Image Data,

    Y. Gal, R. Islam, and Z. Ghahramani, “Deep Bayesian Active Learning with Image Data,” arXiv:1703.02910, pp. 1–10, 2017

  92. [100]

    VEEGAN: Reducing Mode Collapse in GANs using Implicit Variational Learning,

    A. Srivastava, C. Russell, L. Valkov, M. U. Gutmann, and C. Sutton, “VEEGAN: Reducing Mode Collapse in GANs using Implicit Variational Learning,” arXiv:1705.07761, pp. 1–17, 2017

  93. [101]

    Learning how to actively learn: A deep imitation learning approach,

    M. Liu, W. Buntine, and G. Haffari, “Learning how to actively learn: A deep imitation learning approach,” ACL 2018 - 56th Annu. Meet. Assoc. Comput. Linguist. Proc. Conf. (Long Pap. Melbourne, Aust. July 15 - 20, 2018, pp. 1874–1883, 2018

  94. [102]

    Deep Active Learning with Adaptive Acquisition,

    M. Haußmann, F. Hamprecht, and M. Kandemir, “Deep Active Learning with Adaptive Acquisition,” arXiv:1906.11471, pp. 1 –7, 2019

  95. [103]

    Deep Reinforcement Active Learning for Human -In-The-Loop Person Re -Identification,

    H. P. Re-identification, Z. Liu, J. Wang, S. Gong, H. Lu, and D. Tao, “Deep Reinforcement Active Learning for Human -In-The-Loop Person Re -Identification,” 2019 IEEE/CVF Int. Conf. Comput. Vis. (ICCV), 27 Oct. 2019 - 02 Novemb. 2019,Seoul, Korea, 2019

  96. [104]

    ActiveLink: Deep Active Learning for Link Prediction in Knowledge Graphs,

    P. Cudré -mauroux, “ActiveLink: Deep Active Learning for Link Prediction in Knowledge Graphs,” Pro- ceedings ofthe 2019 World Wide Web Conf. (WWW’19), May 13 –17, 2019, San Fr. CA, USA , 2019

  97. [105]

    Meta -Learning Transferable Active Learning Policies by Deep Reinforcement Learning,

    K. Pang, M. Dong, Y. Wu, and T. Hospedales, “Meta -Learning Transferable Active Learning Policies by Deep Reinforcement Learning,” arXiv:1806.04798, pp. 1–8, 2018

  98. [106]

    Language Models are Unsupervised Multitask Learners,

    A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, and I. Sutskever, “Language Models are Unsupervised Multitask Learners,” Feb. 2019

  99. [107]

    Generative Adversarial Networks,

    I. J. Goodfellow et al. , “Generative Adversarial Networks,” arXiv:1406.2661v1, pp. 1–9, Jun. 2014

  100. [108]

    Auto -encoding variational bayes,

    D. P. Kingma and M. Welling, “Auto -encoding variational bayes,” 2nd Int. Conf. Learn. Represent. ICLR 2014 - Conf. Track Proc., pp. 1–14, 2014

  101. [109]

    On the Design Fundamentals of Diffusion Models: A Survey,

    Z. Chang, G. A. Koulieris, and H. P. H. Shum, “On the Design Fundamentals of Diffusion Models: A Survey,” arXiv:2306.04542v3, pp. 1–22, Jun. 2023

  102. [110]

    Socially Adaptive Path Planning Based on Generative Adversarial Network,

    Y. Wang, Y. Kong, W. Chi, and L. Sun, “Socially Adaptive Path Planning Based on Generative Adversarial Network,” arXiv:2404.18687v1, pp. 1–12, 2024

  103. [111]

    Maximum Likelihood: An Introduction,

    L. Le Cam, “Maximum Likelihood: An Introduction,” Int. Stat. Rev. / Rev. Int. Stat., vol. 58, no. 2, pp. 153–171, Jul. 1990

  104. [112]

    Denoising Diffusion Probabilistic Models,

    J. Ho, A. Jain, and P. Abbeel, “Denoising Diffusion Probabilistic Models,” arXiv:2006.11239v2, pp. 1–25, Jun. 2020

  105. [113]

    Score -Based Generative Modeling through Stochastic Differential Equations,

    Y. Song, J. Sohl -Dickstein, D. P. Kingma, A. Kumar, S. Ermon, and B. Poole, “Score -Based Generative Modeling through Stochastic Differential Equations,” arXiv:2011.13456v2, pp. 1–36, Nov. 2020

  106. [114]

    U -Net: Convolutional Networks for Biomedical Image Segmentation,

    O. Ronneberger, P. Fischer, and T. Brox, “U -Net: Convolutional Networks for Biomedical Image Segmentation,” May 2015

  107. [115]

    Transformers are Meta -Reinforcement Learners,

    L. C. Melo, “Transformers are Meta -Reinforcement Learners,” arXiv:2206.06614, pp. 1–20, 2022

  108. [116]

    Amortized Bayesian Meta -Learning,

    S. Ravi and A. Beatson, “Amortized Bayesian Meta -Learning,” Seventh Int. Conf. Learn. Represent. (ICLR 2019),Mon May 6th through Thu 9th, Ernest N. Morial Conv. Center, New Orleans, US , pp. 1–14, 2019

  109. [117]

    Model -Agnostic Meta-Learning for Fast Adaptation of Deep Networks,

    C. Finn, P. Abbeel, and S. Levine, “Model -Agnostic Meta-Learning for Fast Adaptation of Deep Networks,” arXiv:1703.03400, pp. 1–13, 2017

  110. [118]

    One -Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning,

    T. Yu et al. , “One -Shot Imitation from Observing Humans via Domain-Adaptive Meta-Learning,” arXiv:1802.01557, pp. 1–12, 2018

  111. [119]

    Meta -Learning Recipe, Black -Box Adaptation, Optimization-Based Approaches,

    C. Finn, “Meta -Learning Recipe, Black -Box Adaptation, Optimization-Based Approaches,” Lect. Notes Stanford Univ. , pp. 1 – 20 Zhou, C. et al. Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review 34, 2019

  112. [120]

    Conditional neural processes,

    M. Gamelo et al. , “Conditional neural processes,” 35th Int. Conf. Mach. Learn. ICML 2018, vol. 4, pp. 2738–2747, 2018

  113. [121]

    Meta-Learning with Memory -Augmented Neural Networks,

    A. Santoro, S. Bartunov, M. Botvinick, D. Wierstra, and T. Lillicrap, “Meta-Learning with Memory -Augmented Neural Networks,” Proc. 33rd Int. Conf. Int. Conf. Mach. Learn., vol. 48, pp. 1842–1850, 2016

  114. [122]

    Continual lifelong learning with neural networks: A review,

    G. I. Parisi, R. Kemker, J. L. Part, C. Kanan, and S. Wermter, “Continual lifelong learning with neural networks: A review,” Neural Networks, vol. 113. Elsevier Ltd, pp. 54–71, 01-May-2019

  115. [123]

    Progressive Neural Networks,

    A. A. Rusu et al., “Progressive Neural Networks,” Jun. 2016

  116. [124]

    PackNet: Adding Multiple Tasks to a Single Network by Iterative Pruning,

    A. Mallya and S. Lazebnik, “PackNet: Adding Multiple Tasks to a Single Network by Iterative Pruning,” Nov. 2017

  117. [125]

    A Dirichlet Process Mixture of Robust Task Models for Scalable Lifelong Reinforcement Learning,

    Z. Wang, C. Chen, and D. Dong, “A Dirichlet Process Mixture of Robust Task Models for Scalable Lifelong Reinforcement Learning,” IEEE Trans. Cybern., vol. 53, no. 12, pp. 7509–7520, Dec. 2023

  118. [126]

    Lifelong Incremental Reinforcement Learning With Online Bayesian Inference,

    Z. Wang, C. Chen, and D. Dong, “Lifelong Incremental Reinforcement Learning With Online Bayesian Inference,” IEEE Trans. Neural Networks Learn. Syst. , vol. 33, no. 8, pp. 4003 –4016, 2022

  119. [127]

    A Bayesian Mixture Model of Temporal Point Processes with Determinantal Point Process Prior,

    Y. Dong, S. Ye, Y. Cao, Q. Han, H. Xu, and H. Yang, “A Bayesian Mixture Model of Temporal Point Processes with Determinantal Point Process Prior,” Nov. 2024

  120. [128]

    A New Approach to Linear Filtering and Prediction Problems,

    R. E. Kalman, “A New Approach to Linear Filtering and Prediction Problems,” Trans. ASME–Journal Basic Eng. , vol. 82, no. Series D, pp. 35–45, 1960

  121. [129]

    [Re] The Discriminative Kalman Filter for Bayesian Filtering with Nonlinear and Non-Gaussian Observation Models,

    J. Casco -Rodriguez, C. Kemere, and R. G. Baraniuk, “[Re] The Discriminative Kalman Filter for Bayesian Filtering with Nonlinear and Non-Gaussian Observation Models,” arXiv:2401.14429v1, pp. 1– 12, 2024

  122. [130]

    Stochastic processes and filtering theory,

    K. Senne, “Stochastic processes and filtering theory,” IEEE Trans. Automat. Contr., vol. 17, no. 5, pp. 752–753, 1972

  123. [131]

    Cooperation-Aware Reinforcement Learning for Merging in Dense Traffic,

    M. Bouton, A. Nakhaei, K. Fujimura, and M. J. Kochenderfer, “Cooperation-Aware Reinforcement Learning for Merging in Dense Traffic,” 2019 IEEE Intell. Transp. Syst. Conf. (ITSC),27 -30 Oct. 2019,Auckland, New Zeal., 2019

  124. [132]

    Chasing as Ghosts Bayesian State Tracking Chasing Ghosts : Instruction Following Chasing Ghosts,

    B. S. Tracking, “Chasing as Ghosts Bayesian State Tracking Chasing Ghosts : Instruction Following Chasing Ghosts,” arXiv:1907.02022, pp. 1–11, 2019

  125. [133]

    Spiking Neural Network on Neuromorphic Hardware for Energy -Efficient Unidimensional SLAM,

    G. Tang, A. Shah, and K. P. Michmizos, “Spiking Neural Network on Neuromorphic Hardware for Energy -Efficient Unidimensional SLAM,” 2019 IEEE/RSJ Int. Conf. Intell. Robot. Syst. (IROS), 03 -08 Novemb. 2019,Macau, China, 2019

  126. [134]

    Real-time Deep Learning of Robotic Manipulator Inverse Dynamics,

    A. S. Polydoros, L. Nalpantidis, and V. Kr, “Real-time Deep Learning of Robotic Manipulator Inverse Dynamics,” 2015 IEEE/RSJ Int. Conf. Intell. Robot. Syst. (IROS),28 Sept. 2015 - 02 Oct. 2015,Hamburg, Ger., 2015

  127. [135]

    Recursive Bayesian Human Intent Recognition in Shared -Control Robotics,

    S. Jain, B. Argall, S. Jain, C. Science, and S. R. Ability -lab, “Recursive Bayesian Human Intent Recognition in Shared -Control Robotics,” 2018 IEEE/RSJ Int. Conf. Intell. Robot. Syst. (IROS),01-05 Oct. 2018,Madrid, Spain, pp. 3905–3912, 2020

  128. [136]

    Goal Inference Improves Objective and Perceived Performance in Human-Robot Collaboration,

    C. Liu, “Goal Inference Improves Objective and Perceived Performance in Human-Robot Collaboration,” arXiv:1802.01780, pp. 1–9, 2018

  129. [137]

    GLMP- Realtime Pedestrian Path Prediction using Global and Local Movement Patterns,

    A. Bera, S. Kim, T. Randhavane, S. Pratapa, and D. Manocha, “GLMP- Realtime Pedestrian Path Prediction using Global and Local Movement Patterns,” 2016 IEEE Int. Conf. Robot. Autom. (ICRA),16 - 21 May 2016,Stockholm, Sweden, 2016

  130. [138]

    Probabilistic inference of human arm reaching target for effective human -robot collaboration,

    A. M. Zanchettin and P. Rocco, “Probabilistic inference of human arm reaching target for effective human -robot collaboration,” 2017 IEEE/RSJ Int. Conf. Intell. Robot. Syst. (IROS),24 -28 Sept. 2017,Vancouver, BC, Canada, 2017

  131. [139]

    Kalman Filter and Its Application,

    Q. Li, R. Li, K. Ji, and W. Dai, “Kalman Filter and Its Application,” in 2015 8th International Conference on Intelligent Networks and Intelligent Systems (ICINIS), 2015, pp. 74–77

  132. [140]

    The Monte Carlo Method,

    N. Metropolis and S. Ulam, “The Monte Carlo Method,” J. Am. Stat. Assoc., vol. 44, no. 247, pp. 335–341, Jul. 1949

  133. [141]

    Monte Carlo Sampling Methods Using Markov Chains and Their Applications,

    W. K. Hastings, “Monte Carlo Sampling Methods Using Markov Chains and Their Applications,” Biometrika, vol. 57, no. 1, pp. 97 – 109, Jul. 1970

  134. [142]

    Risk -aware Control for Robots with Non - Gaussian Belief Spaces,

    M. Vahs and J. Tumova, “Risk -aware Control for Robots with Non - Gaussian Belief Spaces,” arXiv:2309.12857v2, pp. 1–7, 2023

  135. [143]

    Inducing Cooperation via Team Regret Minimization based Multi-Agent Deep Reinforcement Learning,

    R. Yu et al., “Inducing Cooperation via Team Regret Minimization based Multi-Agent Deep Reinforcement Learning,” arXiv:1911.07712, pp. 1–8, 2019

  136. [144]

    Robust Monte Carlo localization for mobile robots,

    S. Thrun, D. Fox, W. Burgard, and F. Dellaert, “Robust Monte Carlo localization for mobile robots,” Artif. Intell., vol. 128, no. 1, pp. 99 – 141, 2001

  137. [145]

    Batch Nonlinear Continuous-Time Trajectory Estimation as Exactly Sparse Gaussian Process Regression,

    S. Anderson, T. D. Barfoot, C. Hay, and T. Simo, “Batch Nonlinear Continuous-Time Trajectory Estimation as Exactly Sparse Gaussian Process Regression,” arXiv:1412.0630, pp. 1–16, 2015

  138. [146]

    A Sliding Window Filter for SLAM,

    G. Sibley, “A Sliding Window Filter for SLAM,” Tech. report, Univ. South. Calif., pp. 1–17, 2006

  139. [147]

    Model - based Reinforcement Learning : A Survey,

    T. M. Moerland, J. Broekens, A. Plaat, and C. M. Jonker, “Model - based Reinforcement Learning : A Survey,” arXiv:2006.16712v4, pp. 1–120, 2022

  140. [148]

    Reinforcement Learning for Partially Observable Linear Gaussian Systems Using Batch Dynamics of Noisy Observations,

    F. A. Yaghmaie, H. Modares, and F. Gustafsson, “Reinforcement Learning for Partially Observable Linear Gaussian Systems Using Batch Dynamics of Noisy Observations,” IEEE Trans. Automat. Contr., 2024

  141. [149]

    The Iterated Sigma Point Kalman Filter with Applications to Long Range Stereo,

    G. Sibley, G. Sukhatme, and L. Matthies, “The Iterated Sigma Point Kalman Filter with Applications to Long Range Stereo,” Proc. Robot. Sci. Syst., pp. 1–8, 2006

  142. [150]

    An Introduction to MCMC for Machine Learning,

    C. ANDRIEU, “An Introduction to MCMC for Machine Learning,” IEEE Int. Conf. Intell. Robot. Syst. , vol. 2017 -Septe, pp. 4144 –4151, 2017

  143. [151]

    Gibbs sampler and coordinate ascent variational inference : A set -theoretical review,

    A. Texas and C. Station, “Gibbs sampler and coordinate ascent variational inference : A set -theoretical review,” arXiv:2008.01006, pp. 1–19, 2020

  144. [152]

    Evidential reasoning using stochastic simulation of causal models,

    J. Pearl, “Evidential reasoning using stochastic simulation of causal models,” Artif. Intell., vol. 32, no. 2, pp. 245–257, 1987

  145. [153]

    Maximum likelihood from incomplete Ddata via the EM algorithm,

    A. P. Dempster, N. M. Laird, and D. B. Rubin, “Maximum likelihood from incomplete Ddata via the EM algorithm,” J. R. Stat. Soc. Ser. B , vol. 39, no. 1, pp. 1–38, 1977

  146. [154]

    The EM algorithm: an old folk -song sung to a fast new tune,

    X. Meng and D. Van Dyk, “The EM algorithm: an old folk -song sung to a fast new tune,” J. R. Stat. Soc. B, 1997

  147. [155]

    The stochastic EM algorithm : Estimation and asymptotic results,

    S. F. Nielsen, “The stochastic EM algorithm : Estimation and asymptotic results,” Bernoulli, vol. 6, no. 3, pp. 457–489, 2000

  148. [156]

    Implementations of the Monte Carlo EM Algorithm,

    R. A. Levine and G. Casella, “Implementations of the Monte Carlo EM Algorithm,” J. Comput. Graph. Stat., vol. 10, no. 3, pp. 422–439, Jul. 2001

  149. [157]

    A Markov Chain Monte Carlo Expectation Maximization Algorithm for Statistical Analysis of DNA Sequence Evolution with Neighbor -Dependent Substitution Rates,

    A. Hobolth, “A Markov Chain Monte Carlo Expectation Maximization Algorithm for Statistical Analysis of DNA Sequence Evolution with Neighbor -Dependent Substitution Rates,” J. Comput. Graph. Stat., vol. 17, no. 1, pp. 138–162, Jul. 2008

  150. [158]

    Model -Based Bayesian Reinforcement Learning in Large Structured Domains,

    S. Ross and J. Pineau, “Model -Based Bayesian Reinforcement Learning in Large Structured Domains,” arXiv:1206.3281, pp. 1 –8, 2012

  151. [159]

    Approximate Bayesian reinforcement learning based on estimation of plant,

    K. Senda, T. Hishinuma, and Y. Tani, “Approximate Bayesian reinforcement learning based on estimation of plant,” Auton. Robots, vol. 44, no. 5, pp. 845–857, 2020

  152. [160]

    Task - Agnostic Online Reinforcement Learning with an Infinite Mixture of Gaussian Processes,

    M. Xu, W. Ding, J. Zhu, Z. Liu, B. Chen, and D. Zhao, “Task - Agnostic Online Reinforcement Learning with an Infinite Mixture of Gaussian Processes,” arXiv:2006.11441, pp. 1–12, 2020

  153. [161]

    Monte Carlo Tree Search : A Review of Recent Modifications and Applications,

    K. Godlewski and B. Sawicki, “Monte Carlo Tree Search : A Review of Recent Modifications and Applications,” arXiv:2103.04931, pp. 1– 99, 2022

  154. [162]

    Efficient Bayes -Adaptive Reinforcement Learning using Sample -Based Search,

    A. Guez, D. Silver, and P. Dayan, “Efficient Bayes -Adaptive Reinforcement Learning using Sample -Based Search,” arXiv:1205.3109, pp. 1–14, 2012

  155. [163]

    The Grand Challenge of Computer Go:Monte Carlo Tree Search and Extensions,

    S. Gelly et al., “The Grand Challenge of Computer Go:Monte Carlo Tree Search and Extensions,” Commun. ACM, vol. 55, no. 3, pp. 106– 113, 2012

  156. [164]

    Near -Bayesian Exploration in Polynomial Time,

    J. Z. Kolter and A. Y. Ng, “Near -Bayesian Exploration in Polynomial Time,” Proc. 26 th Int. Conf. Mach. Learn. Montr. Canada, 2009

  157. [165]

    Variance -based rewards for approximate Bayesian reinforcement learning,

    J. Sorg and R. L. Lewis, “Variance -based rewards for approximate Bayesian reinforcement learning,” arXiv:1203.3518, pp. 1–8, 2012

  158. [166]

    Learning to predict by the methods of Temporal Differences,

    R. S. Sutton, “Learning to predict by the methods of Temporal Differences,” Mach. Learn., vol. 3, pp. 9–44, 1988

  159. [167]

    Policy gradient methods for reinforcement learning with function approximation,

    R. S. Sutton, D. McAllester, S. Singh, and Y. Mansour, “Policy gradient methods for reinforcement learning with function approximation,” Adv. Neural Inf. Process. Syst., pp. 1057–1063, 2000

  160. [168]

    Actor -critic algorithms,

    V. R. Konda and J. N. Tsitsiklis, “Actor -critic algorithms,” Adv. Neural Inf. Process. Syst., pp. 1008–1014, 2000

  161. [169]

    Bayes meets Bellman : The Gaussian process difference learning approach temporal difference learning,

    Y. Engel and R. Melr, “Bayes meets Bellman : The Gaussian process difference learning approach temporal difference learning,” Procedings 20th Int. Conf. Mach. Learn. (ICML-2003), Washingt. DC, 2003

  162. [170]

    A Bayesian Approach to Reinforcement Learning of Vision-Based Vehicular Control,

    Z. Gharaee, “A Bayesian Approach to Reinforcement Learning of Vision-Based Vehicular Control,” arXiv:2104.03807, pp. 1–8, 2021

  163. [171]

    Implicit Posterior Sampling Reinforcement Learning for Continuous Control,

    S. Wang and B. L. B, “Implicit Posterior Sampling Reinforcement Learning for Continuous Control,” Yang, H., Pasupa, K., Leung, A.CS., Kwok, J.T., Chan, J.H., King, I. Neural Inf. Process. ICONIP

  164. [172]

    Bayes –Hermite quadrature,

    A. O’Hagan, “Bayes –Hermite quadrature,” J. Stat. Plan. Inference , vol. 29, no. 3, pp. 245–260, 1991

  165. [173]

    Bayesian Policy Gradient and Actor-Critic Algorithms,

    M. Valko, “Bayesian Policy Gradient and Actor-Critic Algorithms,” J. Mach. Learn. Res., vol. 17, pp. 1–53, 2016

  166. [174]

    AC -Teach : A Bayesian Actor -Critic Method for Policy Learning with an Ensemble of Suboptimal Teachers,

    A. Kurenkov, A. Mandlekar, R. Martin -martin, S. Savarese, and A. Garg, “AC -Teach : A Bayesian Actor -Critic Method for Policy Learning with an Ensemble of Suboptimal Teachers,” arXiv:1909.04121, pp. 1–19, 2019

  167. [175]

    I Know What You Meant : Learning Human Objectives by (Under) estimating Their Choice Set,

    A. Jonnavittula and D. P. Losey, “I Know What You Meant : Learning Human Objectives by (Under) estimating Their Choice Set,” arXiv:2011.06118, pp. 1–7, 2020. 21 Zhou, C. et al. Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review

  168. [176]

    Learning Human Objectives from Sequences of Physical Corrections,

    M. Li, A. Canberk, D. P. Losey, and D. Sadigh, “Learning Human Objectives from Sequences of Physical Corrections,” arXiv:2104.00078, pp. 1–7, 2021

  169. [177]

    Prediction of Reward Functions for Deep Reinforcement Learning via Gaussian Process Regression,

    J. Lim, S. Ha, and J. Choi, “Prediction of Reward Functions for Deep Reinforcement Learning via Gaussian Process Regression,” IEEE/ASME Trans. Mechatronics , vol. 25, no. 4, pp. 1739 –1746, 2020

  170. [178]

    Motion Planning via Bayesian Learning in the Dark,

    C. Quintero -pe, C. Chamzas, V. Unhelkar, and L. E. Kavraki, “Motion Planning via Bayesian Learning in the Dark,” Work. Mach. Learn. Motion Plan. ICRA2021, 2021

  171. [179]

    Bayesian Inference of Temporal Task Specifications from Demonstrations,

    A. Shah, P. Kamath, S. Li, and J. Shah, “Bayesian Inference of Temporal Task Specifications from Demonstrations,” Adv. Neural Inf. Process. Syst. 31 (NeurIPS 2018), 2018

  172. [180]

    PlaNet of the Bayesians : Reconsidering and Improving Deep Planning Network by Incorporating Bayesian Inference,

    M. Okada, N. Kosaka, and T. Taniguchi, “PlaNet of the Bayesians : Reconsidering and Improving Deep Planning Network by Incorporating Bayesian Inference,” arXiv:2003.00370, pp. 1–8, 2020

  173. [181]

    EDGE : Explaining Deep Reinforcement Learning Policies,

    W. Guo, “EDGE : Explaining Deep Reinforcement Learning Policies,” Adv. Neural Inf. Process. Syst. 34, 2021

  174. [182]

    Curiosity - Driven Exploration via Latent Bayesian Surprise,

    P. Mazzaglia, O. Catal, T. Verbelen, and B. Dhoedt, “Curiosity - Driven Exploration via Latent Bayesian Surprise,” arXiv:2104.07495, pp. 1–9, 2022

  175. [183]

    Bayesian Optimization for Iterative Learning,

    M. A. Osborne, “Bayesian Optimization for Iterative Learning,” arXiv:1909.09593, pp. 1–11, 2019

  176. [184]

    Sample -Efficient Robot Motion Learning using Gaussian Process Latent Variable Models

    J. A. Delgado-guerrero, A. Colomé, and C. Torras, “Sample -Efficient Robot Motion Learning using Gaussian Process Latent Variable Models.”

  177. [185]

    Cautious Bayesian Optimization for Efficient and Scalable Policy Search,

    L. P. Fr and M. N. Zeilinger, “Cautious Bayesian Optimization for Efficient and Scalable Policy Search,” Proc. Mach. Learn. Res. , vol. 144, pp. 1–14, 2021

  178. [186]

    A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning,

    E. Brochu, V. M. Cora, and N. De Freitas, “A Tutorial on Bayesian Optimization of Expensive Cost Functions, with Application to Active User Modeling and Hierarchical Reinforcement Learning,” arXiv:10122599v1, pp. 1–49, 2010

  179. [187]

    Two -Stage Bayesian Optimization for Scalable Inference in State Space Models,

    M. Imani and S. F. Ghoreishi, “Two -Stage Bayesian Optimization for Scalable Inference in State Space Models,” IEEE Trans. Neural Networks Learn. Syst., vol. 33, no. 10, pp. 5138–5149, 2022

  180. [188]

    Uncertainty - Guided Active Reinforcement Learning with Bayesian Neural Uncertainty-Guided Active Reinforcement Learning with Bayesian Neural Networks,

    X. Wu, M. El -shamouty, C. Nitsche, and M. F. Huber, “Uncertainty - Guided Active Reinforcement Learning with Bayesian Neural Uncertainty-Guided Active Reinforcement Learning with Bayesian Neural Networks,” 2023 Int. Conf. Robot. Autom. 29, 2023 –Jun 2, 2023, Excel London., 2023

  181. [189]

    VIME: Variational Information Maximizing Exploration,

    R. Houthooft, X. Chen, Y. Duan, J. Schulman, F. De Turck, and P. Abbeel, “VIME: Variational Information Maximizing Exploration,” arXiv:1605.09674, pp. 1–11, 2017

  182. [190]

    Bayesian Exploration in Deep Reinforcement Learning,

    L. Killingberg and H. Langseth, “Bayesian Exploration in Deep Reinforcement Learning,” 2023 Symp. Nor. AI Soc. June 14-15, 2023, Bergen, Norway., 2023

  183. [191]

    Playing Atari with Deep Reinforcement Learning,

    V. Mnih et al., “Playing Atari with Deep Reinforcement Learning,” arXiv, pp. 1–9, 2013

  184. [192]

    Uncertainty Weighted Actor -Critic for Offline Reinforcement Learning,

    Y. Wu, S. Zhai, N. Srivastava, J. Susskind, J. Zhang, and R. Salakhutdinov, “Uncertainty Weighted Actor -Critic for Offline Reinforcement Learning,” arXiv:2105.08140, pp. 1–22, 2021

  185. [193]

    Learning and Policy Search in Stochastic Dynamical Systems with Bayesian Neural Networks,

    S. Depeweg, J. M. Hernández -lobato, F. Doshi -Velez, and S. Udluft, “Learning and Policy Search in Stochastic Dynamical Systems with Bayesian Neural Networks,” arXiv:1605.07127, pp. 1–14, 2017

  186. [194]

    Batch Active Learning with Graph Neural Networks via Multi -Agent Deep Reinforcement Learning,

    Y. Zhang, H. Tong, Y. Xia, Y. Zhu, Y. Chi, and L. Ying, “Batch Active Learning with Graph Neural Networks via Multi -Agent Deep Reinforcement Learning,” Proc. AAAI Conf. Artif. Intell. , vol. 36, no. 8, pp. 9118–9126, 2022

  187. [195]

    S. Chen, Y. Li, and N. M. Kwok, Active vision in robotic systems: A survey of recent developments, vol. 30, no. 11. 2011

  188. [196]

    Active Learning in Robotics: A Review of Control Principles,

    A. T. Taylor, T. A. Berrueta, and T. D. Murphey, “Active Learning in Robotics: A Review of Control Principles,” arXiv:2106.13697, pp. 1– 25, 2021

  189. [197]

    Active Exploration for Inverse Reinforcement Learning,

    D. Lindner, A. Krause, and G. Ramponi, “Active Exploration for Inverse Reinforcement Learning,” arXiv:2207.08645, pp. 1–31, 2023

  190. [198]

    Active imitation learning with noisy guidance,

    K. Brantley, A. Sharaf, and H. Daumé, “Active imitation learning with noisy guidance,” arXiv:2005.12801, pp. 1–14, 2020

  191. [199]

    Active learning from demonstration for robust autonomous navigation,

    D. Silver, J. A. Bagnell, and A. Stentz, “Active learning from demonstration for robust autonomous navigation,” Proc. - IEEE Int. Conf. Robot. Autom., pp. 200–207, 2012

  192. [200]

    APRIL: Active Preference Learning - Based Reinforcement Learning,

    R. Akrour and M. Schoenauer, “APRIL: Active Preference Learning - Based Reinforcement Learning,” arXiv:1208.0984, pp. 1–16, 2012

  193. [201]

    Two kinds of memory signals in neurons of the human hippocampus,

    Z. J. Urgolites et al., “Two kinds of memory signals in neurons of the human hippocampus,” Proc. Natl. Acad. Sci. , vol. 119, no. 19, p. e2115128119, May 2022

  194. [202]

    Soft actor -critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,

    T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft actor -critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor,” 35th Int. Conf. Mach. Learn. ICML 2018, vol. 5, pp. 2976–2989, 2018

  195. [203]

    Self -Consistent Trajectory Autoencoder : Hierarchical Reinforcement Learning with Trajectory Embeddings,

    J. D. Co -Reyes, Y. Liu, A. Gupta, B. Eysenbach, P. Abbeel, and S. Levine, “Self -Consistent Trajectory Autoencoder : Hierarchical Reinforcement Learning with Trajectory Embeddings,” arXiv:1806.02813v1, pp. 1–11, 2018

  196. [204]

    Accelerating Reinforcement Learning with Learned Skill Priors,

    K. Pertsch and J. J. Lim, “Accelerating Reinforcement Learning with Learned Skill Priors,” 4th Conf. Robot Learn. (CoRL 2020), Cambridge MA, USA., pp. 1–17, 2020

  197. [205]

    Efficient Off-Policy Meta -Reinforcement Learning via Probabilistic Context Variables,

    K. Rakelly, A. Zhou, D. Quillen, C. Finn, and S. Levine, “Efficient Off-Policy Meta -Reinforcement Learning via Probabilistic Context Variables,” arXiv:1903.08254, vol. 2019, pp. 1–11

  198. [206]

    Skill - based meta -reinforcement learning,

    T. Nam, S. -H. Sun, K. Pertsch, S. J. Hwang, and J. J. Lim, “Skill - based meta -reinforcement learning,” Tenth Int. Conf. Learn. Represent. (Virtual), Monday, April 25th., pp. 1–23, 2022

  199. [207]

    A Gentle Introduction to Bayesian Analysis: Applications to Developmental Research,

    R. Van de Schoot, D. Kaplan, J. Denissen, J. B. Asendorpf, F. J. Neyer, and M. A. G. van Aken, “A Gentle Introduction to Bayesian Analysis: Applications to Developmental Research,” Child Dev., vol. 85, no. 3, pp. 842–860, 2014

  200. [208]

    Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,

    C. Chi et al. , “Diffusion Policy: Visuomotor Policy Learning via Action Diffusion,” arXiv:2303.04137v5, pp. 1–19, Mar. 2024

  201. [209]

    Planning with Diffusion for Flexible Behavior Synthesis,

    M. Janner, Y. Du, J. B. Tenenbaum, and S. Levine, “Planning with Diffusion for Flexible Behavior Synthesis,” arXiv:2205.09991v2, pp. 1–14, May 2022

  202. [210]

    Hyper -SAMARL: Hypergraph- based Coordinated Task Allocation and Socially-aware Navigation for Multi-Robot Systems,

    W. Wang, A. Bera, and B. -C. Min, “Hyper -SAMARL: Hypergraph- based Coordinated Task Allocation and Socially-aware Navigation for Multi-Robot Systems,” Sep. 2024

  203. [211]

    Learning to reinforcement learn,

    J. X. Wang et al. , “Learning to reinforcement learn,” arXiv:1611.05763, pp. 1–17, 2017

  204. [212]

    Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAML,

    A. Raghu, M. Raghu, S. Bengio, and O. Vinyals, “Rapid Learning or Feature Reuse? Towards Understanding the Effectiveness of MAML,” arXiv:1909.09157, pp. 1–21, 2020

  205. [213]

    Fast Context Adaptation via Meta -Learning,

    L. Zintgraf, K. Shiarlis, V. Kurin, K. Hofmann, and S. Whiteson, “Fast Context Adaptation via Meta -Learning,” arXiv:1810.03642, pp. 1–15, 2019

  206. [214]

    Continuous Adaptation via Meta - Learning in Nonstationary and Competitive Environments,

    I. Mordatch and P. Abbeel, “Continuous Adaptation via Meta - Learning in Nonstationary and Competitive Environments,” arXiv:1710.03641, pp. 1–21, 2018

  207. [215]

    Some Considerations on Learning to Explore via Meta -Reinforcement Learning,

    B. C. Stadie, P. Abbeel, and X. Chen, “Some Considerations on Learning to Explore via Meta -Reinforcement Learning,” arXiv:1803.01118, pp. 1–11, 2019

  208. [216]

    On First -Order Meta - Learning Algorithms,

    A. Nichol, J. Achiam, and J. Schulman, “On First -Order Meta - Learning Algorithms,” arXiv:1803.02999, pp. 1–15, 2018

  209. [217]

    Introducing Symmetries to Black Box Meta Reinforcement Learning,

    L. Kirsch, S. Flennerhag, H. Van Hasselt, A. Friesen, J. Oh, and Y. Chen, “Introducing Symmetries to Black Box Meta Reinforcement Learning,” arXiv:2109.10781, pp. 1–12, 2022

  210. [218]

    Learning to Learn: Meta -Critic Networks for Sample Efficient Learning,

    F. Sung, L. Zhang, T. Xiang, T. Hospedales, and Y. Yang, “Learning to Learn: Meta -Critic Networks for Sample Efficient Learning,” arXiv:1706.09529, pp. 1–12, 2017

  211. [219]

    Hypernetworks in Meta-Reinforcement Learning,

    J. Beck, R. Vuorio, M. Jackson, and S. Whiteson, “Hypernetworks in Meta-Reinforcement Learning,” arXiv:2210.11348, pp. 1–14, 2022

  212. [220]

    Meta Reinforcement Learning As Task Inference,

    J. Humplik, A. Galashov, L. Hasenclever, and N. Heess, “Meta Reinforcement Learning As Task Inference,” arXiv:1905.06424, pp. 1–22, 2019

  213. [221]

    Decoupling Exploration and Exploitation for Meta -Reinforcement Learning without Sacrifices,

    E. Z. Liu, A. Raghunathan, P. Liang, and C. Finn, “Decoupling Exploration and Exploitation for Meta -Reinforcement Learning without Sacrifices,” Proc. Mach. Learn. Res. , vol. 139, pp. 6925 – 6935, 2021

  214. [222]

    Fast adaptation to new environments via policy -dynamics value functions,

    R. Raileanu, M. Goldstein, A. Szlam, and R. Fergus, “Fast adaptation to new environments via policy -dynamics value functions,” 37th Int. Conf. Mach. Learn. ICML 2020, vol. 119, pp. 7876–7887, 2020

  215. [223]

    Data -Efficient Task Generalization via Probabilistic Model -based Meta Reinforcement Learning,

    A. Bhardwaj et al. , “Data -Efficient Task Generalization via Probabilistic Model -based Meta Reinforcement Learning,” IEEE Robot. Autom. Lett., vol. 9, no. 4, pp. 3918–3925, 2024

  216. [224]

    VariBAD: a Very Good Method for Bayes - Adaptive Deep RL Via Meta -Learning,

    L. Zintgraf et al. , “VariBAD: a Very Good Method for Bayes - Adaptive Deep RL Via Meta -Learning,” 8th Int. Conf. Learn. Represent. ICLR 2020, pp. 1–20, 2020

  217. [225]

    Environment probing interaction policies,

    W. Zhou, L. Pinto, and A. Gupta, “Environment probing interaction policies,” arXiv:1907.11740, pp. 1–13, 2019

  218. [226]

    MAME: Model -Agnostic Meta-Exploration,

    S. Gurumurthy, S. Kumar, and K. Sycara, “MAME: Model -Agnostic Meta-Exploration,” arXiv:1911.04024, pp. 1–13, 2019

  219. [227]

    Exploration in Approximate Hyper-State Space for Meta Reinforcement Learning,

    L. Zintgraf et al., “Exploration in Approximate Hyper-State Space for Meta Reinforcement Learning,” Proc. 38th Int. Conf. Mach. Learn. , vol. 139, pp. 12991–13001, 2021

  220. [228]

    Hindsight Foresight Relabeling for Meta -Reinforcement Learning,

    M. Wan, J. Peng, and T. Gangwani, “Hindsight Foresight Relabeling for Meta -Reinforcement Learning,” ICLR 2022 - 10th Int. Conf. Learn. Represent., pp. 1–18, 2022

  221. [229]

    Model -based adversarial meta-reinforcement learning,

    Z. Lin, G. Thomas, G. Yang, and T. Ma, “Model -based adversarial meta-reinforcement learning,” arXiv:2006.08875, pp. 1–19, 2021

  222. [230]

    Where do rewards come from?,

    R. L. Lewis, S. Singh, and A. G. Barto, “Where do rewards come from?,” Proc. Int. Symp. AI-Inspired Biol. Jackie Chappell, Susannah Thorpe, Nick Hawes Aaron Sloman (Eds.),at AISB 2010 Conv. 29 March – 1 April 2010, Montfort Univ. Leicester, UK , pp. 111 –116, 2010

  223. [231]

    A survey on intrinsic motivation in reinforcement learning,

    A. Aubret, L. Matignon, and S. Hassas, “A survey on intrinsic motivation in reinforcement learning,” arXiv:1908.06976, pp. 1 –39, 22 Zhou, C. et al. Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review 2019

  224. [232]

    Learning Task - Distribution Reward Shaping with Meta -Learning,

    H. Zou, T. Ren, D. Yan, H. Su, and J. Zhu, “Learning Task - Distribution Reward Shaping with Meta -Learning,” 35th AAAI Conf. Artif. Intell. AAAI 2021, vol. 12B, pp. 11210–11218, 2021

  225. [233]

    On learning intrinsic rewards for policy gradient methods,

    Z. Zheng, J. Oh, and S. Singh, “On learning intrinsic rewards for policy gradient methods,” arXiv:1804.06459, pp. 1–15, 2018

  226. [234]

    How should an agent practice?,

    J. Rajendran, R. Lewis, V. Veeriah, H. Lee, and S. Singh, “How should an agent practice?,” AAAI 2020 - 34th AAAI Conf. Artif. Intell., pp. 5454–5461, 2020

  227. [235]

    Off -Policy Meta - Reinforcement Learning with Belief -Based Task Inference,

    T. Imagawa, T. Hiraoka, and Y. Tsuruoka, “Off -Policy Meta - Reinforcement Learning with Belief -Based Task Inference,” IEEE Access, vol. 10, pp. 49494–49507, 2022

  228. [236]

    Adaptive auxiliary task weighting for reinforcement learning,

    X. Lin, H. S. Baweja, G. Kantor, and D. Held, “Adaptive auxiliary task weighting for reinforcement learning,” 33rd Conf. Neural Inf. Process. Syst. (NeurIPS 2019), Vancouver, Canada., 2019

  229. [237]

    Between MDPs and semi - MDPs: A framework for temporal abstraction in reinforcement learning,

    R. S. Sutton, D. Precup, and S. Singh, “Between MDPs and semi - MDPs: A framework for temporal abstraction in reinforcement learning,” Artif. Intell., vol. 112, pp. 181–211, 1999

  230. [238]

    Meta Learning Shared Hierarchies,

    J. Schulman, J. Ho, X. Chen, and P. Abbeel, “Meta Learning Shared Hierarchies,” arXiv:1710.09767, pp. 1–11, 2017

  231. [239]

    Discovery of Options via Meta-Learned Subgoals,

    V. Veeriah et al., “Discovery of Options via Meta-Learned Subgoals,” arXiv:2102.06741, pp. 1–19, 2021

  232. [240]

    Meta Reinforcement Learning for Fast Adaptation of Hierarchical Policies,

    A. Author, “Meta Reinforcement Learning for Fast Adaptation of Hierarchical Policies,” Prepr. Submitt. to 35th Conf. Neural Inf. Process. Syst. (NeurIPS 2021), pp. 1–14, 2021

  233. [241]

    Meta-gradient reinforcement learning with an objective discovered online,

    Z. Xu, H. van Hasselt, M. Hessel, J. Oh, S. Singh, and D. Silver, “Meta-gradient reinforcement learning with an objective discovered online,” arXiv:2007.08433, pp. 1–18, 2020

  234. [242]

    Improving Generalization in Meta Reinforcement Learning using Learned Objectives,

    L. Kirsch, S. van Steenkiste, and J. Schmidhuber, “Improving Generalization in Meta Reinforcement Learning using Learned Objectives,” arXiv:1910.04098, pp. 1–21, 2020

  235. [243]

    Discovered Policy Optimisation,

    C. Lu, J. G. Kuba, A. Letcher, L. Metz, C. S. de Witt, and J. Foerster, “Discovered Policy Optimisation,” arXiv:2210.05639, pp. 1–18, 2022

  236. [244]

    Bayesian Controller Fusion : Leveraging Control Priors In Deep Reinforcement Learning for Robotics,

    K. Rana, V. Dasagi, J. Haviland, B. Talbot, M. Milford, and S. Niko, “Bayesian Controller Fusion : Leveraging Control Priors In Deep Reinforcement Learning for Robotics,” arXiv:2107.09822, pp. 1 –19, 2021

  237. [245]

    Bayesian Curiosity for Efficient Exploration in Reinforcement Learning,

    T. Blau, L. Ott, F. Ramos, and L. G. Nov, “Bayesian Curiosity for Efficient Exploration in Reinforcement Learning,” arXiv:1911.08701, pp. 1–7, 2019

  238. [246]

    Mixed Reinforcement Learning for Efficient Policy Optimization in Stochastic Environments,

    Y. Mu et al. , “Mixed Reinforcement Learning for Efficient Policy Optimization in Stochastic Environments,” 2020 20th Int. Conf. Control. Autom. Syst. (ICCAS), Busan, Korea, 2020

  239. [247]

    PAC -Bayesian Soft Actor-Critic Learning,

    B. Tasdighi, K. K. Brink, and M. Kandemir, “PAC -Bayesian Soft Actor-Critic Learning,” arXiv:2301.12776, pp. 1–12, 2023

  240. [248]

    Probably Approximately Correct Learning,

    D. Haussler, “Probably Approximately Correct Learning,” Natl. Conf. Artif. Intell. (AAAI-1990), Boston, Massachusetts., 1990

  241. [249]

    Probabilistic Model Checking of Robots Deployed in Extreme Environments,

    X. Zhao, V. Robu, D. Flynn, F. Dinmohammadi, M. Fisher, and M. Webster, “Probabilistic Model Checking of Robots Deployed in Extreme Environments,” arXiv:1812.04128, pp. 1–9, 2018

  242. [250]

    Uncertainty Quantification with Statistical Guarantees in End -to-End Autonomous Driving Control,

    R. Michelmore, M. Wicker, L. Laurenti, L. Cardelli, Y. Gal, and M. Kwiatkowska, “Uncertainty Quantification with Statistical Guarantees in End -to-End Autonomous Driving Control,” in 2020 IEEE International Conference on Robotics and Automation (ICRA) , 2020, pp. 7344–7350

  243. [251]

    A Bayesian Deep Neural Network for Safe Visual Servoing in Human -Robot Interaction,

    L. Shi, C. Copot, and S. Vanlanduit, “A Bayesian Deep Neural Network for Safe Visual Servoing in Human -Robot Interaction,” Front. Robot. AI, vol. 8, pp. 1–13, 2021

  244. [252]

    Infinite Time Horizon Safety of Bayesian Neural Networks,

    M. Lechner, “Infinite Time Horizon Safety of Bayesian Neural Networks,” arXiv:2111.03165, pp. 1–15, 2021

  245. [253]

    Risk Averse Bayesian Reward Learning for Autonomous Navigation from Human Demonstration,

    C. Ellis, M. Wigness, J. Rogers, C. Lennon, and L. Fiondella, “Risk Averse Bayesian Reward Learning for Autonomous Navigation from Human Demonstration,” 2021 IEEE/RSJ Int. Conf. Intell. Robot. Syst. (IROS), Sept. 2021 - 01 Oct. 2021,Prague, Czech Repub., 2021

  246. [254]

    Modeling and Predicting Trust Dynamics in Human – Robot Teaming : A Bayesian Inference Approach,

    Y. Guo, X. J. Yang, and X. J. Yang, “Modeling and Predicting Trust Dynamics in Human – Robot Teaming : A Bayesian Inference Approach,” Int. J. Soc. Robot., vol. 13, no. 8, pp. 1899–1909, 2021

  247. [255]

    Neural Dynamic Policies for End-to-End Sensorimotor Learning,

    S. Bahl, M. Mukadam, A. Gupta, and D. Pathak, “Neural Dynamic Policies for End-to-End Sensorimotor Learning,” arXiv:2012.02788v1, pp. 1–16, 2020

  248. [256]

    Offline Contextual Bayesian Optimization,

    I. Char et al., “Offline Contextual Bayesian Optimization,” 33rd Conf. Neural Inf. Process. Syst. (NeurIPS-2019), Vancouver, Canada, 2019

  249. [257]

    Bayesian optimization with safety constraints : safe and automatic parameter tuning in robotics,

    F. Berkenkamp, A. Krause, and A. P. Schoellig, “Bayesian optimization with safety constraints : safe and automatic parameter tuning in robotics,” Mach. Learn., 2021

  250. [258]

    GoSafe : Globally Optimal Safe Robot Learning,

    D. Baumann, A. Marco, M. Turchetta, S. Trimpe, and R. O. May, “GoSafe : Globally Optimal Safe Robot Learning,” 2021 IEEE Int. Conf. Robot. Autom. (ICRA),30 May 2021 - 05 June 2021,Xi’an, China, 2021

  251. [259]

    Bayesian Safe Learning and Control with Sum -of-Squares Analysis and Polynomial Kernels,

    A. Devonport, H. Yin, and M. Arcak, “Bayesian Safe Learning and Control with Sum -of-Squares Analysis and Polynomial Kernels,” 2020 59th IEEE Conf. Decis. Control (CDC), 14 -18 December 2020, Jeju, Korea, 2020

  252. [260]

    Performance and safety of Bayesian model predictive control: Scalable model -based RL with guarantees,

    K. P. Wabersich, “Performance and safety of Bayesian model predictive control: Scalable model -based RL with guarantees,” arXiv:2006.03483v1, pp. 1–17, 2020

  253. [261]

    Decision making of autonomous vehicles in lane change scenarios : Deep reinforcement learning approaches with risk awareness,

    G. Li, Y. Yang, S. Li, X. Qu, N. Lyu, and S. Eben, “Decision making of autonomous vehicles in lane change scenarios : Deep reinforcement learning approaches with risk awareness,” Transp. Res. Part C , vol. 134, p. 103452, 2022

  254. [262]

    CONSTRAINED POLICY OPTIMIZATION VIA BAYESIAN WORLD MODELS,

    Y. As, “CONSTRAINED POLICY OPTIMIZATION VIA BAYESIAN WORLD MODELS,” arXiv:2201.09802, pp. 1–24, 2022

  255. [263]

    Risk -Averse Bayes-Adaptive Reinforcement Learning,

    M. Rigter, B. Lacerda, and N. Hawes, “Risk -Averse Bayes-Adaptive Reinforcement Learning,” arXiv:2102.05762, pp. 1–13, 2021

  256. [264]

    Bayesian Learning -Based Adaptive Control for Safety Critical Systems,

    D. D. Fan, J. Nguyen, R. Thakker, N. Alatur, and E. A. Theodorou, “Bayesian Learning -Based Adaptive Control for Safety Critical Systems,” 2020 IEEE Int. Conf. Robot. Autom. (ICRA),Paris, Fr. 31 May 2020 - 31 August 2020, 2020

  257. [265]

    L1-GP: L1 Adaptive Control with Bayesian Learning,

    A. Gahlawat, “L1-GP: L1 Adaptive Control with Bayesian Learning,” Proc. 2nd Conf. Learn. Dyn. Control. PMLR , vol. 120, pp. 826 –837, 2020

  258. [266]

    Data -Efficient Domain Randomization With Bayesian Optimization,

    F. Muratore, C. Eilers, M. Gienger, and J. Peters, “Data -Efficient Domain Randomization With Bayesian Optimization,” arXiv:2003.02471, pp. 1–8, 2020

  259. [267]

    A User’s Guide to Calibrating Robotics Simulators,

    B. Mehta, D. Fox, and F. Ramos, “A User’s Guide to Calibrating Robotics Simulators,” Proc. 2020 Conf. Robot Learn. PMLR, vol. 155, pp. 1326–1340, 2021

  260. [268]

    Interpretability and Explainability: A Machine Learning Zoo Mini-tour,

    J. E. Vogt, “Interpretability and Explainability: A Machine Learning Zoo Mini-tour,” arXiv:2012.01805, pp. 1–24, 2023

  261. [269]

    Bayes -Adaptive POMDPs,

    J. Pineau, “Bayes -Adaptive POMDPs,” Neural Inf. Process. Syst. (NIPS-2007),Vancouver, Br. Columbia, Canada, 2007

  262. [270]

    Rethinking the implementation tricks and monotonicity constraint in cooperative multi -agent reinforcement learning,

    S. A. Harding, “Rethinking the implementation tricks and monotonicity constraint in cooperative multi -agent reinforcement learning,” arXiv:2102.03479, pp. 1–20, 2021

  263. [271]

    MAMBPO: Sample-efficient multi -robot reinforcement learning using learned world models,

    D. Willemsen, M. Coppola, and G. C. H. E. de Croon, “MAMBPO: Sample-efficient multi -robot reinforcement learning using learned world models,” arXiv:2103.03662, pp. 1–6, 2021

  264. [272]

    Monte -Carlo Planning in Large POMDPs,

    D. Silver and J. Veness, “Monte -Carlo Planning in Large POMDPs,” Neural Inf. Process. Syst. Br. Columbia, Canada, 2010

  265. [273]

    Scalable Planning and Learning for Multiagent POMDPs,

    C. Amato and F. A. Oliehoek, “Scalable Planning and Learning for Multiagent POMDPs,” Adv. Artif. Intell. (AAAI -2015),Austin, Texas, USA, 2015

  266. [274]

    Decentralized Patrolling Under Constraints in Dynamic Environments,

    S. Chen, F. Wu, L. Shen, J. Chen, and S. D. Ramchurn, “Decentralized Patrolling Under Constraints in Dynamic Environments,” IEEE Trans. Cybern., vol. 46, no. 12, pp. 3364–3376, 2016

  267. [275]

    Bayesian Reinforcement Learning for Multi -Robot Decentralized Patrolling in Uncertain Environments,

    X. Zhou, W. Wang, T. Wang, Y. Lei, and F. Zhong, “Bayesian Reinforcement Learning for Multi -Robot Decentralized Patrolling in Uncertain Environments,” IEEE Trans. Veh. Technol., vol. 68, no. 12, pp. 11691–11703, 2019

  268. [276]

    Multi -Agent Reinforcement Learning with Multi-Step Generative Models,

    O. Krupnik, I. Mordatch, and A. Tamar, “Multi -Agent Reinforcement Learning with Multi-Step Generative Models,” arXiv:1901.10251, pp. 1–15, 2019

  269. [277]

    Tesseract: Tensorised Actors for Multi -Agent Reinforcement Learning,

    A. Mahajan, M. Samvelyan, L. Mao, V. Makoviychuk, A. Garg, and J. Kossaifi, “Tesseract: Tensorised Actors for Multi -Agent Reinforcement Learning,” arXiv:2106.00136, pp. 1–21, 2021

  270. [278]

    Model based Multi -agent Reinforcement Learning with Tensor Decompositions,

    P. Van Der Vaart and A. Mahajan, “Model based Multi -agent Reinforcement Learning with Tensor Decompositions,” arXiv:2110.14524, pp. 1–12, 2021

  271. [279]

    Mean Field Multi -Agent Reinforcement Learning,

    Y. Yang, R. Luo, M. Li, M. Zhou, W. Zhang, and J. Wang, “Mean Field Multi -Agent Reinforcement Learning,” arXiv:1802.05438v5, 2018

  272. [280]

    Efficient Model -Based Multi -Agent Mean-Field Reinforcement Learning,

    M. L. May and A. Krause, “Efficient Model -Based Multi -Agent Mean-Field Reinforcement Learning,” arXiv:2107.04050, pp. 1 –35, 2023

  273. [281]

    Model -based Multi - agent Policy Optimization with Adaptive Opponent -wise Rollouts,

    W. Zhang, X. Wang, J. Shen, and M. Zhou, “Model -based Multi - agent Policy Optimization with Adaptive Opponent -wise Rollouts,” arXiv:2105.03363, pp. 1–26, 2022

  274. [282]

    Multi - Agent Actor -Critic for Mixed Cooperative -Competitive Environments,

    R. Lowe, J. Harb, A. Tamar, P. Abbeel, and I. Mordatch, “Multi - Agent Actor -Critic for Mixed Cooperative -Competitive Environments,” arXiv:1706.02275v4, 2020

  275. [283]

    Shared Experience Actor-Critic for Multi -Agent Reinforcement Learning,

    F. Christianos, L. Schäfer, and S. V Albrecht, “Shared Experience Actor-Critic for Multi -Agent Reinforcement Learning,” arXiv:2006.07169, 2020

  276. [284]

    Model-based Reinforcement Learning for Decentralized Multiagent Rendezvous,

    R. E. Wang, T. E. Lee, J. C. Kew, D. Lee, B. Ichter, and T. Zhang, “Model-based Reinforcement Learning for Decentralized Multiagent Rendezvous,” arXiv:2003.06906, pp. 1–15, 2020

  277. [285]

    Decentralized Planning under Uncertainty for Teams of Communicating Agents,

    M. T. J. Spaan, G. J. Gordon, and N. Vlassis, “Decentralized Planning under Uncertainty for Teams of Communicating Agents,” Int. Jt. Conf. Auton. Agents Multiagent Syst. (AAMAS-2006), May 8–12, Hakodate, Hokkaido, Japan, 2006

  278. [286]

    Learning to communicate through imagination with model- based deep multi-agent reinforcement learning,

    A. Pretorius et al. , “Learning to communicate through imagination with model- based deep multi-agent reinforcement learning,” 2020

  279. [287]

    W. Kim, J. Park, and Y. Sung, “Communication in multi -agent 23 Zhou, C. et al. Combining Bayesian Inference and Reinforcement Learning for Agent Decision Making: A Review reinforcement learning: Intention sharing,” Ninth Int. Conf. Learn. Represent. Mon May 3rd through Fri 7t...

  280. [288]

    Scaling Multi -Agent Reinforcement Learning with Selective Parameter Sharing,

    F. Christianos, G. Papoudakis, A. Rahman, and S. V Albrecht, “Scaling Multi -Agent Reinforcement Learning with Selective Parameter Sharing,” arXiv:2102.07475, pp. 1–10, 2021

  281. [289]

    Distral : Robust Multitask Reinforcement Learning,

    Y. W. Teh et al., “Distral : Robust Multitask Reinforcement Learning,” arXiv:1707.04175, pp. 1–13, 2017

  282. [290]

    IMPALA: Scalable Distributed Deep -RL with Importance Weighted Actor -Learner Architectures,

    A. Architectures et al. , “IMPALA: Scalable Distributed Deep -RL with Importance Weighted Actor -Learner Architectures,” arXiv:1802.01561, pp. 1–22, 2018

  283. [291]

    Asynchronous Methods for Deep Reinforcement Learning,

    V. Mnih et al. , “Asynchronous Methods for Deep Reinforcement Learning,” in Proceedings of Machine Learning Research , 2016, vol. 48, pp. 1928–1937

  284. [292]

    Multi-task Deep Reinforcement Learning with PopArt,

    M. Hessel, H. Soyer, L. Espeholt, W. Czarnecki, S. Schmitt, and H. van Hasselt, “Multi-task Deep Reinforcement Learning with PopArt,” arXiv:1809.04474, pp. 1–12, 2018

  285. [293]

    Non -Gaussian Risk Bounded Trajectory Optimization for Stochastic Nonlinear Systems in Uncertain Environments,

    W. Han, A. Jasour, and B. Williams, “Non -Gaussian Risk Bounded Trajectory Optimization for Stochastic Nonlinear Systems in Uncertain Environments,” Proc. - IEEE Int. Conf. Robot. Autom. May 23-27, 2022. Philadelphia, PA, USA, pp. 11044–11050, 2022

  286. [294]

    Efficient Probabilistic Collision Detection for Non-Gaussian Noise Distributions,

    J. S. Park and D. Manocha, “Efficient Probabilistic Collision Detection for Non-Gaussian Noise Distributions,” IEEE Robot. Autom. Lett., vol. 5, no. 2, pp. 1024–1031, 2020

  287. [295]

    Non -Gaussian Chance - Constrained Trajectory Planning for Autonomous Vehicles under Agent Uncertainty,

    A. Wang, A. Jasour, and B. C. Williams, “Non -Gaussian Chance - Constrained Trajectory Planning for Autonomous Vehicles under Agent Uncertainty,” IEEE Robot. Autom. Lett. , vol. 5, no. 4, pp. 6041–6048, 2020

  288. [296]

    Adaptive Message Passing For Cooperative Positioning Under Unknown Non -Gaussian Noises,

    J. Xiong, X. peng Xie, Z. Xiong, Y. Zhuang, Y. Zheng, and C. Wang, “Adaptive Message Passing For Cooperative Positioning Under Unknown Non -Gaussian Noises,” IEEE Trans. Instrum. Meas. , vol. 73, pp. 1–14, 2024

  289. [297]

    Tightly coupled distributed Kalman filter under non -Gaussian noises,

    Y. Fu, M. Sun, and Y. Gao, “Tightly coupled distributed Kalman filter under non -Gaussian noises,” Signal Processing , vol. 200, p. 108678, 2022

  290. [298]

    Recent Advances in Non -Gaussian Stochastic Systems Control Theory and Its Applications,

    Q. Zhang and Y. Zhou, “Recent Advances in Non -Gaussian Stochastic Systems Control Theory and Its Applications,” Int. J. Netw. Dyn. Intell., vol. 1, no. 1, pp. 111–119, 2022

  291. [299]

    Hierarchical Reinforcement Learning with Model - Based Planning for Finding Sparse Rewards,

    T. D. Bartley, “Hierarchical Reinforcement Learning with Model - Based Planning for Finding Sparse Rewards,” UC Irvine Electron. Theses Diss., 2023

  292. [2020]

    Notes Comput

    Lect. Notes Comput. Sci. vol 12533. Springer, Cham, 2020

Pith tools

Reviewed August 15, 2026 · model on record in the stance chip above.