REVIEW 4 minor 300 references
Effective Reward Specification in Deep Reinforcement Learning
T0 review · 0 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read Reward specification is the bottleneck in deep RL, and no single tool fixes it.
desk verdict A competent thesis compilation of four solid papers; valuable as a map of reward specification, not as a new research contribution. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
Four mechanisms carry the argument. Adversarial Soft Advantage Fitting uses a structured discriminator of the form $D(\tau) = \tilde{p}(\tau)/(\tilde{p}(\tau)+p_G(\tau))$, built from two evaluable policies, so that optimizing the discriminator simultaneously solves the generator's problem and yields the expert trajectory distribution without a reinforcement-learning loop. TeamReg adds team-spirit losses—each agent predicts its teammate's action and is regularized to be predictable—while CoachReg introduces a central coach that outputs a policy mask, with agents regularized to match the mask, producing synchronized sub-policy switches. The constrained-RL framework uses indicator cost functions with normalized Lagrange multipliers and a bootstrap constraint that keeps the multiplier scale from dominating the main objective. Goal-conditioned GFlowNets extend the flow-matching objective (trajectory balance) to condition on objective-space subregions, with a learned goal sampler that broadens coverage of the Pareto front.
What would settle it
A concrete test of the load-bearing assumption: train ASAF-1 on a task where the learner's initial policy and the expert visit largely disjoint state regions, and measure the state-occupancy divergence between them during training; if the method still recovers expert-level performance despite the divergence remaining large early on, the occupancy assumption is not the load-bearing part of the argument. A separate thesis-level test: if a single reward-specification method matched all four specialized tools without modification, the paper's central claim of no universal solution would fail.
Extended reading notes
Core claim
On the paper's own terms, the central discovery is that effective reward specification is not a single problem but a family of problems, and that each family has a characteristic failure mode that a purpose-built mechanism can address. In inverse reinforcement learning, the thesis shows that a discriminator conditioned on two policies—the previous generator and a learnable policy—can directly recover the expert policy, so the usual inner RL loop can be dropped entirely. In multi-agent learning, it shows that auxiliary objectives enforcing inter-agent predictability (TeamReg) and synchronized sub-policy selection (CoachReg) act as inductive biases that help agents discover coordinated strategies under sparse rewards. In constrained RL, it proposes indicator cost functions, multiplier normalization, and a bootstrap constraint so that hard behavioral requirements can be specified directly rather than through reward shaping. In molecular design, it shows that goal-conditioned GFlowNets, trained with a learned goal distribution over objective subregions, can generate molecules along the entire Pareto front. The thesis concludes that these results collectively support the absence of a universal solution to reward specification.
Load-bearing premise
The practical imitation results depend on a windowed approximation that assumes the learner and the expert visit the same states with the same frequencies, which is false until imitation has actually succeeded.
Editorial extensions
If this is right
- If ASAF is right, adversarial imitation learning can be implemented and trained at roughly half the complexity, because the discriminator update itself produces the new policy; this removes the unstable alternation between RL and reward fitting.
- If TeamReg and CoachReg are right, coordination-promoting regularizers can replace task-specific reward shaping and curriculum design in sparse-reward cooperative tasks, and can even improve hyperparameter robustness in multi-agent training.
- If the constrained-RL framework is right, designers can monitor whether an agent satisfies hard behavioral requirements directly, since constraints are stated as costs rather than buried in a weighted reward.
- If goal-conditioned GFlowNets are right, one trained model can be steered to different regions of the objective space at deployment time, making multi-objective molecular design a matter of choosing a goal rather than retraining.
- Taken together, the thesis's four results imply that reward specification should be treated as a toolbox selection problem: demonstrations, auxiliary objectives, constraints, and goal conditioning each fit different applications.
Reading between the lines
- Beyond the thesis's claims, the ASAF windowed approximation suggests a practical diagnostic: tracking the divergence between learner and expert state-occupancy during training could tell a practitioner when the imitation signal is trustworthy and when it is fitting to the wrong distribution.
- The thesis treats its four contributions separately, but they compose naturally: an ASAF-style reward model could seed the constrained-RL framework, and the multi-agent coordination regularizers could supply the behavioral constraints that the framework monitors.
- The goal-conditioned GFlowNet approach implies a testable extension to other generative design problems—materials, circuits, or biological sequences—wherever a designer cares about covering trade-offs rather than a single optimum.
- The thesis's 'no universal solution' claim, if correct, predicts that benchmark comparisons of reward-specification methods will keep showing environment-dependent winners; a meta-analysis of such comparisons would be a direct test.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The manuscript is a PhD thesis that frames reward specification as a fundamental obstacle in deep reinforcement learning. It provides a technical background on deep learning and RL, a literature review organized around reward composition and reward modeling, and four original contribution chapters reproducing peer-reviewed published articles: ASAF (adversarial imitation learning without policy optimization), TeamReg/CoachReg (policy regularization for multi-agent coordination), constrained RL for direct behavior specification, and goal-conditioned GFlowNets for multi-objective molecular design. The thesis concludes that there is no universal reward specification solution and that practitioners should select tools according to the requirements of each application.
Significance. The four contribution chapters are already peer-reviewed, and the thesis adds value by placing them in a common taxonomy and by discussing their strengths and limitations. The ASAF formulation is a genuine simplification of adversarial imitation learning, with a theoretical statement for full trajectories and empirical support on several benchmarks. CoachReg provides a useful inductive bias for sparse-reward multi-agent coordination, and the constrained RL and goal-conditioned GFlowNet contributions offer practical tools for behavior specification and controllable multi-objective generation. The thesis is honest about its limitations, including the ASAF-1 windowed approximation (Section 5.3.3) and the negative TeamReg result on COMPROMISE (Table 6.1), which strengthens the credibility of the concluding 'no universal solution' claim. That claim is a qualitative synthesis rather than a formal theorem; within that scope, the evidence is adequate.
minor comments (4)
- [Abstract and Section 1.2] The term 'alignment' is used as a central success criterion in the abstract and introduction, but it is never defined operationally; a sentence specifying what counts as alignment in each contribution (e.g., constraint satisfaction rate, behavioral predictability, or Pareto coverage) would make the thesis's central claim easier to evaluate.
- [Section 5.3.3] The ASAF-1 approximation is described in a paragraph inside the algorithm box; because this is the main caveat of the practical variant, I suggest moving it to a dedicated 'Limitations' subsection and stating explicitly that Theorem 1 applies to full trajectories only.
- [Table 6.1 and Section 6.7.1] The negative result of TeamReg on COMPROMISE is visible in the table but not annotated; a footnote explaining that this failure mode is associated with an adversarial component, and stating whether significance testing was performed, would help readers weigh the cross-task claims.
- [Chapter 9] The heading 'Sucesses and Limitations' contains a typo and should read 'Successes and Limitations'.
Circularity Check
No significant circularity; the thesis is a compilation of independent, peer-reviewed contributions with disclosed approximations.
full rationale
I walked the derivation chains of the four contributions. Article 1 (ASAF) derives the structured-discriminator optimum from the standard logistic/GAN objective; the result that the optimal discriminator parameter equals the expert distribution is a mathematical identity of that objective, not a case of fitting a parameter and then predicting it back. The practical windowed variant ASAF-1 is explicitly acknowledged in Section 5.3.3 to assume equal state-occupancy measures until the expert policy is recovered; that is a disclosed limitation rather than a concealed input-to-output reduction. Articles 2-4 are empirical and algorithmic contributions: TeamReg/CoachReg add auxiliary objectives to MADDPG and are validated against baselines on independent environments; the constrained-RL framework and goal-conditioned GFlowNet approach are method proposals evaluated with external benchmarks. No load-bearing step is justified solely by a self-citation: the thesis reproduces the author's own published articles, but their results are experiment- and code-based, and the thesis does not invoke any self-authored uniqueness theorem to forbid alternatives. The meta-claim that no universal reward-specification solution exists is a survey-style conclusion, not a formal derivation from an assumption that already contains the conclusion. The self-citation present is structural (a doctoral thesis compiled from the author's papers), not load-bearing, and therefore does not constitute circularity.
Assumptions & free parameters
free parameters (4)
- ASAF window size w =
32 to 200 (tuned per environment)
- TeamReg coefficients lambda1, lambda2 =
tuned per environment (see Appendix B.5)
- CoachReg coefficients lambda1, lambda2, lambda3 and mask count K =
tuned per environment
- Goal-conditioned GFlowNet hyperparameter m_g =
set to control reward profile (Appendix D)
assumptions (4)
- domain assumption Markov property and stationary policy assumption
- domain assumption The reward hypothesis: any goal can be captured by a scalar reward function
- domain assumption Expert demonstrations are drawn from an optimal or near-optimal policy
- standard math Neural networks can approximate the required functions
invented entities (3)
-
Coach entity
independent evidence
-
Policy masks
independent evidence
-
Tabular Goal-Sampler
independent evidence
Cite this review
Pith. "Pith review of Effective Reward Specification in Deep Reinforcement Learning." pith.science (2026). https://pith.science/paper/AP45VC5I
@misc{pith2026241207177,
author = {Pith},
title = {Pith review of: Effective Reward Specification in Deep Reinforcement Learning},
year = {2026},
howpublished = {\url{https://pith.science/paper/AP45VC5I}},
note = {Machine review of arXiv:2412.07177}
}
read the original abstract
In the last decade, Deep Reinforcement Learning has evolved into a powerful tool for complex sequential decision-making problems. It combines deep learning's proficiency in processing rich input signals with reinforcement learning's adaptability across diverse control tasks. At its core, an RL agent seeks to maximize its cumulative reward, enabling AI algorithms to uncover novel solutions previously unknown to experts. However, this focus on reward maximization also introduces a significant difficulty: improper reward specification can result in unexpected, misaligned agent behavior and inefficient learning. The complexity of accurately specifying the reward function is further amplified by the sequential nature of the task, the sparsity of learning signals, and the multifaceted aspects of the desired behavior. In this thesis, we survey the literature on effective reward specification strategies, identify core challenges relating to each of these approaches, and propose original contributions addressing the issue of sample efficiency and alignment in deep reinforcement learning. Reward specification represents one of the most challenging aspects of applying reinforcement learning in real-world domains. Our work underscores the absence of a universal solution to this complex and nuanced challenge; solving it requires selecting the most appropriate tools for the specific requirements of each unique application.
Figures
Figures from the paper (21 more)
Reference graph
Works this paper leans on
-
[1]
write newline
" write newline "" initialize.prev.this.status FUNCTION begin.bib " write newline preamble empty 'skip preamble write newline if " thebibliography " longest.label * " " * write newline " [1] #1 " write newline " url@samestyle " write newline " " write newline " [2] #2 " write newline " =0pt " write newline " " ALTinterwordstretchfactor * " " * write newli...
-
[2]
Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al. (2016). \ TensorFlow \ : a system for \ Large-Scale \ machine learning. In 12th USENIX symposium on operating systems design and implementation (OSDI 16) , pages 265--283
2016
-
[3]
Abbeel, P., Coates, A., and Ng, A. Y. (2010). Autonomous helicopter aerobatics through apprenticeship learning. The International Journal of Robotics Research , 29(13):1608--1639
2010
-
[4]
Abbeel, P., Coates, A., Quigley, M., and Ng, A. (2006). An application of reinforcement learning to aerobatic helicopter flight. Advances in neural information processing systems , 19
2006
-
[5]
and Ng, A
Abbeel, P. and Ng, A. Y. (2004). Apprenticeship learning via inverse reinforcement learning. In Proceedings of the twenty-first international conference on Machine learning , page 1
2004
-
[6]
and Ng, A
Abbeel, P. and Ng, A. Y. (2005). Exploration and apprenticeship learning in reinforcement learning. In Proceedings of the 22nd international conference on Machine learning , pages 1--8
2005
-
[7]
K., Littman, M., Precup, D., and Singh, S
Abel, D., Dabney, W., Harutyunyan, A., Ho, M. K., Littman, M., Precup, D., and Singh, S. (2021). On the expressivity of markov reward. Advances in Neural Information Processing Systems , 34:7799--7812
2021
-
[8]
Abels, A., Roijers, D., Lenaerts, T., Now \'e , A., and Steckelmacher, D. (2019). Dynamic weights in multi-objective deep reinforcement learning. In International conference on machine learning , pages 11--20. PMLR
2019
Show all 300 references
-
[9]
Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017). Constrained policy optimization. In International conference on machine learning , pages 22--31. PMLR
2017
-
[10]
Adams, S., Cody, T., and Beling, P. A. (2022). A survey of inverse reinforcement learning. Artificial Intelligence Review , 55(6):4307--4346
2022
-
[11]
and Dayan, P
Ahilan, S. and Dayan, P. (2019). Feudal multi-agent hierarchies for cooperative reinforcement learning. arXiv preprint arXiv:1901.08492
2019 arXiv
-
[12]
Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al. (2022). Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691
2022 arXiv
-
[13]
Aissani, N., Beldjilali, B., and Trentesaux, D. (2008). Efficient and effective reactive scheduling of manufacturing system using sarsa-multi-objective agents. In MOSIM’08: 7th Conference Internationale de Modelisation et Simulation , pages 698--707
2008
-
[14]
Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al. (2019). Solving rubik's cube with a robot hand. arXiv preprint arXiv:1910.07113
2019 arXiv
-
[15]
Akrour, R., Schoenauer, M., and Sebag, M. (2011). Preference-based policy learning. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2011, Athens, Greece, September 5-9, 2011. Proceedings, Part I 11 , pages 12--27. Springer
2011
-
[16]
Alonso, E., Peter, M., Goumard, D., and Romoff, J. (2020). Deep reinforcement learning for navigation in aaa video games. arXiv preprint arXiv:2011.04764
2020 arXiv
-
[17]
Altman, E. (1999). Constrained Markov Decision Processes . Stochastic Modeling Series. Taylor & Francis
1999
-
[18]
Amin, S., Gomrokchi, M., Satija, H., van Hoof, H., and Precup, D. (2021). A survey of exploration methods in reinforcement learning. arXiv preprint arXiv:2109.00157
2021 arXiv
-
[19]
Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Man \'e , D. (2016). Concrete problems in ai safety. arXiv preprint arXiv:1606.06565
2016 arXiv
-
[20]
Anderson, A., Dodge, J., Sadarangani, A., Juozapaitis, Z., Newman, E., Irvine, J., Chattopadhyay, S., Fern, A., and Burnett, M. (2019). Explaining reinforcement learning to mere mortals: An empirical study. arXiv preprint arXiv:1903.09708
2019 arXiv
-
[21]
P., and Zaremba, W
Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W. (2017). Hindsight experience replay. In Advances in neural information processing systems , pages 5048--5058
2017
-
[22]
M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al
Andrychowicz, O. M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al. (2020). Learning dexterous in-hand manipulation. The International Journal of Robotics Research , 39(1):3--20
2020
-
[23]
and Doshi, P
Arora, S. and Doshi, P. (2021). A survey of inverse reinforcement learning: Challenges, methods and progress. Artificial Intelligence , 297:103500
2021
-
[24]
L., and Tellex, S
Arumugam, D., Karamcheti, S., Gopalan, N., Wong, L. L., and Tellex, S. (2017). Accurately and efficiently interpreting human-robot instructions of varying granularities. arXiv preprint arXiv:1704.06616
2017 arXiv
-
[25]
Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. (2021). A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861
2021 arXiv
-
[26]
J., Wang, B., and Bengio, Y
Atanackovic, L., Tong, A., Hartford, J., Lee, L. J., Wang, B., and Bengio, Y. (2023). Dyngfn: Bayesian dynamic causal discovery using generative flow networks. arXiv preprint arXiv:2302.04178
2023 arXiv
-
[27]
Audet, C., Bigeon, J., Cartier, D., Le Digabel, S., and Salomon, L. (2021). Performance indicators in multiobjective optimization. European journal of operational research , 292(2):397--422
2021
-
[28]
L., Kiros, J
Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016). Layer normalization. arXiv preprint arXiv:1607.06450
2016 arXiv
-
[29]
Bacon, P.-L., Harb, J., and Precup, D. (2017). The option-critic architecture. In Thirty-First AAAI Conference on Artificial Intelligence
2017
-
[30]
Bahdanau, D., Hill, F., Leike, J., Hughes, E., Hosseini, A., Kohli, P., and Grefenstette, E. (2018). Learning to understand goal specifications by modelling reward. arXiv preprint arXiv:1806.01946
2018 arXiv
-
[31]
Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. (2022). Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862
2022 arXiv
-
[32]
P., O’malley, M
Bajcsy, A., Losey, D. P., O’malley, M. K., and Dragan, A. D. (2017). Learning robot objectives from physical human interaction. In Conference on Robot Learning , pages 217--226. PMLR
2017
-
[33]
and Narayanan, S
Barrett, L. and Narayanan, S. (2008). Learning all optimal policies with multiple criteria. In Proceedings of the 25th international conference on Machine learning , pages 41--47
2008
-
[34]
L., Waytowich, N
Barton, S. L., Waytowich, N. R., Zaroukian, E., and Asher, D. E. (2018). Measuring collaborative emergent behavior in multi-agent reinforcement learning. In International Conference on Human Systems Engineering and Design: Future Trends and Applications , pages 422--427. Springer
2018
-
[35]
Beeching, E., Peter, M., Marcotte, P., Debangoye, J., Simonin, O., Romoff, J., and Wolf, C. (2021). Graph augmented deep reinforcement learning in the gamerland3d environment. arXiv preprint arXiv:2112.11731
2021 arXiv
-
[36]
G., Candido, S., Castro, P
Bellemare, M. G., Candido, S., Castro, P. S., Gong, J., Machado, M. C., Moitra, S., Ponda, S. S., and Wang, Z. (2020). Autonomous navigation of stratospheric balloons using reinforcement learning. Nature , 588(7836):77--82
2020
-
[37]
G., Dabney, W., and Munos, R
Bellemare, M. G., Dabney, W., and Munos, R. (2017). A distributional perspective on reinforcement learning. arXiv preprint arXiv:1707.06887
2017 arXiv
-
[38]
G., Naddaf, Y., Veness, J., and Bowling, M
Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013). The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research , 47:253--279
2013
-
[39]
Bellman, R. (1966). Dynamic programming. science , 153(3731):34--37
1966
-
[40]
Bengio, E., Jain, M., Korablyov, M., Precup, D., and Bengio, Y. (2021). Flow network based generative models for non-iterative diverse candidate generation. Advances in Neural Information Processing Systems , 34:27381--27394
2021
-
[41]
J., Tiwari, M., and Bengio, E
Bengio, Y., Lahlou, S., Deleu, T., Hu, E. J., Tiwari, M., and Bengio, E. (2023). Gflownet foundations. Journal of Machine Learning Research , 24(210):1--55
2023
-
[42]
Bergdahl, J., Gordillo, C., Tollmar, K., and Gissl \'e n, L. (2020). Augmenting automated game testing with deep reinforcement learning. In 2020 IEEE Conference on Games (CoG) , pages 600--603. IEEE
2020
-
[43]
Berner, C., Brockman, G., Chan, B., Cheung, V., D e biak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al. (2019). Dota 2 with large scale deep reinforcement learning. arXiv preprint arXiv:1912.06680
2019 arXiv
-
[44]
Bertsekas, D. P. (1997). Nonlinear programming. Journal of the Operational Research Society , 48(3):334--334
1997
-
[45]
R., Paolini, G
Bickerton, G. R., Paolini, G. V., Besnard, J., Muresan, S., and Hopkins, A. L. (2012). Quantifying the chemical beauty of drugs. Nature chemistry , 4(2):90--98
2012
-
[46]
Bohez, S., Abdolmaleki, A., Neunert, M., Buchli, J., Heess, N., and Hadsell, R. (2019). Value constrained model-free continuous control. arXiv preprint arXiv:1902.04623
2019 arXiv
-
[47]
B., Shah, J., Niekum, S., Stone, P., and Allievi, A
Booth, S., Knox, W. B., Shah, J., Niekum, S., Stone, P., and Allievi, A. (2023). The perils of trial-and-error reward design: misdesign through overfitting and invalid task specifications. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 37, pages 5920--5929
2023
-
[48]
Borkar, V. S. (2005). An actor-critic algorithm for constrained markov decision processes. Systems & control letters , 54(3):207--213
2005
-
[49]
Boularias, A., Kober, J., and Peters, J. (2011). Relative entropy inverse reinforcement learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics , pages 182--189. JMLR Workshop and Conference Proceedings
2011
-
[50]
D., Abel, D., and Dabney, W
Bowling, M., Martin, J. D., Abel, D., and Dabney, W. (2023). Settling the reward hypothesis. In International Conference on Machine Learning , pages 3003--3020. PMLR
2023
-
[51]
Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016). Openai gym. arXiv preprint arXiv:1606.01540
2016 arXiv
-
[52]
S., Goo, W., Nagarajan, P., and Niekum, S
Brown, D. S., Goo, W., Nagarajan, P., and Niekum, S. (2019a). Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations. arXiv preprint arXiv:1904.06387
2019 arXiv
-
[53]
H., and Vaucher, A
Brown, N., Fiscato, M., Segler, M. H., and Vaucher, A. C. (2019b). Guacamol: benchmarking models for de novo molecular design. Journal of chemical information and modeling , 59(3):1096--1108
2019
-
[54]
Brown, N., McKay, B., and Gasteiger, J. (2006). A novel workflow for the inverse qspr problem using multiobjective optimization. Journal of computer-aided molecular design , 20:333--341
2006
-
[55]
A., Mankowitz, D
Calian, D. A., Mankowitz, D. J., Zahavy, T., Xu, Z., Oh, J., Levine, N., and Mann, T. (2020). Balancing constraints and rewards with meta-gradient d4pg. arXiv preprint arXiv:2010.06324
2020 arXiv
-
[56]
Calinon, S., Evrard, P., Gribovskaya, E., Billard, A., and Kheddar, A. (2009). Learning collaborative manipulation tasks by demonstration using a haptic interface. In 2009 International Conference on Advanced Robotics , pages 1--6. IEEE
2009
-
[57]
Castelletti, A., Corani, G., Rizzolli, A., Soncinie-Sessa, R., and Weber, E. (2002). Reinforcement learning in the operational management of a water system. In IFAC workshop on modeling and control in environmental issues , pages 325--330. Keio University Yokohama
2002
-
[58]
Chai, J., Zeng, H., Li, A., and Ngai, E. W. (2021). Deep learning in computer vision: A critical review of emerging techniques and application scenarios. Machine Learning with Applications , 6:100134
2021
-
[59]
u rnkranz, J., H \
Cheng, W., F \"u rnkranz, J., H \"u llermeier, E., and Park, S.-H. (2011). Preference-based policy iteration: Leveraging preference learning for reinforcement learning. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2011, Athens, Greec...
2011
-
[60]
G., and Singh, S
Chentanez, N., Barto, A. G., and Singh, S. P. (2005). Intrinsically motivated reinforcement learning. In Advances in neural information processing systems , pages 1281--1288
2005
-
[61]
H., and Bengio, Y
Chevalier-Boisvert, M., Bahdanau, D., Lahlou, S., Willems, L., Saharia, C., Nguyen, T. H., and Bengio, Y. (2018). Babyai: A platform to study the sample efficiency of grounded language learning. arXiv preprint arXiv:1810.08272
2018 arXiv
-
[62]
and Kim, K.-E
Choi, J. and Kim, K.-E. (2011). Map inference for bayesian inverse reinforcement learning. Advances in neural information processing systems , 24
2011
-
[63]
Chow, Y., Ghavamzadeh, M., Janson, L., and Pavone, M. (2017). Risk-constrained reinforcement learning with percentile risk criteria. The Journal of Machine Learning Research , 18(1):6070--6120
2017
-
[64]
Chow, Y., Nachum, O., Duenez-Guzman, E., and Ghavamzadeh, M. (2018). A lyapunov-based approach to safe reinforcement learning. Advances in neural information processing systems , 31
2018
-
[65]
Chow, Y., Nachum, O., Faust, A., Duenez-Guzman, E., and Ghavamzadeh, M. (2019). Lyapunov-based safe policy optimization for continuous control. arXiv preprint arXiv:1901.10031
2019 arXiv
-
[66]
F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D
Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017). Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems , pages 4299--4307
2017
-
[67]
and Amodei, D
Clark, J. and Amodei, D. (2016). Faulty reward functions in the wild. Open AI
2016
-
[68]
Coello, C. A. C. and Cort \'e s, N. C. (2005). Solving multiobjective optimization problems using an artificial immune system. Genetic programming and evolvable machines , 6:163--190
2005
-
[69]
and Niekum, S
Cui, Y. and Niekum, S. (2018). Active reward learning from critiques. In 2018 IEEE international conference on robotics and automation (ICRA) , pages 6907--6914. IEEE
2018
-
[70]
Cybenko, G. (1989). Approximation by superpositions of a sigmoidal function. Mathematics of control, signals and systems , 2(4):303--314
1989
-
[71]
G., and Silver, D
Dabney, W., Barreto, A., Rowland, M., Dadashi, R., Quan, J., Bellemare, M. G., and Silver, D. (2021). The value-improvement path: Towards better representations for reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 35, pages 7160--7168
2021
-
[72]
G., and Munos, R
Dabney, W., Rowland, M., Bellemare, M. G., and Munos, R. (2018). Distributional reinforcement learning with quantile regression. In Thirty-Second AAAI Conference on Artificial Intelligence
2018
-
[73]
Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y. (2018). Safe exploration in continuous action spaces. arXiv preprint arXiv:1801.08757
2018 arXiv
-
[74]
B., Abelian, J., Abeyruwan, S., Ahn, M., Bewley, A., Boyd, J., Choromanski, K., Cortes, O., Coumans, E., Ding, T., et al
D'Ambrosio, D. B., Abelian, J., Abeyruwan, S., Ahn, M., Bewley, A., Boyd, J., Choromanski, K., Cortes, O., Coumans, E., Ding, T., et al. (2023). Robotic table tennis: A case study into a high speed learning system. arXiv preprint arXiv:2309.03315
2023 arXiv
-
[75]
Daniel, C., Viering, M., Metz, J., Kroemer, O., and Peters, J. (2014). Active reward learning. In Robotics: Science and systems , volume 98
2014
-
[76]
and Dennis, J
Das, I. and Dennis, J. E. (1997). A closer look at drawbacks of minimizing weighted sums of objectives for pareto set generation in multicriteria optimization problems. Structural optimization , 14:63--69
1997
-
[77]
David, H. A. (1963). The method of paired comparisons , volume 12. London
1963
-
[78]
and Hinton, G
Dayan, P. and Hinton, G. E. (1992). Feudal reinforcement learning. Advances in neural information processing systems , 5
1992
-
[79]
de Woillemont, P. L. P., Labory, R., and Corruble, V. (2021). Configurable agent with reward as input: A play-style continuum generation. In 2021 IEEE Conference on Games (CoG) , pages 1--8. IEEE
2021
-
[80]
Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B., Carpanese, F., Ewalds, T., Hafner, R., Abdolmaleki, A., de Las Casas, D., et al. (2022). Magnetic control of tokamak plasmas through deep reinforcement learning. Nature , 602(7897):414--419
2022
-
[81]
Degris, T., White, M., and Sutton, R. S. (2012). Off-policy actor-critic. arXiv preprint arXiv:1205.4839
2012 arXiv
-
[82]
Deleu, T., G \'o is, A., Emezue, C., Rankawat, M., Lacoste-Julien, S., Bauer, S., and Bengio, Y. (2022). Bayesian structure learning with generative flow networks. In Uncertainty in Artificial Intelligence , pages 518--528. PMLR
2022
-
[83]
Devlin, S., Georgescu, R., Momennejad, I., Rzepecki, J., Zuniga, E., Costello, G., Leroy, G., Shaw, A., and Hofmann, K. (2021). Navigation turing test (ntt): Learning to evaluate human-like navigation. arXiv preprint arXiv:2105.09637
2021 arXiv
-
[84]
Devlin, S. M. and Kudenko, D. (2012). Dynamic potential-based reward shaping. In Proceedings of the 11th international conference on autonomous agents and multiagent systems , pages 433--440. IFAAMAS
2012
-
[85]
Dewey, D. (2014). Reinforcement learning and the reward engineering principle. In 2014 AAAI Spring Symposium Series
2014
-
[86]
Ding, Y., Florensa, C., Abbeel, P., and Phielipp, M. (2019). Goal-conditioned imitation learning. In Advances in Neural Information Processing Systems (NeurIPS) , pages 15298--15309
2019
-
[87]
Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V. (2017). Carla: An open urban driving simulator. In Conference on robot learning , pages 1--16. PMLR
2017
-
[88]
M., Jayakumar, S
Du, Y., Czarnecki, W. M., Jayakumar, S. M., Farajtabar, M., Pascanu, R., and Lakshminarayanan, B. (2018). Adapting auxiliary losses using gradient similarity. arXiv preprint arXiv:1812.02224
2018 arXiv
-
[89]
J., Li, J., Paduraru, C., Gowal, S., and Hester, T
Dulac-Arnold, G., Levine, N., Mankowitz, D. J., Li, J., Paduraru, C., Gowal, S., and Hester, T. (2021). Challenges of real-world reinforcement learning: definitions, benchmarks and analysis. Machine Learning , 110(9):2419--2468
2021
-
[90]
Dulac-Arnold, G., Mankowitz, D., and Hester, T. (2019). Challenges of real-world reinforcement learning. arXiv preprint arXiv:1904.12901
2019 arXiv
-
[91]
Ehrgott, M. (2005). Multicriteria optimization , volume 491. Springer Science & Business Media
2005
-
[92]
Emmerich, M. T. and Deutz, A. H. (2018). A tutorial on multiobjective optimization: fundamentals and evolutionary methods. Natural computing , 17:585--609
2018
-
[93]
Erhan, D., Courville, A., Bengio, Y., and Vincent, P. (2010). Why does unsupervised pre-training help deep learning? In Proceedings of the thirteenth international conference on artificial intelligence and statistics , pages 201--208. JMLR Workshop and Conference Proceedings
2010
-
[94]
and Schuffenhauer, A
Ertl, P. and Schuffenhauer, A. (2009). Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. Journal of cheminformatics , 1:1--11
2009
-
[95]
and Gao, J
Evans, R. and Gao, J. (2016). Deepmind ai reduces google data centre cooling bill by 40 DeepMind
2016
-
[96]
G., and Larochelle, H
Fedus, W., Gelada, C., Bengio, Y., Bellemare, M. G., and Larochelle, H. (2019). Hyperbolic discounting and learning over multiple horizons. arXiv preprint arXiv:1902.06865
2019 arXiv
-
[97]
H., Bertsch, A., de Souza, J
Fernandes, P., Madaan, A., Liu, E., Farinhas, A., Martins, P. H., Bertsch, A., de Souza, J. G., Zhou, S., Wu, T., Neubig, G., et al. (2023). Bridging the gap: A survey on integrating (human) feedback for natural language generation. arXiv preprint arXiv:2305.00955
2023 arXiv
-
[98]
Finn, C., Christiano, P., Abbeel, P., and Levine, S. (2016a). A connection between generative adversarial networks, inverse reinforcement learning, and energy-based models. arXiv preprint arXiv:1611.03852
2016 arXiv
-
[99]
Finn, C., Levine, S., and Abbeel, P. (2016b). Guided cost learning: Deep inverse optimal control via policy optimization. In Proceedings of the 33rd International Conference on Machine Learning (ICML) , pages 49--58
2016
-
[100]
A., de Freitas, N., and Whiteson, S
Foerster, J., Assael, I. A., de Freitas, N., and Whiteson, S. (2016). Learning to communicate with deep multi-agent reinforcement learning. In Advances in Neural Information Processing Systems , pages 2137--2145
2016
-
[101]
Foerster, J., Song, F., Hughes, E., Burch, N., Dunning, I., Whiteson, S., Botvinick, M., and Bowling, M. (2019). Bayesian action decoder for deep multi-agent reinforcement learning. International Conference on Machine Learning
2019
-
[102]
N., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S
Foerster, J. N., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S. (2018). Counterfactual multi-agent policy gradients. In Thirty-Second AAAI Conference on Artificial Intelligence
2018
-
[103]
Fortnow, L. (2009). The status of the p versus np problem. Communications of the ACM , 52(9):78--86
2009
-
[104]
Fu, J., Korattikara, A., Levine, S., and Guadarrama, S. (2019). From language to goals: Inverse reinforcement learning for vision-based instruction following. arXiv preprint arXiv:1902.07742
2019 arXiv
-
[105]
Fu, J., Luo, K., and Levine, S. (2017). Learning robust rewards with adversarial inverse reinforcement learning. arXiv preprint arXiv:1710.11248
2017 arXiv
-
[106]
Fujimoto, S., Van Hoof, H., and Meger, D. (2018). Addressing function approximation error in actor-critic methods. arXiv preprint arXiv:1802.09477
2018 arXiv
-
[107]
G \'a bor, Z., Kalm \'a r, Z., and Szepesv \'a ri, C. (1998). Multi-criteria reinforcement learning. In ICML , volume 98, pages 197--205
1998
-
[108]
Ghasemipour, S. K. S., Zemel, R., and Gu, S. (2019). A divergence minimization perspective on imitation learning methods. In Proceedings of the 3rd Conference on Robot Learning (CoRL)
2019
-
[109]
Ghasemipour, S. K. S., Zemel, R., and Gu, S. (2020). A divergence minimization perspective on imitation learning methods. In Conference on Robot Learning , pages 1259--1277. PMLR
2020
-
[110]
Gissl \'e n, L., Eakins, A., Gordillo, C., Bergdahl, J., and Tollmar, K. (2021). Adversarial reinforcement learning for procedural content generation. arXiv preprint arXiv:2103.04847
2021 arXiv
-
[111]
Glaese, A., McAleese, N., Tr e bacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al. (2022). Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375
2022 arXiv
-
[112]
Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep learning . MIT press
2016
-
[113]
Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014). Generative adversarial nets. Advances in neural information processing systems , 27
2014
-
[114]
Gordillo, C., Bergdahl, J., Tollmar, K., and Gissl \'e n, L. (2021). Improving playtesting coverage via curiosity driven reinforcement learning agents. In 2021 IEEE Conference on Games (CoG) , pages 1--8. IEEE
2021
-
[115]
Goyal, P., Niekum, S., and Mooney, R. (2021). Pixl2r: Guiding reinforcement learning using natural language by mapping pixels to rewards. In Conference on Robot Learning , pages 485--497. PMLR
2021
-
[116]
Goyal, P., Niekum, S., and Mooney, R. J. (2019). Using natural language for reward shaping in reinforcement learning. arXiv preprint arXiv:1903.02020
2019 arXiv
-
[117]
Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. C. (2017). Improved training of W asserstein GAN s. In Advances in Neural Information Processing Systems (NeurIPS) , pages 5767--5777
2017
-
[118]
Gupta, A., Pacchiano, A., Zhai, Y., Kakade, S., and Levine, S. (2022). Unpacking reward shaping: Understanding the benefits of reward engineering on sample complexity. Advances in Neural Information Processing Systems , 35:15281--15295
2022
-
[119]
K., Egorov, M., and Kochenderfer, M
Gupta, J. K., Egorov, M., and Kochenderfer, M. J. (2017). Cooperative multi-agent control using deep reinforcement learning. In AAMAS Workshops
2017
-
[120]
Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. (2017). Reinforcement learning with deep energy-based policies. arXiv preprint arXiv:1702.08165
2017 arXiv
-
[121]
Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al. (2018). Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905
2018 arXiv
-
[122]
J., and Dragan, A
Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J., and Dragan, A. (2017). Inverse reward design. Advances in neural information processing systems , 30
2017
-
[123]
Harutyunyan, A., Devlin, S., Vrancx, P., and Now \'e , A. (2015). Expressing arbitrary reward functions as potential-based advice. In Proceedings of the AAAI conference on artificial intelligence , volume 29
2015
-
[124]
F., Howley, E., and Mannion, P
Hayes, C. F., Howley, E., and Mannion, P. (2020). Dynamic thresholded lexicograpic ordering. In Adaptive and Learning Agents Workshop (AAMAS 2020)
2020
-
[125]
a llstr \
Hayes, C. F., R a dulescu, R., Bargiacchi, E., K \"a llstr \"o m, J., Macfarlane, M., Reymond, M., Verstraeten, T., Zintgraf, L. M., Dazeley, R., Heintz, F., et al. (2022). A practical guide to multi-objective reinforcement learning and planning. Autonomous Agents and Multi-Ag...
2022
-
[126]
M., Singh, K., and Van Soest, A
Hazan, E., Kakade, S. M., Singh, K., and Van Soest, A. (2018). Provably efficient maximum entropy exploration. arXiv preprint arXiv:1812.02690
2018 arXiv
-
[127]
He, H., Boyd-Graber, J., Kwok, K., and Daum \'e III, H. (2016). Opponent modeling in deep reinforcement learning. In International Conference on Machine Learning , pages 1804--1813
2016
-
[128]
Hernandez-Leal, P., Kartal, B., and Taylor, M. E. (2018). Is multiagent deep reinforcement learning the answer or the question? a brief survey. arXiv preprint arXiv:1810.05587
2018 arXiv
-
[129]
Hernandez-Leal, P., Kartal, B., and Taylor, M. E. (2019a). Agent modeling as auxiliary task for deep reinforcement learning. In Proceedings of the AAAI conference on artificial intelligence and interactive digital entertainment , volume 15, pages 31--37
2019
-
[130]
Hernandez-Leal, P., Kartal, B., and Taylor, M. E. (2019b). Agent Modeling as Auxiliary Task for Deep Reinforcement Learning . In AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment
2019
-
[131]
Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D. (2018). Rainbow: Combining improvements in deep reinforcement learning. In Thirty-Second AAAI Conference on Artificial Intelligence
2018
-
[132]
and Ermon, S
Ho, J. and Ermon, S. (2016). Generative adversarial imitation learning. Advances in neural information processing systems , 29:4565--4573
2016
-
[133]
Hong, Z.-W., Su, S.-Y., Shann, T.-Y., Chang, Y.-H., and Lee, C.-Y. (2017). A deep policy inference q-network for multi-agent systems. arXiv preprint arXiv:1712.07893
2017 arXiv
-
[134]
Hu, E., Malkin, N., Jain, M., Everett, K., Graikos, A., and Bengio, Y. (2023). Gflownet-em for learning compositional latent variable models. arXiv preprint arXiv:2302.06576
2023 arXiv
-
[135]
Hu, Y., Wang, W., Jia, H., Wang, Y., Chen, Y., Hao, J., Wu, F., and Fan, C. (2020). Learning to utilize shaping rewards: A new approach of reward shaping. Advances in Neural Information Processing Systems , 33:15931--15941
2020
-
[136]
W., Xiao, C., Sun, J., and Zitnik, M
Huang, K., Fu, T., Gao, W., Zhao, Y., Roohani, Y., Leskovec, J., Coley, C. W., Xiao, C., Sun, J., and Zitnik, M. (2021). Therapeutics data commons: Machine learning datasets and tasks for drug discovery and development. arXiv preprint arXiv:2102.09548
2021 arXiv
-
[137]
Huang, W., Abbeel, P., Pathak, D., and Mordatch, I. (2022). Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning , pages 9118--9147. PMLR
2022
-
[138]
Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D. (2018). Reward learning from human preferences and demonstrations in atari. Advances in neural information processing systems , 31
2018
-
[139]
T., Klassen, T
Icarte, R. T., Klassen, T. Q., Valenzano, R., and McIlraith, S. A. (2022). Reward machines: Exploiting reward function structure in reinforcement learning. Journal of Artificial Intelligence Research , 73:173--208
2022
-
[140]
and Kuroe, Y
Iima, H. and Kuroe, Y. (2014). Multi-objective reinforcement learning for acquiring all pareto optimal policies simultaneously-method of determining scalarization weights. In 2014 IEEE International Conference on Systems, Man, and Cybernetics (SMC) , pages 876--881. IEEE
2014
-
[141]
and Szegedy, C
Ioffe, S. and Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167
2015 arXiv
-
[142]
and Sha, F
Iqbal, S. and Sha, F. (2019). Actor-attention-critic for multi-agent reinforcement learning. In International Conference on Machine Learning , pages 2961--2970
2019
-
[143]
it’s unwieldy and it takes a lot of time
Jacob, M., Devlin, S., and Hofmann, K. (2020). “it’s unwieldy and it takes a lot of time”—challenges and opportunities for creating agents in commercial games. In Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment , volume 16, p...
2020
-
[144]
M., Schaul, T., Leibo, J
Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. (2016). Reinforcement learning with unsupervised auxiliary tasks. arXiv preprint arXiv:1611.05397
2016 arXiv
-
[145]
Jain, A., Wojcik, B., Joachims, T., and Saxena, A. (2013). Learning trajectory preferences for manipulators via iterative improvement. Advances in neural information processing systems , 26
2013
-
[146]
F., Ekbote, C
Jain, M., Bengio, E., Hernandez-Garcia, A., Rector-Brooks, J., Dossou, B. F., Ekbote, C. A., Fu, J., Zhang, T., Kilgour, M., Zhang, D., et al. (2022a). Biological sequence design with gflownets. In International Conference on Machine Learning , pages 9786--9801. PMLR
2022
-
[147]
C., Hernandez-Garcia, A., Rector-Brooks, J., Bengio, Y., Miret, S., and Bengio, E
Jain, M., Raparthy, S. C., Hernandez-Garcia, A., Rector-Brooks, J., Bengio, Y., Miret, S., and Bengio, E. (2022b). Multi-objective gflownets. arXiv preprint arXiv:2210.12765
2022 arXiv
-
[148]
Jang, E., Gu, S., and Poole, B. (2017). Categorical reparametrization with gumble-softmax. In International Conference on Learning Representations (ICLR 2017) . OpenReview. net
2017
-
[149]
H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R
Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R. (2019a). Way off-policy batch deep reinforcement learning of implicit human preferences in dialog. arXiv preprint arXiv:1907.00456
2019 arXiv
-
[150]
Z., and De Freitas, N
Jaques, N., Lazaridou, A., Hughes, E., Gulcehre, C., Ortega, P., Strouse, D., Leibo, J. Z., and De Freitas, N. (2019b). Social influence as intrinsic motivation for multi-agent deep reinforcement learning. In International Conference on Machine Learning , pages 3040--3049
2019
-
[151]
J., Milli, S., and Dragan, A
Jeon, H. J., Milli, S., and Dragan, A. (2020). Reward-rational (implicit) choice: A unifying formalism for reward learning. Advances in Neural Information Processing Systems , 33:4415--4426
2020
-
[152]
and Lu, Z
Jiang, J. and Lu, Z. (2018). Learning attentional communication for multi-agent cooperation. In Advances in Neural Information Processing Systems , pages 7254--7264
2018
-
[153]
Jin, W., Barzilay, R., and Jaakkola, T. (2020). Multi-objective molecule generation using interpretable substructures. In International conference on machine learning , pages 4849--4859. PMLR
2020
-
[154]
Juliani, A., Berges, V.-P., Teng, E., Cohen, A., Harper, J., Elion, C., Goy, C., Gao, Y., Henry, H., Mattar, M., et al. (2018). Unity: A general platform for intelligent agents. arXiv preprint arXiv:1809.02627
2018 arXiv
-
[155]
Juozapaitis, Z., Koul, A., Fern, A., Erwig, M., and Doshi-Velez, F. (2019). Explainable reinforcement learning via reward decomposition. In IJCAI/ECAI Workshop on explainable artificial intelligence
2019
-
[156]
P., Littman, M
Kaelbling, L. P., Littman, M. L., and Moore, A. W. (1996). Reinforcement learning: A survey. Journal of artificial intelligence research , 4:237--285
1996
-
[157]
Kallenberg, L. (2011). Markov decision processes. Lecture Notes. University of Leiden , pages 65--66
2011
-
[158]
B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D
Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020). Scaling laws for neural language models. arXiv preprint arXiv:2001.08361
2020 arXiv
-
[159]
Kaplan, R., Sauer, C., and Sosa, A. (2017). Beating atari with natural language guided reinforcement learning. arXiv preprint arXiv:1704.05539
2017 arXiv
-
[160]
Karpathy, A. (2017). Software 2.0. Data Set
2017
-
[161]
Kartal, B., Hernandez-Leal, P., and Taylor, M. E. (2019). Terminal prediction as an auxiliary task for deep reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment , volume 15, pages 38--44
2019
-
[162]
Keeney, R., Raiffa, H., L, K., and Meyer, R. (1993). Decisions with Multiple Objectives: Preferences and Value Trade-Offs . Wiley series in probability and mathematical statistics. Applied probability and statistics. Cambridge University Press
1993
-
[163]
Kingma, D. P. and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980
2014 arXiv
-
[164]
Kingma, D. P. and Welling, M. (2013). Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114
2013 arXiv
-
[165]
Klissarov, M., D'Oro, P., Sodhani, S., Raileanu, R., Bacon, P.-L., Vincent, P., Zhang, A., and Henaff, M. (2023). Motif: Intrinsic motivation from artificial intelligence feedback. arXiv preprint arXiv:2310.00166
2023 arXiv
-
[166]
B., Allievi, A., Banzhaf, H., Schmitt, F., and Stone, P
Knox, W. B., Allievi, A., Banzhaf, H., Schmitt, F., and Stone, P. (2023). Reward (mis) design for autonomous driving. Artificial Intelligence , 316:103829
2023
-
[167]
Knox, W. B. and Stone, P. (2009). Interactively shaping agents via human reinforcement: The tamer framework. In Proceedings of the fifth international conference on Knowledge capture , pages 9--16
2009
-
[168]
Knox, W. B. and Stone, P. (2012). Reinforcement learning from simultaneous human and mdp reward. In AAMAS , volume 1004, pages 475--482. Valencia
2012
-
[169]
Korpelevich, G. M. (1976). The extragradient method for finding saddle points and other problems. Matecon , 12:747--756
1976
-
[170]
K., Dwibedi, D., Levine, S., and Tompson, J
Kostrikov, I., Agrawal, K. K., Dwibedi, D., Levine, S., and Tompson, J. (2018). Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning. arXiv preprint arXiv:1809.02925
2018 arXiv
-
[171]
K., Dwibedi, D., Levine, S., and Tompson, J
Kostrikov, I., Agrawal, K. K., Dwibedi, D., Levine, S., and Tompson, J. (2019). Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning. In Proceedings of the 7th International Conference on Learning Representations (ICLR)
2019
-
[172]
Kostrikov, I., Nachum, O., and Tompson, J. (2020). Imitation learning via off-policy distribution matching. In Proceedings of the 8th International Conference on Learning Representations (ICLR)
2020
-
[173]
Kreutzer, J., Khadivi, S., Matusov, E., and Riezler, S. (2018). Can neural machine translation be improved with user feedback? arXiv preprint arXiv:1804.05958
2018 arXiv
-
[174]
Kuefler, A., Morton, J., Wheeler, T., and Kochenderfer, M. (2017). Imitating driver behavior with generative adversarial networks. In Proceedings of 2017 IEEE Intelligent Vehicles Symposium (IV) , pages 204--211
2017
-
[175]
Kumar, A., Voet, A., and Zhang, K. Y. (2012). Fragment based drug design: from experimental to computational approaches. Current medicinal chemistry , 19(30):5128--5147
2012
-
[176]
Kurach, K., Raichuk, A., Sta \'n czyk, P., Zajac, M., Bachem, O., Espeholt, L., Riquelme, C., Vincent, D., Michalski, M., Bousquet, O., et al. (2019). Google research football: A novel reinforcement learning environment. arXiv preprint arXiv:1907.11180
2019 arXiv
-
[177]
M., Bullard, K., and Sadigh, D
Kwon, M., Xie, S. M., Bullard, K., and Sadigh, D. (2023). Reward design with language models. arXiv preprint arXiv:2303.00001
2023 arXiv
-
[178]
N., Bengio, Y., and Malkin, N
Lahlou, S., Deleu, T., Lemos, P., Zhang, D., Volokhova, A., Hern \'a ndez-Garc \' a, A., Ezzine, L. N., Bengio, Y., and Malkin, N. (2023). A theory of continuous generative flow networks. arXiv preprint arXiv:2301.12594
2023 arXiv
-
[179]
and Chaplot, D
Lample, G. and Chaplot, D. S. (2017). Playing fps games with deep reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 31
2017
-
[180]
Laskey, M., Lee, J., Fox, R., Dragan, A., and Goldberg, K. (2017). Dart: Noise injection for robust imitation learning. In Conference on robot learning , pages 143--156. PMLR
2017
-
[181]
Laskin, M., Srinivas, A., and Abbeel, P. (2020). Curl: Contrastive unsupervised representations for reinforcement learning. In International Conference on Machine Learning , pages 5639--5650. PMLR
2020
-
[182]
and DeJong, G
Laud, A. and DeJong, G. (2003). The influence of reward on the speed of reinforcement learning: An analysis of shaping. In Proceedings of the 20th International Conference on Machine Learning (ICML-03) , pages 440--447
2003
-
[183]
Lazaridou, A., Peysakhovich, A., and Baroni, M. (2016). Multi-agent cooperation and the emergence of (natural) language. arXiv preprint arXiv:1612.07182
2016 arXiv
-
[184]
LeCun, Y., Bengio, Y., and Hinton, G. (2015). Deep learning. nature , 521(7553):436--444
2015
-
[185]
LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998). Gradient-based learning applied to document recognition. Proceedings of the IEEE , 86(11):2278--2324
1998
-
[186]
Lee, K., Smith, L., and Abbeel, P. (2021). Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training. arXiv preprint arXiv:2106.05091
2021 arXiv
-
[187]
J., Bernard, S., Beslon, G., Bryson, D
Lehman, J., Clune, J., Misevic, D., Adami, C., Altenberg, L., Beaulieu, J., Bentley, P. J., Bernard, S., Beslon, G., Bryson, D. M., et al. (2020). The surprising creativity of digital evolution: A collection of anecdotes from the evolutionary computation and artificial life re...
2020
-
[188]
Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S. (2018). Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871
2018 arXiv
-
[189]
Levine, S., Popovic, Z., and Koltun, V. (2011). Nonlinear inverse reinforcement learning with gaussian processes. Advances in neural information processing systems , 24
2011
-
[190]
and Czarnecki, K
Li, C. and Czarnecki, K. (2018). Urban driving with multi-objective deep reinforcement learning. arXiv preprint arXiv:1811.08586
2018 arXiv
-
[191]
Li, K., Zhang, T., and Wang, R. (2020). Deep reinforcement learning for multiobjective optimization. IEEE transactions on cybernetics , 51(6):3103--3114
2020
-
[192]
Liang, Q., Que, F., and Modiano, E. (2018). Accelerated primal-dual policy optimization for safe reinforcement learning. arXiv preprint arXiv:1802.06480
2018 arXiv
-
[193]
P., Hunt, J
Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971
2015 arXiv
-
[194]
Lin, L.-J. (1993). Reinforcement learning for robots using neural networks. Technical report, Carnegie-Mellon Univ Pittsburgh PA School of Computer Science
1993
-
[195]
Lin, T., Jin, C., and Jordan, M. (2020). On gradient descent ascent for nonconvex-concave minimax problems. In International Conference on Machine Learning , pages 6083--6093. PMLR
2020
-
[196]
Lin, X., Baweja, H., Kantor, G., and Held, D. (2019a). Adaptive auxiliary task weighting for reinforcement learning. Advances in neural information processing systems , 32
2019
-
[197]
Lin, X., Zhen, H.-L., Li, Z., Zhang, Q.-F., and Kwong, S. (2019b). Pareto multi-task learning. Advances in neural information processing systems , 32
2019
-
[198]
Littman, M. L. (1994). Markov games as a framework for multi-agent reinforcement learning. In Machine learning proceedings 1994 , pages 157--163. Elsevier
1994
-
[199]
Liu, Y., Datta, G., Novoseller, E., and Brown, D. S. (2023). Efficient preference-based reinforcement learning using learned dynamics models. arXiv preprint arXiv:2301.04741
2023 arXiv
-
[200]
Liu, Y., Ding, J., and Liu, X. (2020). Ipo: Interior-point policy optimization under constraints. In Proceedings of the AAAI conference on artificial intelligence , volume 34, pages 4940--4947
2020
-
[201]
Liu, Y., Halev, A., and Liu, X. (2021). Policy learning with constraints in model-free reinforcement learning: A survey. In The 30th International Joint Conference on Artificial Intelligence (IJCAI)
2021
-
[202]
P., and Mordatch, I
Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, O. P., and Mordatch, I. (2017). Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems , pages 6379--6390
2017
-
[203]
Luketina, J., Nardelli, N., Farquhar, G., Foerster, J., Andreas, J., Grefenstette, E., Whiteson, S., and Rockt \"a schel, T. (2019). A survey of reinforcement learning informed by natural language. arXiv preprint arXiv:1906.03926
2019 arXiv
-
[204]
Q., Wu, N., et al
Luo, J., Paduraru, C., Voicu, O., Chervonyi, Y., Munns, S., Li, J., Qian, C., Dutta, P., Davis, J. Q., Wu, N., et al. (2022). Controlling commercial cooling systems using reinforcement learning. arXiv preprint arXiv:2211.07357
2022 arXiv
-
[205]
Lyle, C., Rowland, M., Ostrovski, G., and Dabney, W. (2021). On the effect of auxiliary tasks on representation dynamics. In International Conference on Artificial Intelligence and Statistics , pages 1--9. PMLR
2021
-
[206]
J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A
Ma, Y. J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A. (2023). Eureka: Human-level reward design via coding large language models
2023
-
[207]
L., Muresan, S., Squire, S., Tellex, S., Arumugam, D., and Yang, L
MacGlashan, J., Babes-Vroman, M., desJardins, M., Littman, M. L., Muresan, S., Squire, S., Tellex, S., Arumugam, D., and Yang, L. (2015). Grounding english commands to reward functions. In Robotics: Science and Systems
2015
-
[208]
K., Loftin, R., Peng, B., Wang, G., Roberts, D
MacGlashan, J., Ho, M. K., Loftin, R., Peng, B., Wang, G., Roberts, D. L., Taylor, M. E., and Littman, M. L. (2017). Interactive learning from policy-dependent human feedback. In International conference on machine learning , pages 2285--2294. PMLR
2017
-
[209]
Madan, K., Rector-Brooks, J., Korablyov, M., Bengio, E., Jain, M., Nica, A., Bosc, T., Bengio, Y., and Malkin, N. (2022). Learning gflownets from partial episodes for improved convergence and stability. arXiv preprint arXiv:2209.12782
2022 arXiv
-
[210]
C., Bosc, T., Bengio, Y., and Malkin, N
Madan, K., Rector-Brooks, J., Korablyov, M., Bengio, E., Jain, M., Nica, A. C., Bosc, T., Bengio, Y., and Malkin, N. (2023). Learning gflownets from partial episodes for improved convergence and stability. In International Conference on Machine Learning , pages 23467--23483. PMLR
2023
-
[211]
Mahajan, A., Rashid, T., Samvelyan, M., and Whiteson, S. (2019). Maven: Multi-agent variational exploration. In Advances in Neural Information Processing Systems , pages 7613--7624
2019
-
[212]
R., van Hasselt, H
Mahmood, A. R., van Hasselt, H. P., and Sutton, R. S. (2014). Weighted importance sampling for off-policy learning with linear function approximation. In Advances in Neural Information Processing Systems , pages 3014--3022
2014
-
[213]
Malkin, N., Jain, M., Bengio, E., Sun, C., and Bengio, Y. (2022a). Trajectory balance: Improved credit assignment in gflownets. Advances in Neural Information Processing Systems , 35:5955--5967
2022
-
[214]
Malkin, N., Lahlou, S., Deleu, T., Ji, X., Hu, E., Everett, K., Zhang, D., and Bengio, Y. (2022b). Gflownets and variational inference. arXiv preprint arXiv:2210.00580
2022 arXiv
-
[215]
Marchesini, E., Corsi, D., and Farinelli, A. (2022). Exploring safer behaviors for deep reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 36, pages 7701--7709
2022
-
[216]
Mathewson, K. W. and Pilarski, P. M. (2022). A brief guide to designing and evaluating human-centered interactive machine learning. arXiv preprint arXiv:2204.09622
2022 arXiv
-
[217]
Miettinen, K. (2012). Nonlinear multiobjective optimization , volume 12. Springer Science & Business Media
2012
-
[218]
Mindermann, S., Shah, R., Gleave, A., and Hadfield-Menell, D. (2018). Active inverse reward design. arXiv preprint arXiv:1809.03060
2018 arXiv
-
[219]
J., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, K., et al
Mirowski, P., Pascanu, R., Viola, F., Soyer, H., Ballard, A. J., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, K., et al. (2016). Learning to navigate in complex environments. arXiv preprint arXiv:1611.03673
2016 arXiv
-
[220]
Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013). Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602
2013 arXiv
-
[221]
A., Veness, J., Bellemare, M
Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015). Human-level control through deep reinforcement learning. nature , 518(7540):529--533
2015
-
[222]
M., Broekens, J., Plaat, A., Jonker, C
Moerland, T. M., Broekens, J., Plaat, A., Jonker, C. M., et al. (2023). Model-based reinforcement learning: A survey. Foundations and Trends in Machine Learning , 16(1):1--118
2023
-
[223]
and Lakshminarayanan, B
Mohamed, S. and Lakshminarayanan, B. (2016). Learning in implicit generative models. arXiv preprint arXiv:1610.03483
2016 arXiv
-
[224]
and Abbeel, P
Mordatch, I. and Abbeel, P. (2018). Emergence of grounded compositional language in multi-agent populations. In Thirty-Second AAAI Conference on Artificial Intelligence
2018
-
[225]
Mosqueira-Rey, E., Hern \'a ndez-Pereira, E., Alonso-R \' os, D., Bobes-Bascar \'a n, J., and Fern \'a ndez-Leal, \'A . (2023). Human-in-the-loop machine learning: A state of the art. Artificial Intelligence Review , 56(4):3005--3054
2023
-
[226]
M., Roijers, D
Mossalam, H., Assael, Y. M., Roijers, D. M., and Whiteson, S. (2016). Multi-objective deep reinforcement learning. arXiv preprint arXiv:1610.02707
2016 arXiv
-
[227]
Murphy, K. P. (2012). Machine learning: a probabilistic perspective . MIT press
2012
-
[228]
Nachum, O., Chow, Y., Dai, B., and Li, L. (2019). Dual DICE : Behavior-agnostic estimation of discounted stationary distribution corrections. In Advances in Neural Information Processing Systems (NeurIPS) , pages 2318--2328
2019
-
[229]
Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. (2018). Trust- PCL : An off-policy trust region method for continuous control. In Proceedings of the 6th International Conference on Learning Representations (ICLR)
2018
-
[230]
and Hinton, G
Nair, V. and Hinton, G. E. (2010). Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th international conference on machine learning (ICML-10) , pages 807--814
2010
-
[231]
Y., Harada, D., and Russell, S
Ng, A. Y., Harada, D., and Russell, S. (1999). Policy invariance under reward transformations: Theory and application to reward shaping. In ICML , volume 99, pages 278--287
1999
-
[232]
Y., Russell, S., et al
Ng, A. Y., Russell, S., et al. (2000). Algorithms for inverse reinforcement learning. In Icml , volume 1, page 2
2000
-
[233]
O'Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V. (2016). Combining policy gradient and q-learning. arXiv preprint arXiv:1611.01626
2016 arXiv
-
[234]
Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. (2016). Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499
2016 arXiv
-
[235]
W., Medina, J
Otter, D. W., Medina, J. R., and Kalita, J. K. (2020). A survey of the usages of deep learning for natural language processing. IEEE transactions on neural networks and learning systems , 32(2):604--624
2020
-
[236]
Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems , 35:27730--27744
2022
-
[237]
C., Shevchuk, G., and Sadigh, D
Palan, M., Landolfi, N. C., Shevchuk, G., and Sadigh, D. (2019). Learning reward functions by integrating human demonstrations and preferences. arXiv preprint arXiv:1906.08928
2019 arXiv
-
[238]
Pan, A., Bhatia, K., and Steinhardt, J. (2022). The effects of reward misspecification: Mapping and mitigating misaligned models. arXiv preprint arXiv:2201.03544
2022 arXiv
-
[239]
Pan, L., Malkin, N., Zhang, D., and Bengio, Y. (2023). Better training of gflownets with local credit and incomplete trajectories. arXiv preprint arXiv:2302.01687
2023 arXiv
-
[240]
Papadopoulos, A. I. and Linke, P. (2006). Multiobjective molecular design for integrated process-solvent systems synthesis. AIChE Journal , 52(3):1057--1070
2006
-
[241]
M., Z ilinskas, A., Z ilinskas, J., et al
Pardalos, P. M., Z ilinskas, A., Z ilinskas, J., et al. (2017). Non-convex multi-objective optimization . Springer
2017
-
[242]
Parisi, S., Pirotta, M., Smacchia, N., Bascetta, L., and Restelli, M. (2014). Policy gradient approaches for multi-objective sequential decision making. In 2014 International Joint Conference on Neural Networks (IJCNN) , pages 2323--2330. IEEE
2014
-
[243]
Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019). Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems , 32
2019
-
[244]
M., Dawson, M
Pilarski, P. M., Dawson, M. R., Degris, T., Fahimi, F., Carey, J. P., and Sutton, R. S. (2011). Online human training of a myoelectric prosthesis controller via actor-critic reinforcement learning. In 2011 IEEE international conference on rehabilitation robotics , pages 1--7. IEEE
2011
-
[245]
E., Wray, K
Pineda, L. E., Wray, K. H., and Zilberstein, S. (2015). Revisiting multi-objective mdps with relaxed lexicographic preferences. In 2015 AAAI Fall Symposium Series
2015
-
[246]
Polyak, B. (1970). Iterative methods using lagrange multipliers for solving extremal problems with constraints of the equation type. USSR Computational Mathematics and Mathematical Physics , 10(5):42--52
1970
-
[247]
Pomerleau, D. A. (1991). Efficient training of artificial neural networks for autonomous navigation. Neural computation , 3(1):88--97
1991
-
[248]
Popova, M., Isayev, O., and Tropsha, A. (2018). Deep reinforcement learning for de novo drug design. Science advances , 4(7):eaap7885
2018
-
[249]
Puterman, M. L. (1990). Markov decision processes. Handbooks in operations research and management science , 2:331--434
1990
-
[250]
and Amir, E
Ramachandran, D. and Amir, E. (2007). Bayesian inverse reinforcement learning. In IJCAI , volume 7, pages 2586--2591
2007
-
[251]
P., Luu, A
Ramp \'a s ek, L., Galkin, M., Dwivedi, V. P., Luu, A. T., Wolf, G., and Beaini, D. (2022). Recipe for a general, powerful, scalable graph transformer. Advances in Neural Information Processing Systems , 35:14501--14515
2022
-
[252]
and Alstr m, P
Randl v, J. and Alstr m, P. (1998). Learning to drive a bicycle using reinforcement learning and shaping. In ICML , volume 98, pages 463--471
1998
-
[253]
S., Farquhar, G., Foerster, J., and Whiteson, S
Rashid, T., Samvelyan, M., Witt, C. S., Farquhar, G., Foerster, J., and Whiteson, S. (2018). Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning. In International Conference on Machine Learning , pages 4292--4301
2018
-
[254]
D., Bagnell, J
Ratliff, N. D., Bagnell, J. A., and Zinkevich, M. A. (2006). Maximum margin planning. In Proceedings of the 23rd international conference on Machine learning , pages 729--736
2006
-
[255]
D., Silver, D., and Bagnell, J
Ratliff, N. D., Silver, D., and Bagnell, J. A. (2009). Learning to search: Functional gradient techniques for imitation learning. Autonomous Robots , 27:25--53
2009
-
[256]
Ratner, E., Hadfield-Menell, D., and Dragan, A. D. (2018). Simplifying reward design through divide-and-conquer. arXiv preprint arXiv:1806.02501
2018 arXiv
-
[257]
Ray, A., Achiam, J., and Amodei, D. (2019). Benchmarking safe exploration in deep reinforcement learning. arXiv preprint arXiv:1910.01708 , 7
2019 arXiv
-
[258]
D., and Levine, S
Reddy, S., Dragan, A. D., and Levine, S. (2019). SQIL : Imitation learning via reinforcement learning with sparse rewards
2019
-
[259]
Resnick, C., Eldridge, W., Ha, D., Britz, D., Foerster, J., Togelius, J., Cho, K., and Bruna, J. (2018). Pommerman: A multi-agent playground. arXiv preprint arXiv:1809.07124
2018 arXiv
-
[260]
F., Steckelmacher, D., Roijers, D
Reymond, M., Hayes, C. F., Steckelmacher, D., Roijers, D. M., and Now \'e , A. (2023). Actor-critic multi-objective reinforcement learning for non-linear utility functions. Autonomous Agents and Multi-Agent Systems , 37(2):23
2023
-
[261]
and Now \'e , A
Reymond, M. and Now \'e , A. (2019). Pareto-dqn: Approximating the pareto front in complex multi-objective decision problems. In Proceedings of the adaptive and learning agents workshop (ALA-19) at AAMAS
2019
-
[262]
and Mohamed, S
Rezende, D. and Mohamed, S. (2015). Variational inference with normalizing flows. In International conference on machine learning , pages 1530--1538. PMLR
2015
-
[263]
M., Vamplew, P., Whiteson, S., and Dazeley, R
Roijers, D. M., Vamplew, P., Whiteson, S., and Dazeley, R. (2013). A survey of multi-objective sequential decision-making. Journal of Artificial Intelligence Research , 48:67--113
2013
-
[264]
Rosenbaum, C., Klinger, T., and Riemer, M. (2017). Routing networks: Adaptive selection of non-linear functions for multi-task learning. arXiv preprint arXiv:1711.01239
2017 arXiv
-
[265]
Rosenblatt, F. (1958). The perceptron: a probabilistic model for information storage and organization in the brain. Psychological review , 65(6):386
1958
-
[266]
and Bagnell, D
Ross, S. and Bagnell, D. (2010). Efficient reductions for imitation learning. In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS) , pages 661--668
2010
-
[267]
Ross, S., Gordon, G., and Bagnell, D. (2011). A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics , pages 627--635. JMLR Workshop and Confe...
2011
-
[268]
Roy, J., Girgis, R., Romoff, J., Bacon, P.-L., and Pal, C. (2021). Direct behavior specification via constrained reinforcement learning. arXiv preprint arXiv:2112.12228
2021 arXiv
-
[269]
E., Hinton, G
Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986). Learning representations by back-propagating errors. nature , 323(6088):533--536
1986
-
[270]
Russell, S. (1998). Learning agents for uncertain environments. In Proceedings of the eleventh annual conference on Computational learning theory , pages 101--103
1998
-
[271]
Russell, S. J. and Zimdars, A. (2003). Q-decomposition for reinforcement learning agents. In Proceedings of the 20th International Conference on Machine Learning (ICML-03) , pages 656--663
2003
-
[272]
Rust, J. (2008). Dynamic programming. The new Palgrave dictionary of economics , 1:8
2008
-
[273]
D., Sastry, S., and Seshia, S
Sadigh, D., Dragan, A. D., Sastry, S., and Seshia, S. A. (2017). Active preference-based learning of reward functions
2017
-
[274]
Sasaki, F., Yohira, T., and Kawaguchi, A. (2018). Sample efficient imitation learning for continuous control. In Proceedings of the 6th International Conference on Learning Representations (ICLR)
2018
-
[275]
Saunders, W., Sastry, G., Stuhlmueller, A., and Evans, O. (2017). Trial without error: Towards safe reinforcement learning via human intervention. arXiv preprint arXiv:1707.05173
2017 arXiv
-
[276]
Schadd, F., Bakkes, S., and Spronck, P. (2007). Opponent modeling in real-time strategy games. In GAMEON , pages 61--70
2007
-
[277]
Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015a). Universal value function approximators. In International conference on machine learning , pages 1312--1320
2015
-
[278]
Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2015b). Prioritized experience replay. arXiv preprint arXiv:1511.05952
2015 arXiv
-
[279]
Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015). Trust region policy optimization. In Proceedings of the 32nd International Conference on Machine Learning (ICML) , pages 1889--1897
2015
-
[280]
Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347
2017 arXiv
-
[281]
and Amir, O
Septon, Y. and Amir, O. (2022). Integrating policy summaries with reward decomposition explanations. In ICAPS 2022 Workshop on Explainable AI Planning
2022
-
[282]
Settles, B. (2009). Active learning literature survey
2009
-
[283]
and Ghobaei-Arani, M
Shahidinejad, A. and Ghobaei-Arani, M. (2020). Joint computation offloading and resource provisioning for e dge-cloud computing environment: A machine learning-based approach. Software: Practice and Experience , 50(12):2212--2230
2020
-
[284]
Shao, K., Tang, Z., Zhu, Y., Li, N., and Zhao, D. (2019). A survey of deep reinforcement learning in video games. arXiv preprint arXiv:1912.10944
2019 arXiv
-
[285]
Shelhamer, E., Mahmoudieh, P., Argus, M., and Darrell, T. (2016). Loss is its own reward: Self-supervision for reinforcement learning. arXiv preprint arXiv:1612.07307
2016 arXiv
-
[286]
Siddique, U., Weng, P., and Zimmer, M. (2020). Learning fair policies in multi-objective (deep) reinforcement learning with average and discounted rewards. In International Conference on Machine Learning , pages 8905--8915. PMLR
2020
-
[287]
Silver, D., Bagnell, J., and Stentz, A. (2008). High performance outdoor navigation from overhead data using imitation learning. Robotics: Science and Systems IV, Zurich, Switzerland , 1
2008
-
[288]
J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al
Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016). Mastering the game of go with deep neural networks and tree search. nature , 529(7587):484--489
2016
-
[289]
Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014). Deterministic policy gradient algorithms. In International conference on machine learning , pages 387--395. Pmlr
2014
-
[290]
Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. (2017). Mastering the game of go without human knowledge. nature , 550(7676):354--359
2017
-
[291]
Silver, D., Singh, S., Precup, D., and Sutton, R. S. (2021). Reward is enough. Artificial Intelligence , 299:103535
2021
-
[292]
L., and Barto, A
Singh, S., Lewis, R. L., and Barto, A. G. (2009). Where do rewards come from. In Proceedings of the annual conference of the cognitive science society , pages 2601--2606. Cognitive Science Society
2009
-
[293]
L., Sorg, J., Barto, A
Singh, S., Lewis, R. L., Sorg, J., Barto, A. G., and Helou, A. (2010). On separating agent designer goals from agent goals: Breaking the preferences--parameters confound
2010
-
[294]
Skalse, J., Hammond, L., Griffin, C., and Abate, A. (2022). Lexicographic multi-objective reinforcement learning. arXiv preprint arXiv:2212.13769
2022 arXiv
-
[295]
Song, H., Li, A., Wang, T., and Wang, M. (2021). Multimodal deep reinforcement learning with auxiliary task for obstacle avoidance of indoor mobile robot. Sensors , 21(4):1363
2021
-
[296]
Sorg, J. D. (2011). The optimal reward problem: Designing effective reward for bounded agents . PhD thesis, University of Michigan
2011
-
[297]
Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014). Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research , 15(1):1929--1958
2014
-
[298]
St hl, N., Falkman, G., Karlsson, A., Mathiason, G., and Bostrom, J. (2019). Deep reinforcement learning for multiparameter optimization in de novo drug design. Journal of chemical information and modeling , 59(7):3166--3176
2019
-
[299]
Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. (2020). Learning to summarize with human feedback. Advances in Neural Information Processing Systems , 33:3008--3021
2020
-
[300]
Stooke, A., Achiam, J., and Abbeel, P. (2020). Responsive safety in reinforcement learning by pid lagrangian methods. In International Conference on Machine Learning , pages 9133--9143. PMLR
2020
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.