Pith. sign in

REVIEW 4 minor 300 references

Effective Reward Specification in Deep Reinforcement Learning

T0 review · 0 major / 4 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read Reward specification is the bottleneck in deep RL, and no single tool fixes it.

desk verdict A competent thesis compilation of four solid papers; valuable as a map of reward specification, not as a new research contribution. read the letter →

arxiv 2412.07177 v1 pith:AP45VC5I submitted 2024-12-10 cs.LG

classification cs.LG
keywords rewardspecificationdeepreinforcementlearningimitationmulti-agentcoordinationconstrainedGFlowNetsmoleculardesignsampleefficiency
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This thesis tries to establish that reward specification—the act of turning a human intention into a reward function—is one of the hardest parts of applying deep reinforcement learning, and that the field's many methods are complementary rather than interchangeable. The thesis makes the case by contributing four algorithms, each targeting a different failure mode: imitation learning that removes the policy-optimization loop, policy regularizers that promote multi-agent coordination, a constrained-RL framework that lets designers specify behavior directly, and goal-conditioned generative flow networks that explore the whole objective space in molecular design. Each contribution is designed to improve either sample efficiency or alignment, and the thesis's closing claim is that choosing the right specification tool for the application matters more than searching for a universal recipe.

What carries the argument

Four mechanisms carry the argument. Adversarial Soft Advantage Fitting uses a structured discriminator of the form $D(\tau) = \tilde{p}(\tau)/(\tilde{p}(\tau)+p_G(\tau))$, built from two evaluable policies, so that optimizing the discriminator simultaneously solves the generator's problem and yields the expert trajectory distribution without a reinforcement-learning loop. TeamReg adds team-spirit losses—each agent predicts its teammate's action and is regularized to be predictable—while CoachReg introduces a central coach that outputs a policy mask, with agents regularized to match the mask, producing synchronized sub-policy switches. The constrained-RL framework uses indicator cost functions with normalized Lagrange multipliers and a bootstrap constraint that keeps the multiplier scale from dominating the main objective. Goal-conditioned GFlowNets extend the flow-matching objective (trajectory balance) to condition on objective-space subregions, with a learned goal sampler that broadens coverage of the Pareto front.

What would settle it

A concrete test of the load-bearing assumption: train ASAF-1 on a task where the learner's initial policy and the expert visit largely disjoint state regions, and measure the state-occupancy divergence between them during training; if the method still recovers expert-level performance despite the divergence remaining large early on, the occupancy assumption is not the load-bearing part of the argument. A separate thesis-level test: if a single reward-specification method matched all four specialized tools without modification, the paper's central claim of no universal solution would fail.

Watch

Extended reading notes

Core claim

On the paper's own terms, the central discovery is that effective reward specification is not a single problem but a family of problems, and that each family has a characteristic failure mode that a purpose-built mechanism can address. In inverse reinforcement learning, the thesis shows that a discriminator conditioned on two policies—the previous generator and a learnable policy—can directly recover the expert policy, so the usual inner RL loop can be dropped entirely. In multi-agent learning, it shows that auxiliary objectives enforcing inter-agent predictability (TeamReg) and synchronized sub-policy selection (CoachReg) act as inductive biases that help agents discover coordinated strategies under sparse rewards. In constrained RL, it proposes indicator cost functions, multiplier normalization, and a bootstrap constraint so that hard behavioral requirements can be specified directly rather than through reward shaping. In molecular design, it shows that goal-conditioned GFlowNets, trained with a learned goal distribution over objective subregions, can generate molecules along the entire Pareto front. The thesis concludes that these results collectively support the absence of a universal solution to reward specification.

Load-bearing premise

The practical imitation results depend on a windowed approximation that assumes the learner and the expert visit the same states with the same frequencies, which is false until imitation has actually succeeded.

Editorial extensions

If this is right

  • If ASAF is right, adversarial imitation learning can be implemented and trained at roughly half the complexity, because the discriminator update itself produces the new policy; this removes the unstable alternation between RL and reward fitting.
  • If TeamReg and CoachReg are right, coordination-promoting regularizers can replace task-specific reward shaping and curriculum design in sparse-reward cooperative tasks, and can even improve hyperparameter robustness in multi-agent training.
  • If the constrained-RL framework is right, designers can monitor whether an agent satisfies hard behavioral requirements directly, since constraints are stated as costs rather than buried in a weighted reward.
  • If goal-conditioned GFlowNets are right, one trained model can be steered to different regions of the objective space at deployment time, making multi-objective molecular design a matter of choosing a goal rather than retraining.
  • Taken together, the thesis's four results imply that reward specification should be treated as a toolbox selection problem: demonstrations, auxiliary objectives, constraints, and goal conditioning each fit different applications.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • Beyond the thesis's claims, the ASAF windowed approximation suggests a practical diagnostic: tracking the divergence between learner and expert state-occupancy during training could tell a practitioner when the imitation signal is trustworthy and when it is fitting to the wrong distribution.
  • The thesis treats its four contributions separately, but they compose naturally: an ASAF-style reward model could seed the constrained-RL framework, and the multi-agent coordination regularizers could supply the behavioral constraints that the framework monitors.
  • The goal-conditioned GFlowNet approach implies a testable extension to other generative design problems—materials, circuits, or biological sequences—wherever a designer cares about covering trade-offs rather than a single optimum.
  • The thesis's 'no universal solution' claim, if correct, predicts that benchmark comparisons of reward-specification methods will keep showing environment-dependent winners; a meta-analysis of such comparisons would be a direct test.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

0 major / 4 minor

Summary. The manuscript is a PhD thesis that frames reward specification as a fundamental obstacle in deep reinforcement learning. It provides a technical background on deep learning and RL, a literature review organized around reward composition and reward modeling, and four original contribution chapters reproducing peer-reviewed published articles: ASAF (adversarial imitation learning without policy optimization), TeamReg/CoachReg (policy regularization for multi-agent coordination), constrained RL for direct behavior specification, and goal-conditioned GFlowNets for multi-objective molecular design. The thesis concludes that there is no universal reward specification solution and that practitioners should select tools according to the requirements of each application.

Significance. The four contribution chapters are already peer-reviewed, and the thesis adds value by placing them in a common taxonomy and by discussing their strengths and limitations. The ASAF formulation is a genuine simplification of adversarial imitation learning, with a theoretical statement for full trajectories and empirical support on several benchmarks. CoachReg provides a useful inductive bias for sparse-reward multi-agent coordination, and the constrained RL and goal-conditioned GFlowNet contributions offer practical tools for behavior specification and controllable multi-objective generation. The thesis is honest about its limitations, including the ASAF-1 windowed approximation (Section 5.3.3) and the negative TeamReg result on COMPROMISE (Table 6.1), which strengthens the credibility of the concluding 'no universal solution' claim. That claim is a qualitative synthesis rather than a formal theorem; within that scope, the evidence is adequate.

minor comments (4)
  1. [Abstract and Section 1.2] The term 'alignment' is used as a central success criterion in the abstract and introduction, but it is never defined operationally; a sentence specifying what counts as alignment in each contribution (e.g., constraint satisfaction rate, behavioral predictability, or Pareto coverage) would make the thesis's central claim easier to evaluate.
  2. [Section 5.3.3] The ASAF-1 approximation is described in a paragraph inside the algorithm box; because this is the main caveat of the practical variant, I suggest moving it to a dedicated 'Limitations' subsection and stating explicitly that Theorem 1 applies to full trajectories only.
  3. [Table 6.1 and Section 6.7.1] The negative result of TeamReg on COMPROMISE is visible in the table but not annotated; a footnote explaining that this failure mode is associated with an adversarial component, and stating whether significance testing was performed, would help readers weigh the cross-task claims.
  4. [Chapter 9] The heading 'Sucesses and Limitations' contains a typo and should read 'Successes and Limitations'.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity; the thesis is a compilation of independent, peer-reviewed contributions with disclosed approximations.

full rationale

I walked the derivation chains of the four contributions. Article 1 (ASAF) derives the structured-discriminator optimum from the standard logistic/GAN objective; the result that the optimal discriminator parameter equals the expert distribution is a mathematical identity of that objective, not a case of fitting a parameter and then predicting it back. The practical windowed variant ASAF-1 is explicitly acknowledged in Section 5.3.3 to assume equal state-occupancy measures until the expert policy is recovered; that is a disclosed limitation rather than a concealed input-to-output reduction. Articles 2-4 are empirical and algorithmic contributions: TeamReg/CoachReg add auxiliary objectives to MADDPG and are validated against baselines on independent environments; the constrained-RL framework and goal-conditioned GFlowNet approach are method proposals evaluated with external benchmarks. No load-bearing step is justified solely by a self-citation: the thesis reproduces the author's own published articles, but their results are experiment- and code-based, and the thesis does not invoke any self-authored uniqueness theorem to forbid alternatives. The meta-claim that no universal reward-specification solution exists is a survey-style conclusion, not a formal derivation from an assumption that already contains the conclusion. The self-citation present is structural (a doctoral thesis compiled from the author's papers), not load-bearing, and therefore does not constitute circularity.

Assumptions & free parameters 4 free parameters · 4 assumptions · 3 invented entities

The thesis's central claims rest on standard RL assumptions and a set of hyperparameters tuned per environment. The algorithmic entities (coach, policy masks, goal sampler) are new constructs with empirical support. No physical or mathematical entity is postulated without experimental backing.

free parameters (4)
  • ASAF window size w = 32 to 200 (tuned per environment)
    The windowed trajectory approximation ASAF-w treats sub-trajectories as full trajectories, and w was selected via hyperparameter search.
  • TeamReg coefficients lambda1, lambda2 = tuned per environment (see Appendix B.5)
    These weights balance predictability and predictability losses in the team-spirit objective.
  • CoachReg coefficients lambda1, lambda2, lambda3 and mask count K = tuned per environment
    Regularization coefficients and number of sub-policy masks are chosen by search.
  • Goal-conditioned GFlowNet hyperparameter m_g = set to control reward profile (Appendix D)
    Adjusts the conditioning sharpness for molecular objectives.
assumptions (4)
  • domain assumption Markov property and stationary policy assumption
    The thesis builds on MDPs and assumes stationary policies throughout the technical background (Section 2.2.1).
  • domain assumption The reward hypothesis: any goal can be captured by a scalar reward function
    The framing of reward specification relies on Sutton and Barto's reward hypothesis, which the thesis notes is a subject of debate (Section 3.1).
  • domain assumption Expert demonstrations are drawn from an optimal or near-optimal policy
    The imitation learning contributions (Chapter 5) assume the expert is optimal, as stated in the introduction to Article 1.
  • standard math Neural networks can approximate the required functions
    The thesis invokes the universal approximation theorem to justify the use of neural network policies and critics (Section 2.1.1).
invented entities (3)
  • Coach entity independent evidence
    purpose: Central model in CoachReg that outputs synchronized policy masks to coordinate agents during training
    Evaluated empirically in four multi-agent tasks; removed at test time.
  • Policy masks independent evidence
    purpose: One-hot vectors modulating hidden units to switch between sub-policies
    Used to implement synchronized sub-policy selection; shown to improve coordination in experiments.
  • Tabular Goal-Sampler independent evidence
    purpose: A learned goal distribution for conditioning GFlowNets on objective-space subregions
    Introduced in Article 4 to improve controllable molecular design; ablated in experiments.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Effective Reward Specification in Deep Reinforcement Learning." pith.science (2026). https://pith.science/paper/AP45VC5I

@misc{pith2026241207177,
  author       = {Pith},
  title        = {Pith review of: Effective Reward Specification in Deep Reinforcement Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/AP45VC5I}},
  note         = {Machine review of arXiv:2412.07177}
}
read the original abstract

In the last decade, Deep Reinforcement Learning has evolved into a powerful tool for complex sequential decision-making problems. It combines deep learning's proficiency in processing rich input signals with reinforcement learning's adaptability across diverse control tasks. At its core, an RL agent seeks to maximize its cumulative reward, enabling AI algorithms to uncover novel solutions previously unknown to experts. However, this focus on reward maximization also introduces a significant difficulty: improper reward specification can result in unexpected, misaligned agent behavior and inefficient learning. The complexity of accurately specifying the reward function is further amplified by the sequential nature of the task, the sparsity of learning signals, and the multifaceted aspects of the desired behavior. In this thesis, we survey the literature on effective reward specification strategies, identify core challenges relating to each of these approaches, and propose original contributions addressing the issue of sample efficiency and alignment in deep reinforcement learning. Reward specification represents one of the most challenging aspects of applying reinforcement learning in real-world domains. Our work underscores the absence of a universal solution to this complex and nuanced challenge; solving it requires selecting the most appropriate tools for the specific requirements of each unique application.

Figures

Figures reproduced from arXiv: 2412.07177 by the authors.

Figure 2
Figure 2. Markov Chain over state-action pairs . . . . . . . . . . . . . . . . . . 10 [PITH_FULL_IMAGE:figures/full_fig_p011_2.png] view at source ↗
Figure 7
Figure 7. Effect of the multiplier normalization . . . . . . . . . . . . . . . . . . 89 [PITH_FULL_IMAGE:figures/full_fig_p011_7.png] view at source ↗
Figure 8
Figure 8. Depiction of a GFlowNet Goal Sampler . . . . . . . . . . . . . . . . . 101 [PITH_FULL_IMAGE:figures/full_fig_p011_8.png] view at source ↗
Figures from the paper (21 more)
Figure 2
Figure 2. Figure 2: Markov Chain over state-action pairs [PITH_FULL_IMAGE:figures/full_fig_p025_2.png]
Figure 3
Figure 3. Figure 3: Overview of reward composition strategies. Reward composition consists in inte [PITH_FULL_IMAGE:figures/full_fig_p045_3.png]
Figure 3
Figure 3. Figure 3: ) [PITH_FULL_IMAGE:figures/full_fig_p051_3.png]
Figure 3
Figure 3. Figure 3: Overview of reward modeling strategies. Reward modeling consists in learning [PITH_FULL_IMAGE:figures/full_fig_p052_3.png]
Figure 5
Figure 5. Figure 5: Results on classic control and Box2D tasks for 10 expert demonstrations. First [PITH_FULL_IMAGE:figures/full_fig_p072_5.png]
Figure 5
Figure 5. Figure 5: Results on MuJoCo tasks for 25 expert demonstrations. [PITH_FULL_IMAGE:figures/full_fig_p073_5.png]
Figure 5
Figure 5. Figure 5: Results on Pommerman Random-Tag: (Left) Snapshot of the environment. (Cen [PITH_FULL_IMAGE:figures/full_fig_p074_5.png]
Figure 6
Figure 6. Figure 6: (Top) The tabular Q-learning [PITH_FULL_IMAGE:figures/full_fig_p080_6.png]
Figure 6
Figure 6. Figure 6: Illustration of CoachReg with two [PITH_FULL_IMAGE:figures/full_fig_p083_6.png]
Figure 6
Figure 6. Figure 6: reports the average learning curves and Table 6.1 presents the final performance. [PITH_FULL_IMAGE:figures/full_fig_p086_6.png]
Figure 6
Figure 6. Figure 6: Learning curves (mean return over agents) for our two proposed algorithms, two [PITH_FULL_IMAGE:figures/full_fig_p087_6.png]
Figure 6
Figure 6. Figure 6: shows the average entropy of the mask distributions for each environment com [PITH_FULL_IMAGE:figures/full_fig_p088_6.png]
Figure 6
Figure 6. Figure 6: Snapshot of the google research [PITH_FULL_IMAGE:figures/full_fig_p089_6.png]
Figure 7
Figure 7. Figure 7: Depictions of our setup to evaluate direct behavior specification using con [PITH_FULL_IMAGE:figures/full_fig_p094_7.png]
Figure 7
Figure 7. Figure 7: Enforcing behavioral constraints using reward engineering. Each grid represents [PITH_FULL_IMAGE:figures/full_fig_p095_7.png]
Figure 7
Figure 7. Figure 7: The multiplier normalisation keeps the learning dynamics stable when discovering [PITH_FULL_IMAGE:figures/full_fig_p104_7.png]
Figure 7
Figure 7. Figure 7: Each [PITH_FULL_IMAGE:figures/full_fig_p105_7.png]
Figure 7
Figure 7. Figure 7: A SAC-Lagrangian agent trained to solve the navigation problem in the Open [PITH_FULL_IMAGE:figures/full_fig_p106_7.png]
Figure 8
Figure 8. Figure 8: The diagram on the left depicts the state space of a GFlowNet molecule generator [PITH_FULL_IMAGE:figures/full_fig_p111_8.png]
Figure 8
Figure 8. Figure 8: Comparisons of the same sampling distributions depicted in Figure 8.2. Now the [PITH_FULL_IMAGE:figures/full_fig_p114_8.png]
Figure 8
Figure 8. Figure 8: Depiction of a GFlowNet Goal Sampler (GFN-GS) gradually building goal direc [PITH_FULL_IMAGE:figures/full_fig_p116_8.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

300 extracted references · 13 canonical work pages

  1. [1]

    write newline

    " write newline "" initialize.prev.this.status FUNCTION begin.bib " write newline preamble empty 'skip preamble write newline if " thebibliography " longest.label * " " * write newline " [1] #1 " write newline " url@samestyle " write newline " " write newline " [2] #2 " write newline " =0pt " write newline " " ALTinterwordstretchfactor * " " * write newli...

  2. [2]

    Abadi, M., Barham, P., Chen, J., Chen, Z., Davis, A., Dean, J., Devin, M., Ghemawat, S., Irving, G., Isard, M., et al. (2016). \ TensorFlow \ : a system for \ Large-Scale \ machine learning. In 12th USENIX symposium on operating systems design and implementation (OSDI 16) , pages 265--283

  3. [3]

    Abbeel, P., Coates, A., and Ng, A. Y. (2010). Autonomous helicopter aerobatics through apprenticeship learning. The International Journal of Robotics Research , 29(13):1608--1639

  4. [4]

    Abbeel, P., Coates, A., Quigley, M., and Ng, A. (2006). An application of reinforcement learning to aerobatic helicopter flight. Advances in neural information processing systems , 19

  5. [5]

    and Ng, A

    Abbeel, P. and Ng, A. Y. (2004). Apprenticeship learning via inverse reinforcement learning. In Proceedings of the twenty-first international conference on Machine learning , page 1

  6. [6]

    and Ng, A

    Abbeel, P. and Ng, A. Y. (2005). Exploration and apprenticeship learning in reinforcement learning. In Proceedings of the 22nd international conference on Machine learning , pages 1--8

  7. [7]

    K., Littman, M., Precup, D., and Singh, S

    Abel, D., Dabney, W., Harutyunyan, A., Ho, M. K., Littman, M., Precup, D., and Singh, S. (2021). On the expressivity of markov reward. Advances in Neural Information Processing Systems , 34:7799--7812

  8. [8]

    Abels, A., Roijers, D., Lenaerts, T., Now \'e , A., and Steckelmacher, D. (2019). Dynamic weights in multi-objective deep reinforcement learning. In International conference on machine learning , pages 11--20. PMLR

Show all 300 references
  1. [9]

    Achiam, J., Held, D., Tamar, A., and Abbeel, P. (2017). Constrained policy optimization. In International conference on machine learning , pages 22--31. PMLR

  2. [10]

    Adams, S., Cody, T., and Beling, P. A. (2022). A survey of inverse reinforcement learning. Artificial Intelligence Review , 55(6):4307--4346

  3. [11]

    and Dayan, P

    Ahilan, S. and Dayan, P. (2019). Feudal multi-agent hierarchies for cooperative reinforcement learning. arXiv preprint arXiv:1901.08492

  4. [12]

    Ahn, M., Brohan, A., Brown, N., Chebotar, Y., Cortes, O., David, B., Finn, C., Fu, C., Gopalakrishnan, K., Hausman, K., et al. (2022). Do as i can, not as i say: Grounding language in robotic affordances. arXiv preprint arXiv:2204.01691

  5. [13]

    Aissani, N., Beldjilali, B., and Trentesaux, D. (2008). Efficient and effective reactive scheduling of manufacturing system using sarsa-multi-objective agents. In MOSIM’08: 7th Conference Internationale de Modelisation et Simulation , pages 698--707

  6. [14]

    Akkaya, I., Andrychowicz, M., Chociej, M., Litwin, M., McGrew, B., Petron, A., Paino, A., Plappert, M., Powell, G., Ribas, R., et al. (2019). Solving rubik's cube with a robot hand. arXiv preprint arXiv:1910.07113

  7. [15]

    Akrour, R., Schoenauer, M., and Sebag, M. (2011). Preference-based policy learning. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2011, Athens, Greece, September 5-9, 2011. Proceedings, Part I 11 , pages 12--27. Springer

  8. [16]

    Alonso, E., Peter, M., Goumard, D., and Romoff, J. (2020). Deep reinforcement learning for navigation in aaa video games. arXiv preprint arXiv:2011.04764

  9. [17]

    Altman, E. (1999). Constrained Markov Decision Processes . Stochastic Modeling Series. Taylor & Francis

  10. [18]

    Amin, S., Gomrokchi, M., Satija, H., van Hoof, H., and Precup, D. (2021). A survey of exploration methods in reinforcement learning. arXiv preprint arXiv:2109.00157

  11. [19]

    Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., and Man \'e , D. (2016). Concrete problems in ai safety. arXiv preprint arXiv:1606.06565

  12. [20]

    Anderson, A., Dodge, J., Sadarangani, A., Juozapaitis, Z., Newman, E., Irvine, J., Chattopadhyay, S., Fern, A., and Burnett, M. (2019). Explaining reinforcement learning to mere mortals: An empirical study. arXiv preprint arXiv:1903.09708

  13. [21]

    P., and Zaremba, W

    Andrychowicz, M., Wolski, F., Ray, A., Schneider, J., Fong, R., Welinder, P., McGrew, B., Tobin, J., Abbeel, O. P., and Zaremba, W. (2017). Hindsight experience replay. In Advances in neural information processing systems , pages 5048--5058

  14. [22]

    M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al

    Andrychowicz, O. M., Baker, B., Chociej, M., Jozefowicz, R., McGrew, B., Pachocki, J., Petron, A., Plappert, M., Powell, G., Ray, A., et al. (2020). Learning dexterous in-hand manipulation. The International Journal of Robotics Research , 39(1):3--20

  15. [23]

    and Doshi, P

    Arora, S. and Doshi, P. (2021). A survey of inverse reinforcement learning: Challenges, methods and progress. Artificial Intelligence , 297:103500

  16. [24]

    L., and Tellex, S

    Arumugam, D., Karamcheti, S., Gopalan, N., Wong, L. L., and Tellex, S. (2017). Accurately and efficiently interpreting human-robot instructions of varying granularities. arXiv preprint arXiv:1704.06616

  17. [25]

    Askell, A., Bai, Y., Chen, A., Drain, D., Ganguli, D., Henighan, T., Jones, A., Joseph, N., Mann, B., DasSarma, N., et al. (2021). A general language assistant as a laboratory for alignment. arXiv preprint arXiv:2112.00861

  18. [26]

    J., Wang, B., and Bengio, Y

    Atanackovic, L., Tong, A., Hartford, J., Lee, L. J., Wang, B., and Bengio, Y. (2023). Dyngfn: Bayesian dynamic causal discovery using generative flow networks. arXiv preprint arXiv:2302.04178

  19. [27]

    Audet, C., Bigeon, J., Cartier, D., Le Digabel, S., and Salomon, L. (2021). Performance indicators in multiobjective optimization. European journal of operational research , 292(2):397--422

  20. [28]

    L., Kiros, J

    Ba, J. L., Kiros, J. R., and Hinton, G. E. (2016). Layer normalization. arXiv preprint arXiv:1607.06450

  21. [29]

    Bacon, P.-L., Harb, J., and Precup, D. (2017). The option-critic architecture. In Thirty-First AAAI Conference on Artificial Intelligence

  22. [30]

    Bahdanau, D., Hill, F., Leike, J., Hughes, E., Hosseini, A., Kohli, P., and Grefenstette, E. (2018). Learning to understand goal specifications by modelling reward. arXiv preprint arXiv:1806.01946

  23. [31]

    Bai, Y., Jones, A., Ndousse, K., Askell, A., Chen, A., DasSarma, N., Drain, D., Fort, S., Ganguli, D., Henighan, T., et al. (2022). Training a helpful and harmless assistant with reinforcement learning from human feedback. arXiv preprint arXiv:2204.05862

  24. [32]

    P., O’malley, M

    Bajcsy, A., Losey, D. P., O’malley, M. K., and Dragan, A. D. (2017). Learning robot objectives from physical human interaction. In Conference on Robot Learning , pages 217--226. PMLR

  25. [33]

    and Narayanan, S

    Barrett, L. and Narayanan, S. (2008). Learning all optimal policies with multiple criteria. In Proceedings of the 25th international conference on Machine learning , pages 41--47

  26. [34]

    L., Waytowich, N

    Barton, S. L., Waytowich, N. R., Zaroukian, E., and Asher, D. E. (2018). Measuring collaborative emergent behavior in multi-agent reinforcement learning. In International Conference on Human Systems Engineering and Design: Future Trends and Applications , pages 422--427. Springer

  27. [35]

    Beeching, E., Peter, M., Marcotte, P., Debangoye, J., Simonin, O., Romoff, J., and Wolf, C. (2021). Graph augmented deep reinforcement learning in the gamerland3d environment. arXiv preprint arXiv:2112.11731

  28. [36]

    G., Candido, S., Castro, P

    Bellemare, M. G., Candido, S., Castro, P. S., Gong, J., Machado, M. C., Moitra, S., Ponda, S. S., and Wang, Z. (2020). Autonomous navigation of stratospheric balloons using reinforcement learning. Nature , 588(7836):77--82

  29. [37]

    G., Dabney, W., and Munos, R

    Bellemare, M. G., Dabney, W., and Munos, R. (2017). A distributional perspective on reinforcement learning. arXiv preprint arXiv:1707.06887

  30. [38]

    G., Naddaf, Y., Veness, J., and Bowling, M

    Bellemare, M. G., Naddaf, Y., Veness, J., and Bowling, M. (2013). The arcade learning environment: An evaluation platform for general agents. Journal of Artificial Intelligence Research , 47:253--279

  31. [39]

    Bellman, R. (1966). Dynamic programming. science , 153(3731):34--37

  32. [40]

    Bengio, E., Jain, M., Korablyov, M., Precup, D., and Bengio, Y. (2021). Flow network based generative models for non-iterative diverse candidate generation. Advances in Neural Information Processing Systems , 34:27381--27394

  33. [41]

    J., Tiwari, M., and Bengio, E

    Bengio, Y., Lahlou, S., Deleu, T., Hu, E. J., Tiwari, M., and Bengio, E. (2023). Gflownet foundations. Journal of Machine Learning Research , 24(210):1--55

  34. [42]

    Bergdahl, J., Gordillo, C., Tollmar, K., and Gissl \'e n, L. (2020). Augmenting automated game testing with deep reinforcement learning. In 2020 IEEE Conference on Games (CoG) , pages 600--603. IEEE

  35. [43]

    Berner, C., Brockman, G., Chan, B., Cheung, V., D e biak, P., Dennison, C., Farhi, D., Fischer, Q., Hashme, S., Hesse, C., et al. (2019). Dota 2 with large scale deep reinforcement learning. arXiv preprint arXiv:1912.06680

  36. [44]

    Bertsekas, D. P. (1997). Nonlinear programming. Journal of the Operational Research Society , 48(3):334--334

  37. [45]

    R., Paolini, G

    Bickerton, G. R., Paolini, G. V., Besnard, J., Muresan, S., and Hopkins, A. L. (2012). Quantifying the chemical beauty of drugs. Nature chemistry , 4(2):90--98

  38. [46]

    Bohez, S., Abdolmaleki, A., Neunert, M., Buchli, J., Heess, N., and Hadsell, R. (2019). Value constrained model-free continuous control. arXiv preprint arXiv:1902.04623

  39. [47]

    B., Shah, J., Niekum, S., Stone, P., and Allievi, A

    Booth, S., Knox, W. B., Shah, J., Niekum, S., Stone, P., and Allievi, A. (2023). The perils of trial-and-error reward design: misdesign through overfitting and invalid task specifications. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 37, pages 5920--5929

  40. [48]

    Borkar, V. S. (2005). An actor-critic algorithm for constrained markov decision processes. Systems & control letters , 54(3):207--213

  41. [49]

    Boularias, A., Kober, J., and Peters, J. (2011). Relative entropy inverse reinforcement learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics , pages 182--189. JMLR Workshop and Conference Proceedings

  42. [50]

    D., Abel, D., and Dabney, W

    Bowling, M., Martin, J. D., Abel, D., and Dabney, W. (2023). Settling the reward hypothesis. In International Conference on Machine Learning , pages 3003--3020. PMLR

  43. [51]

    Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., and Zaremba, W. (2016). Openai gym. arXiv preprint arXiv:1606.01540

  44. [52]

    S., Goo, W., Nagarajan, P., and Niekum, S

    Brown, D. S., Goo, W., Nagarajan, P., and Niekum, S. (2019a). Extrapolating beyond suboptimal demonstrations via inverse reinforcement learning from observations. arXiv preprint arXiv:1904.06387

  45. [53]

    H., and Vaucher, A

    Brown, N., Fiscato, M., Segler, M. H., and Vaucher, A. C. (2019b). Guacamol: benchmarking models for de novo molecular design. Journal of chemical information and modeling , 59(3):1096--1108

  46. [54]

    Brown, N., McKay, B., and Gasteiger, J. (2006). A novel workflow for the inverse qspr problem using multiobjective optimization. Journal of computer-aided molecular design , 20:333--341

  47. [55]

    A., Mankowitz, D

    Calian, D. A., Mankowitz, D. J., Zahavy, T., Xu, Z., Oh, J., Levine, N., and Mann, T. (2020). Balancing constraints and rewards with meta-gradient d4pg. arXiv preprint arXiv:2010.06324

  48. [56]

    Calinon, S., Evrard, P., Gribovskaya, E., Billard, A., and Kheddar, A. (2009). Learning collaborative manipulation tasks by demonstration using a haptic interface. In 2009 International Conference on Advanced Robotics , pages 1--6. IEEE

  49. [57]

    Castelletti, A., Corani, G., Rizzolli, A., Soncinie-Sessa, R., and Weber, E. (2002). Reinforcement learning in the operational management of a water system. In IFAC workshop on modeling and control in environmental issues , pages 325--330. Keio University Yokohama

  50. [58]

    Chai, J., Zeng, H., Li, A., and Ngai, E. W. (2021). Deep learning in computer vision: A critical review of emerging techniques and application scenarios. Machine Learning with Applications , 6:100134

  51. [59]

    u rnkranz, J., H \

    Cheng, W., F \"u rnkranz, J., H \"u llermeier, E., and Park, S.-H. (2011). Preference-based policy iteration: Leveraging preference learning for reinforcement learning. In Machine Learning and Knowledge Discovery in Databases: European Conference, ECML PKDD 2011, Athens, Greec...

  52. [60]

    G., and Singh, S

    Chentanez, N., Barto, A. G., and Singh, S. P. (2005). Intrinsically motivated reinforcement learning. In Advances in neural information processing systems , pages 1281--1288

  53. [61]

    H., and Bengio, Y

    Chevalier-Boisvert, M., Bahdanau, D., Lahlou, S., Willems, L., Saharia, C., Nguyen, T. H., and Bengio, Y. (2018). Babyai: A platform to study the sample efficiency of grounded language learning. arXiv preprint arXiv:1810.08272

  54. [62]

    and Kim, K.-E

    Choi, J. and Kim, K.-E. (2011). Map inference for bayesian inverse reinforcement learning. Advances in neural information processing systems , 24

  55. [63]

    Chow, Y., Ghavamzadeh, M., Janson, L., and Pavone, M. (2017). Risk-constrained reinforcement learning with percentile risk criteria. The Journal of Machine Learning Research , 18(1):6070--6120

  56. [64]

    Chow, Y., Nachum, O., Duenez-Guzman, E., and Ghavamzadeh, M. (2018). A lyapunov-based approach to safe reinforcement learning. Advances in neural information processing systems , 31

  57. [65]

    Chow, Y., Nachum, O., Faust, A., Duenez-Guzman, E., and Ghavamzadeh, M. (2019). Lyapunov-based safe policy optimization for continuous control. arXiv preprint arXiv:1901.10031

  58. [66]

    F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D

    Christiano, P. F., Leike, J., Brown, T., Martic, M., Legg, S., and Amodei, D. (2017). Deep reinforcement learning from human preferences. In Advances in Neural Information Processing Systems , pages 4299--4307

  59. [67]

    and Amodei, D

    Clark, J. and Amodei, D. (2016). Faulty reward functions in the wild. Open AI

  60. [68]

    Coello, C. A. C. and Cort \'e s, N. C. (2005). Solving multiobjective optimization problems using an artificial immune system. Genetic programming and evolvable machines , 6:163--190

  61. [69]

    and Niekum, S

    Cui, Y. and Niekum, S. (2018). Active reward learning from critiques. In 2018 IEEE international conference on robotics and automation (ICRA) , pages 6907--6914. IEEE

  62. [70]

    Cybenko, G. (1989). Approximation by superpositions of a sigmoidal function. Mathematics of control, signals and systems , 2(4):303--314

  63. [71]

    G., and Silver, D

    Dabney, W., Barreto, A., Rowland, M., Dadashi, R., Quan, J., Bellemare, M. G., and Silver, D. (2021). The value-improvement path: Towards better representations for reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 35, pages 7160--7168

  64. [72]

    G., and Munos, R

    Dabney, W., Rowland, M., Bellemare, M. G., and Munos, R. (2018). Distributional reinforcement learning with quantile regression. In Thirty-Second AAAI Conference on Artificial Intelligence

  65. [73]

    Dalal, G., Dvijotham, K., Vecerik, M., Hester, T., Paduraru, C., and Tassa, Y. (2018). Safe exploration in continuous action spaces. arXiv preprint arXiv:1801.08757

  66. [74]

    B., Abelian, J., Abeyruwan, S., Ahn, M., Bewley, A., Boyd, J., Choromanski, K., Cortes, O., Coumans, E., Ding, T., et al

    D'Ambrosio, D. B., Abelian, J., Abeyruwan, S., Ahn, M., Bewley, A., Boyd, J., Choromanski, K., Cortes, O., Coumans, E., Ding, T., et al. (2023). Robotic table tennis: A case study into a high speed learning system. arXiv preprint arXiv:2309.03315

  67. [75]

    Daniel, C., Viering, M., Metz, J., Kroemer, O., and Peters, J. (2014). Active reward learning. In Robotics: Science and systems , volume 98

  68. [76]

    and Dennis, J

    Das, I. and Dennis, J. E. (1997). A closer look at drawbacks of minimizing weighted sums of objectives for pareto set generation in multicriteria optimization problems. Structural optimization , 14:63--69

  69. [77]

    David, H. A. (1963). The method of paired comparisons , volume 12. London

  70. [78]

    and Hinton, G

    Dayan, P. and Hinton, G. E. (1992). Feudal reinforcement learning. Advances in neural information processing systems , 5

  71. [79]

    de Woillemont, P. L. P., Labory, R., and Corruble, V. (2021). Configurable agent with reward as input: A play-style continuum generation. In 2021 IEEE Conference on Games (CoG) , pages 1--8. IEEE

  72. [80]

    Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B., Carpanese, F., Ewalds, T., Hafner, R., Abdolmaleki, A., de Las Casas, D., et al. (2022). Magnetic control of tokamak plasmas through deep reinforcement learning. Nature , 602(7897):414--419

  73. [81]

    Degris, T., White, M., and Sutton, R. S. (2012). Off-policy actor-critic. arXiv preprint arXiv:1205.4839

  74. [82]

    Deleu, T., G \'o is, A., Emezue, C., Rankawat, M., Lacoste-Julien, S., Bauer, S., and Bengio, Y. (2022). Bayesian structure learning with generative flow networks. In Uncertainty in Artificial Intelligence , pages 518--528. PMLR

  75. [83]

    Devlin, S., Georgescu, R., Momennejad, I., Rzepecki, J., Zuniga, E., Costello, G., Leroy, G., Shaw, A., and Hofmann, K. (2021). Navigation turing test (ntt): Learning to evaluate human-like navigation. arXiv preprint arXiv:2105.09637

  76. [84]

    Devlin, S. M. and Kudenko, D. (2012). Dynamic potential-based reward shaping. In Proceedings of the 11th international conference on autonomous agents and multiagent systems , pages 433--440. IFAAMAS

  77. [85]

    Dewey, D. (2014). Reinforcement learning and the reward engineering principle. In 2014 AAAI Spring Symposium Series

  78. [86]

    Ding, Y., Florensa, C., Abbeel, P., and Phielipp, M. (2019). Goal-conditioned imitation learning. In Advances in Neural Information Processing Systems (NeurIPS) , pages 15298--15309

  79. [87]

    Dosovitskiy, A., Ros, G., Codevilla, F., Lopez, A., and Koltun, V. (2017). Carla: An open urban driving simulator. In Conference on robot learning , pages 1--16. PMLR

  80. [88]

    M., Jayakumar, S

    Du, Y., Czarnecki, W. M., Jayakumar, S. M., Farajtabar, M., Pascanu, R., and Lakshminarayanan, B. (2018). Adapting auxiliary losses using gradient similarity. arXiv preprint arXiv:1812.02224

  81. [89]

    J., Li, J., Paduraru, C., Gowal, S., and Hester, T

    Dulac-Arnold, G., Levine, N., Mankowitz, D. J., Li, J., Paduraru, C., Gowal, S., and Hester, T. (2021). Challenges of real-world reinforcement learning: definitions, benchmarks and analysis. Machine Learning , 110(9):2419--2468

  82. [90]

    Dulac-Arnold, G., Mankowitz, D., and Hester, T. (2019). Challenges of real-world reinforcement learning. arXiv preprint arXiv:1904.12901

  83. [91]

    Ehrgott, M. (2005). Multicriteria optimization , volume 491. Springer Science & Business Media

  84. [92]

    Emmerich, M. T. and Deutz, A. H. (2018). A tutorial on multiobjective optimization: fundamentals and evolutionary methods. Natural computing , 17:585--609

  85. [93]

    Erhan, D., Courville, A., Bengio, Y., and Vincent, P. (2010). Why does unsupervised pre-training help deep learning? In Proceedings of the thirteenth international conference on artificial intelligence and statistics , pages 201--208. JMLR Workshop and Conference Proceedings

  86. [94]

    and Schuffenhauer, A

    Ertl, P. and Schuffenhauer, A. (2009). Estimation of synthetic accessibility score of drug-like molecules based on molecular complexity and fragment contributions. Journal of cheminformatics , 1:1--11

  87. [95]

    and Gao, J

    Evans, R. and Gao, J. (2016). Deepmind ai reduces google data centre cooling bill by 40 DeepMind

  88. [96]

    G., and Larochelle, H

    Fedus, W., Gelada, C., Bengio, Y., Bellemare, M. G., and Larochelle, H. (2019). Hyperbolic discounting and learning over multiple horizons. arXiv preprint arXiv:1902.06865

  89. [97]

    H., Bertsch, A., de Souza, J

    Fernandes, P., Madaan, A., Liu, E., Farinhas, A., Martins, P. H., Bertsch, A., de Souza, J. G., Zhou, S., Wu, T., Neubig, G., et al. (2023). Bridging the gap: A survey on integrating (human) feedback for natural language generation. arXiv preprint arXiv:2305.00955

  90. [98]

    Finn, C., Christiano, P., Abbeel, P., and Levine, S. (2016a). A connection between generative adversarial networks, inverse reinforcement learning, and energy-based models. arXiv preprint arXiv:1611.03852

  91. [99]

    Finn, C., Levine, S., and Abbeel, P. (2016b). Guided cost learning: Deep inverse optimal control via policy optimization. In Proceedings of the 33rd International Conference on Machine Learning (ICML) , pages 49--58

  92. [100]

    A., de Freitas, N., and Whiteson, S

    Foerster, J., Assael, I. A., de Freitas, N., and Whiteson, S. (2016). Learning to communicate with deep multi-agent reinforcement learning. In Advances in Neural Information Processing Systems , pages 2137--2145

  93. [101]

    Foerster, J., Song, F., Hughes, E., Burch, N., Dunning, I., Whiteson, S., Botvinick, M., and Bowling, M. (2019). Bayesian action decoder for deep multi-agent reinforcement learning. International Conference on Machine Learning

  94. [102]

    N., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S

    Foerster, J. N., Farquhar, G., Afouras, T., Nardelli, N., and Whiteson, S. (2018). Counterfactual multi-agent policy gradients. In Thirty-Second AAAI Conference on Artificial Intelligence

  95. [103]

    Fortnow, L. (2009). The status of the p versus np problem. Communications of the ACM , 52(9):78--86

  96. [104]

    Fu, J., Korattikara, A., Levine, S., and Guadarrama, S. (2019). From language to goals: Inverse reinforcement learning for vision-based instruction following. arXiv preprint arXiv:1902.07742

  97. [105]

    Fu, J., Luo, K., and Levine, S. (2017). Learning robust rewards with adversarial inverse reinforcement learning. arXiv preprint arXiv:1710.11248

  98. [106]

    Fujimoto, S., Van Hoof, H., and Meger, D. (2018). Addressing function approximation error in actor-critic methods. arXiv preprint arXiv:1802.09477

  99. [107]

    G \'a bor, Z., Kalm \'a r, Z., and Szepesv \'a ri, C. (1998). Multi-criteria reinforcement learning. In ICML , volume 98, pages 197--205

  100. [108]

    Ghasemipour, S. K. S., Zemel, R., and Gu, S. (2019). A divergence minimization perspective on imitation learning methods. In Proceedings of the 3rd Conference on Robot Learning (CoRL)

  101. [109]

    Ghasemipour, S. K. S., Zemel, R., and Gu, S. (2020). A divergence minimization perspective on imitation learning methods. In Conference on Robot Learning , pages 1259--1277. PMLR

  102. [110]

    Gissl \'e n, L., Eakins, A., Gordillo, C., Bergdahl, J., and Tollmar, K. (2021). Adversarial reinforcement learning for procedural content generation. arXiv preprint arXiv:2103.04847

  103. [111]

    Glaese, A., McAleese, N., Tr e bacz, M., Aslanides, J., Firoiu, V., Ewalds, T., Rauh, M., Weidinger, L., Chadwick, M., Thacker, P., et al. (2022). Improving alignment of dialogue agents via targeted human judgements. arXiv preprint arXiv:2209.14375

  104. [112]

    Goodfellow, I., Bengio, Y., and Courville, A. (2016). Deep learning . MIT press

  105. [113]

    Goodfellow, I., Pouget-Abadie, J., Mirza, M., Xu, B., Warde-Farley, D., Ozair, S., Courville, A., and Bengio, Y. (2014). Generative adversarial nets. Advances in neural information processing systems , 27

  106. [114]

    Gordillo, C., Bergdahl, J., Tollmar, K., and Gissl \'e n, L. (2021). Improving playtesting coverage via curiosity driven reinforcement learning agents. In 2021 IEEE Conference on Games (CoG) , pages 1--8. IEEE

  107. [115]

    Goyal, P., Niekum, S., and Mooney, R. (2021). Pixl2r: Guiding reinforcement learning using natural language by mapping pixels to rewards. In Conference on Robot Learning , pages 485--497. PMLR

  108. [116]

    Goyal, P., Niekum, S., and Mooney, R. J. (2019). Using natural language for reward shaping in reinforcement learning. arXiv preprint arXiv:1903.02020

  109. [117]

    Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., and Courville, A. C. (2017). Improved training of W asserstein GAN s. In Advances in Neural Information Processing Systems (NeurIPS) , pages 5767--5777

  110. [118]

    Gupta, A., Pacchiano, A., Zhai, Y., Kakade, S., and Levine, S. (2022). Unpacking reward shaping: Understanding the benefits of reward engineering on sample complexity. Advances in Neural Information Processing Systems , 35:15281--15295

  111. [119]

    K., Egorov, M., and Kochenderfer, M

    Gupta, J. K., Egorov, M., and Kochenderfer, M. J. (2017). Cooperative multi-agent control using deep reinforcement learning. In AAMAS Workshops

  112. [120]

    Haarnoja, T., Tang, H., Abbeel, P., and Levine, S. (2017). Reinforcement learning with deep energy-based policies. arXiv preprint arXiv:1702.08165

  113. [121]

    Haarnoja, T., Zhou, A., Hartikainen, K., Tucker, G., Ha, S., Tan, J., Kumar, V., Zhu, H., Gupta, A., Abbeel, P., et al. (2018). Soft actor-critic algorithms and applications. arXiv preprint arXiv:1812.05905

  114. [122]

    J., and Dragan, A

    Hadfield-Menell, D., Milli, S., Abbeel, P., Russell, S. J., and Dragan, A. (2017). Inverse reward design. Advances in neural information processing systems , 30

  115. [123]

    Harutyunyan, A., Devlin, S., Vrancx, P., and Now \'e , A. (2015). Expressing arbitrary reward functions as potential-based advice. In Proceedings of the AAAI conference on artificial intelligence , volume 29

  116. [124]

    F., Howley, E., and Mannion, P

    Hayes, C. F., Howley, E., and Mannion, P. (2020). Dynamic thresholded lexicograpic ordering. In Adaptive and Learning Agents Workshop (AAMAS 2020)

  117. [125]

    a llstr \

    Hayes, C. F., R a dulescu, R., Bargiacchi, E., K \"a llstr \"o m, J., Macfarlane, M., Reymond, M., Verstraeten, T., Zintgraf, L. M., Dazeley, R., Heintz, F., et al. (2022). A practical guide to multi-objective reinforcement learning and planning. Autonomous Agents and Multi-Ag...

  118. [126]

    M., Singh, K., and Van Soest, A

    Hazan, E., Kakade, S. M., Singh, K., and Van Soest, A. (2018). Provably efficient maximum entropy exploration. arXiv preprint arXiv:1812.02690

  119. [127]

    He, H., Boyd-Graber, J., Kwok, K., and Daum \'e III, H. (2016). Opponent modeling in deep reinforcement learning. In International Conference on Machine Learning , pages 1804--1813

  120. [128]

    Hernandez-Leal, P., Kartal, B., and Taylor, M. E. (2018). Is multiagent deep reinforcement learning the answer or the question? a brief survey. arXiv preprint arXiv:1810.05587

  121. [129]

    Hernandez-Leal, P., Kartal, B., and Taylor, M. E. (2019a). Agent modeling as auxiliary task for deep reinforcement learning. In Proceedings of the AAAI conference on artificial intelligence and interactive digital entertainment , volume 15, pages 31--37

  122. [130]

    Hernandez-Leal, P., Kartal, B., and Taylor, M. E. (2019b). Agent Modeling as Auxiliary Task for Deep Reinforcement Learning . In AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment

  123. [131]

    Hessel, M., Modayil, J., Van Hasselt, H., Schaul, T., Ostrovski, G., Dabney, W., Horgan, D., Piot, B., Azar, M., and Silver, D. (2018). Rainbow: Combining improvements in deep reinforcement learning. In Thirty-Second AAAI Conference on Artificial Intelligence

  124. [132]

    and Ermon, S

    Ho, J. and Ermon, S. (2016). Generative adversarial imitation learning. Advances in neural information processing systems , 29:4565--4573

  125. [133]

    Hong, Z.-W., Su, S.-Y., Shann, T.-Y., Chang, Y.-H., and Lee, C.-Y. (2017). A deep policy inference q-network for multi-agent systems. arXiv preprint arXiv:1712.07893

  126. [134]

    Hu, E., Malkin, N., Jain, M., Everett, K., Graikos, A., and Bengio, Y. (2023). Gflownet-em for learning compositional latent variable models. arXiv preprint arXiv:2302.06576

  127. [135]

    Hu, Y., Wang, W., Jia, H., Wang, Y., Chen, Y., Hao, J., Wu, F., and Fan, C. (2020). Learning to utilize shaping rewards: A new approach of reward shaping. Advances in Neural Information Processing Systems , 33:15931--15941

  128. [136]

    W., Xiao, C., Sun, J., and Zitnik, M

    Huang, K., Fu, T., Gao, W., Zhao, Y., Roohani, Y., Leskovec, J., Coley, C. W., Xiao, C., Sun, J., and Zitnik, M. (2021). Therapeutics data commons: Machine learning datasets and tasks for drug discovery and development. arXiv preprint arXiv:2102.09548

  129. [137]

    Huang, W., Abbeel, P., Pathak, D., and Mordatch, I. (2022). Language models as zero-shot planners: Extracting actionable knowledge for embodied agents. In International Conference on Machine Learning , pages 9118--9147. PMLR

  130. [138]

    Ibarz, B., Leike, J., Pohlen, T., Irving, G., Legg, S., and Amodei, D. (2018). Reward learning from human preferences and demonstrations in atari. Advances in neural information processing systems , 31

  131. [139]

    T., Klassen, T

    Icarte, R. T., Klassen, T. Q., Valenzano, R., and McIlraith, S. A. (2022). Reward machines: Exploiting reward function structure in reinforcement learning. Journal of Artificial Intelligence Research , 73:173--208

  132. [140]

    and Kuroe, Y

    Iima, H. and Kuroe, Y. (2014). Multi-objective reinforcement learning for acquiring all pareto optimal policies simultaneously-method of determining scalarization weights. In 2014 IEEE International Conference on Systems, Man, and Cybernetics (SMC) , pages 876--881. IEEE

  133. [141]

    and Szegedy, C

    Ioffe, S. and Szegedy, C. (2015). Batch normalization: Accelerating deep network training by reducing internal covariate shift. arXiv preprint arXiv:1502.03167

  134. [142]

    and Sha, F

    Iqbal, S. and Sha, F. (2019). Actor-attention-critic for multi-agent reinforcement learning. In International Conference on Machine Learning , pages 2961--2970

  135. [143]

    it’s unwieldy and it takes a lot of time

    Jacob, M., Devlin, S., and Hofmann, K. (2020). “it’s unwieldy and it takes a lot of time”—challenges and opportunities for creating agents in commercial games. In Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment , volume 16, p...

  136. [144]

    M., Schaul, T., Leibo, J

    Jaderberg, M., Mnih, V., Czarnecki, W. M., Schaul, T., Leibo, J. Z., Silver, D., and Kavukcuoglu, K. (2016). Reinforcement learning with unsupervised auxiliary tasks. arXiv preprint arXiv:1611.05397

  137. [145]

    Jain, A., Wojcik, B., Joachims, T., and Saxena, A. (2013). Learning trajectory preferences for manipulators via iterative improvement. Advances in neural information processing systems , 26

  138. [146]

    F., Ekbote, C

    Jain, M., Bengio, E., Hernandez-Garcia, A., Rector-Brooks, J., Dossou, B. F., Ekbote, C. A., Fu, J., Zhang, T., Kilgour, M., Zhang, D., et al. (2022a). Biological sequence design with gflownets. In International Conference on Machine Learning , pages 9786--9801. PMLR

  139. [147]

    C., Hernandez-Garcia, A., Rector-Brooks, J., Bengio, Y., Miret, S., and Bengio, E

    Jain, M., Raparthy, S. C., Hernandez-Garcia, A., Rector-Brooks, J., Bengio, Y., Miret, S., and Bengio, E. (2022b). Multi-objective gflownets. arXiv preprint arXiv:2210.12765

  140. [148]

    Jang, E., Gu, S., and Poole, B. (2017). Categorical reparametrization with gumble-softmax. In International Conference on Learning Representations (ICLR 2017) . OpenReview. net

  141. [149]

    H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R

    Jaques, N., Ghandeharioun, A., Shen, J. H., Ferguson, C., Lapedriza, A., Jones, N., Gu, S., and Picard, R. (2019a). Way off-policy batch deep reinforcement learning of implicit human preferences in dialog. arXiv preprint arXiv:1907.00456

  142. [150]

    Z., and De Freitas, N

    Jaques, N., Lazaridou, A., Hughes, E., Gulcehre, C., Ortega, P., Strouse, D., Leibo, J. Z., and De Freitas, N. (2019b). Social influence as intrinsic motivation for multi-agent deep reinforcement learning. In International Conference on Machine Learning , pages 3040--3049

  143. [151]

    J., Milli, S., and Dragan, A

    Jeon, H. J., Milli, S., and Dragan, A. (2020). Reward-rational (implicit) choice: A unifying formalism for reward learning. Advances in Neural Information Processing Systems , 33:4415--4426

  144. [152]

    and Lu, Z

    Jiang, J. and Lu, Z. (2018). Learning attentional communication for multi-agent cooperation. In Advances in Neural Information Processing Systems , pages 7254--7264

  145. [153]

    Jin, W., Barzilay, R., and Jaakkola, T. (2020). Multi-objective molecule generation using interpretable substructures. In International conference on machine learning , pages 4849--4859. PMLR

  146. [154]

    Juliani, A., Berges, V.-P., Teng, E., Cohen, A., Harper, J., Elion, C., Goy, C., Gao, Y., Henry, H., Mattar, M., et al. (2018). Unity: A general platform for intelligent agents. arXiv preprint arXiv:1809.02627

  147. [155]

    Juozapaitis, Z., Koul, A., Fern, A., Erwig, M., and Doshi-Velez, F. (2019). Explainable reinforcement learning via reward decomposition. In IJCAI/ECAI Workshop on explainable artificial intelligence

  148. [156]

    P., Littman, M

    Kaelbling, L. P., Littman, M. L., and Moore, A. W. (1996). Reinforcement learning: A survey. Journal of artificial intelligence research , 4:237--285

  149. [157]

    Kallenberg, L. (2011). Markov decision processes. Lecture Notes. University of Leiden , pages 65--66

  150. [158]

    B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D

    Kaplan, J., McCandlish, S., Henighan, T., Brown, T. B., Chess, B., Child, R., Gray, S., Radford, A., Wu, J., and Amodei, D. (2020). Scaling laws for neural language models. arXiv preprint arXiv:2001.08361

  151. [159]

    Kaplan, R., Sauer, C., and Sosa, A. (2017). Beating atari with natural language guided reinforcement learning. arXiv preprint arXiv:1704.05539

  152. [160]

    Karpathy, A. (2017). Software 2.0. Data Set

  153. [161]

    Kartal, B., Hernandez-Leal, P., and Taylor, M. E. (2019). Terminal prediction as an auxiliary task for deep reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment , volume 15, pages 38--44

  154. [162]

    Keeney, R., Raiffa, H., L, K., and Meyer, R. (1993). Decisions with Multiple Objectives: Preferences and Value Trade-Offs . Wiley series in probability and mathematical statistics. Applied probability and statistics. Cambridge University Press

  155. [163]

    Kingma, D. P. and Ba, J. (2014). Adam: A method for stochastic optimization. arXiv preprint arXiv:1412.6980

  156. [164]

    Kingma, D. P. and Welling, M. (2013). Auto-encoding variational bayes. arXiv preprint arXiv:1312.6114

  157. [165]

    Klissarov, M., D'Oro, P., Sodhani, S., Raileanu, R., Bacon, P.-L., Vincent, P., Zhang, A., and Henaff, M. (2023). Motif: Intrinsic motivation from artificial intelligence feedback. arXiv preprint arXiv:2310.00166

  158. [166]

    B., Allievi, A., Banzhaf, H., Schmitt, F., and Stone, P

    Knox, W. B., Allievi, A., Banzhaf, H., Schmitt, F., and Stone, P. (2023). Reward (mis) design for autonomous driving. Artificial Intelligence , 316:103829

  159. [167]

    Knox, W. B. and Stone, P. (2009). Interactively shaping agents via human reinforcement: The tamer framework. In Proceedings of the fifth international conference on Knowledge capture , pages 9--16

  160. [168]

    Knox, W. B. and Stone, P. (2012). Reinforcement learning from simultaneous human and mdp reward. In AAMAS , volume 1004, pages 475--482. Valencia

  161. [169]

    Korpelevich, G. M. (1976). The extragradient method for finding saddle points and other problems. Matecon , 12:747--756

  162. [170]

    K., Dwibedi, D., Levine, S., and Tompson, J

    Kostrikov, I., Agrawal, K. K., Dwibedi, D., Levine, S., and Tompson, J. (2018). Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning. arXiv preprint arXiv:1809.02925

  163. [171]

    K., Dwibedi, D., Levine, S., and Tompson, J

    Kostrikov, I., Agrawal, K. K., Dwibedi, D., Levine, S., and Tompson, J. (2019). Discriminator-actor-critic: Addressing sample inefficiency and reward bias in adversarial imitation learning. In Proceedings of the 7th International Conference on Learning Representations (ICLR)

  164. [172]

    Kostrikov, I., Nachum, O., and Tompson, J. (2020). Imitation learning via off-policy distribution matching. In Proceedings of the 8th International Conference on Learning Representations (ICLR)

  165. [173]

    Kreutzer, J., Khadivi, S., Matusov, E., and Riezler, S. (2018). Can neural machine translation be improved with user feedback? arXiv preprint arXiv:1804.05958

  166. [174]

    Kuefler, A., Morton, J., Wheeler, T., and Kochenderfer, M. (2017). Imitating driver behavior with generative adversarial networks. In Proceedings of 2017 IEEE Intelligent Vehicles Symposium (IV) , pages 204--211

  167. [175]

    Kumar, A., Voet, A., and Zhang, K. Y. (2012). Fragment based drug design: from experimental to computational approaches. Current medicinal chemistry , 19(30):5128--5147

  168. [176]

    Kurach, K., Raichuk, A., Sta \'n czyk, P., Zajac, M., Bachem, O., Espeholt, L., Riquelme, C., Vincent, D., Michalski, M., Bousquet, O., et al. (2019). Google research football: A novel reinforcement learning environment. arXiv preprint arXiv:1907.11180

  169. [177]

    M., Bullard, K., and Sadigh, D

    Kwon, M., Xie, S. M., Bullard, K., and Sadigh, D. (2023). Reward design with language models. arXiv preprint arXiv:2303.00001

  170. [178]

    N., Bengio, Y., and Malkin, N

    Lahlou, S., Deleu, T., Lemos, P., Zhang, D., Volokhova, A., Hern \'a ndez-Garc \' a, A., Ezzine, L. N., Bengio, Y., and Malkin, N. (2023). A theory of continuous generative flow networks. arXiv preprint arXiv:2301.12594

  171. [179]

    and Chaplot, D

    Lample, G. and Chaplot, D. S. (2017). Playing fps games with deep reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 31

  172. [180]

    Laskey, M., Lee, J., Fox, R., Dragan, A., and Goldberg, K. (2017). Dart: Noise injection for robust imitation learning. In Conference on robot learning , pages 143--156. PMLR

  173. [181]

    Laskin, M., Srinivas, A., and Abbeel, P. (2020). Curl: Contrastive unsupervised representations for reinforcement learning. In International Conference on Machine Learning , pages 5639--5650. PMLR

  174. [182]

    and DeJong, G

    Laud, A. and DeJong, G. (2003). The influence of reward on the speed of reinforcement learning: An analysis of shaping. In Proceedings of the 20th International Conference on Machine Learning (ICML-03) , pages 440--447

  175. [183]

    Lazaridou, A., Peysakhovich, A., and Baroni, M. (2016). Multi-agent cooperation and the emergence of (natural) language. arXiv preprint arXiv:1612.07182

  176. [184]

    LeCun, Y., Bengio, Y., and Hinton, G. (2015). Deep learning. nature , 521(7553):436--444

  177. [185]

    LeCun, Y., Bottou, L., Bengio, Y., and Haffner, P. (1998). Gradient-based learning applied to document recognition. Proceedings of the IEEE , 86(11):2278--2324

  178. [186]

    Lee, K., Smith, L., and Abbeel, P. (2021). Pebble: Feedback-efficient interactive reinforcement learning via relabeling experience and unsupervised pre-training. arXiv preprint arXiv:2106.05091

  179. [187]

    J., Bernard, S., Beslon, G., Bryson, D

    Lehman, J., Clune, J., Misevic, D., Adami, C., Altenberg, L., Beaulieu, J., Bentley, P. J., Bernard, S., Beslon, G., Bryson, D. M., et al. (2020). The surprising creativity of digital evolution: A collection of anecdotes from the evolutionary computation and artificial life re...

  180. [188]

    Leike, J., Krueger, D., Everitt, T., Martic, M., Maini, V., and Legg, S. (2018). Scalable agent alignment via reward modeling: a research direction. arXiv preprint arXiv:1811.07871

  181. [189]

    Levine, S., Popovic, Z., and Koltun, V. (2011). Nonlinear inverse reinforcement learning with gaussian processes. Advances in neural information processing systems , 24

  182. [190]

    and Czarnecki, K

    Li, C. and Czarnecki, K. (2018). Urban driving with multi-objective deep reinforcement learning. arXiv preprint arXiv:1811.08586

  183. [191]

    Li, K., Zhang, T., and Wang, R. (2020). Deep reinforcement learning for multiobjective optimization. IEEE transactions on cybernetics , 51(6):3103--3114

  184. [192]

    Liang, Q., Que, F., and Modiano, E. (2018). Accelerated primal-dual policy optimization for safe reinforcement learning. arXiv preprint arXiv:1802.06480

  185. [193]

    P., Hunt, J

    Lillicrap, T. P., Hunt, J. J., Pritzel, A., Heess, N., Erez, T., Tassa, Y., Silver, D., and Wierstra, D. (2015). Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971

  186. [194]

    Lin, L.-J. (1993). Reinforcement learning for robots using neural networks. Technical report, Carnegie-Mellon Univ Pittsburgh PA School of Computer Science

  187. [195]

    Lin, T., Jin, C., and Jordan, M. (2020). On gradient descent ascent for nonconvex-concave minimax problems. In International Conference on Machine Learning , pages 6083--6093. PMLR

  188. [196]

    Lin, X., Baweja, H., Kantor, G., and Held, D. (2019a). Adaptive auxiliary task weighting for reinforcement learning. Advances in neural information processing systems , 32

  189. [197]

    Lin, X., Zhen, H.-L., Li, Z., Zhang, Q.-F., and Kwong, S. (2019b). Pareto multi-task learning. Advances in neural information processing systems , 32

  190. [198]

    Littman, M. L. (1994). Markov games as a framework for multi-agent reinforcement learning. In Machine learning proceedings 1994 , pages 157--163. Elsevier

  191. [199]

    Liu, Y., Datta, G., Novoseller, E., and Brown, D. S. (2023). Efficient preference-based reinforcement learning using learned dynamics models. arXiv preprint arXiv:2301.04741

  192. [200]

    Liu, Y., Ding, J., and Liu, X. (2020). Ipo: Interior-point policy optimization under constraints. In Proceedings of the AAAI conference on artificial intelligence , volume 34, pages 4940--4947

  193. [201]

    Liu, Y., Halev, A., and Liu, X. (2021). Policy learning with constraints in model-free reinforcement learning: A survey. In The 30th International Joint Conference on Artificial Intelligence (IJCAI)

  194. [202]

    P., and Mordatch, I

    Lowe, R., Wu, Y., Tamar, A., Harb, J., Abbeel, O. P., and Mordatch, I. (2017). Multi-agent actor-critic for mixed cooperative-competitive environments. In Advances in Neural Information Processing Systems , pages 6379--6390

  195. [203]

    Luketina, J., Nardelli, N., Farquhar, G., Foerster, J., Andreas, J., Grefenstette, E., Whiteson, S., and Rockt \"a schel, T. (2019). A survey of reinforcement learning informed by natural language. arXiv preprint arXiv:1906.03926

  196. [204]

    Q., Wu, N., et al

    Luo, J., Paduraru, C., Voicu, O., Chervonyi, Y., Munns, S., Li, J., Qian, C., Dutta, P., Davis, J. Q., Wu, N., et al. (2022). Controlling commercial cooling systems using reinforcement learning. arXiv preprint arXiv:2211.07357

  197. [205]

    Lyle, C., Rowland, M., Ostrovski, G., and Dabney, W. (2021). On the effect of auxiliary tasks on representation dynamics. In International Conference on Artificial Intelligence and Statistics , pages 1--9. PMLR

  198. [206]

    J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A

    Ma, Y. J., Liang, W., Wang, G., Huang, D.-A., Bastani, O., Jayaraman, D., Zhu, Y., Fan, L., and Anandkumar, A. (2023). Eureka: Human-level reward design via coding large language models

  199. [207]

    L., Muresan, S., Squire, S., Tellex, S., Arumugam, D., and Yang, L

    MacGlashan, J., Babes-Vroman, M., desJardins, M., Littman, M. L., Muresan, S., Squire, S., Tellex, S., Arumugam, D., and Yang, L. (2015). Grounding english commands to reward functions. In Robotics: Science and Systems

  200. [208]

    K., Loftin, R., Peng, B., Wang, G., Roberts, D

    MacGlashan, J., Ho, M. K., Loftin, R., Peng, B., Wang, G., Roberts, D. L., Taylor, M. E., and Littman, M. L. (2017). Interactive learning from policy-dependent human feedback. In International conference on machine learning , pages 2285--2294. PMLR

  201. [209]

    Madan, K., Rector-Brooks, J., Korablyov, M., Bengio, E., Jain, M., Nica, A., Bosc, T., Bengio, Y., and Malkin, N. (2022). Learning gflownets from partial episodes for improved convergence and stability. arXiv preprint arXiv:2209.12782

  202. [210]

    C., Bosc, T., Bengio, Y., and Malkin, N

    Madan, K., Rector-Brooks, J., Korablyov, M., Bengio, E., Jain, M., Nica, A. C., Bosc, T., Bengio, Y., and Malkin, N. (2023). Learning gflownets from partial episodes for improved convergence and stability. In International Conference on Machine Learning , pages 23467--23483. PMLR

  203. [211]

    Mahajan, A., Rashid, T., Samvelyan, M., and Whiteson, S. (2019). Maven: Multi-agent variational exploration. In Advances in Neural Information Processing Systems , pages 7613--7624

  204. [212]

    R., van Hasselt, H

    Mahmood, A. R., van Hasselt, H. P., and Sutton, R. S. (2014). Weighted importance sampling for off-policy learning with linear function approximation. In Advances in Neural Information Processing Systems , pages 3014--3022

  205. [213]

    Malkin, N., Jain, M., Bengio, E., Sun, C., and Bengio, Y. (2022a). Trajectory balance: Improved credit assignment in gflownets. Advances in Neural Information Processing Systems , 35:5955--5967

  206. [214]

    Malkin, N., Lahlou, S., Deleu, T., Ji, X., Hu, E., Everett, K., Zhang, D., and Bengio, Y. (2022b). Gflownets and variational inference. arXiv preprint arXiv:2210.00580

  207. [215]

    Marchesini, E., Corsi, D., and Farinelli, A. (2022). Exploring safer behaviors for deep reinforcement learning. In Proceedings of the AAAI Conference on Artificial Intelligence , volume 36, pages 7701--7709

  208. [216]

    Mathewson, K. W. and Pilarski, P. M. (2022). A brief guide to designing and evaluating human-centered interactive machine learning. arXiv preprint arXiv:2204.09622

  209. [217]

    Miettinen, K. (2012). Nonlinear multiobjective optimization , volume 12. Springer Science & Business Media

  210. [218]

    Mindermann, S., Shah, R., Gleave, A., and Hadfield-Menell, D. (2018). Active inverse reward design. arXiv preprint arXiv:1809.03060

  211. [219]

    J., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, K., et al

    Mirowski, P., Pascanu, R., Viola, F., Soyer, H., Ballard, A. J., Banino, A., Denil, M., Goroshin, R., Sifre, L., Kavukcuoglu, K., et al. (2016). Learning to navigate in complex environments. arXiv preprint arXiv:1611.03673

  212. [220]

    Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., and Riedmiller, M. (2013). Playing atari with deep reinforcement learning. arXiv preprint arXiv:1312.5602

  213. [221]

    A., Veness, J., Bellemare, M

    Mnih, V., Kavukcuoglu, K., Silver, D., Rusu, A. A., Veness, J., Bellemare, M. G., Graves, A., Riedmiller, M., Fidjeland, A. K., Ostrovski, G., et al. (2015). Human-level control through deep reinforcement learning. nature , 518(7540):529--533

  214. [222]

    M., Broekens, J., Plaat, A., Jonker, C

    Moerland, T. M., Broekens, J., Plaat, A., Jonker, C. M., et al. (2023). Model-based reinforcement learning: A survey. Foundations and Trends in Machine Learning , 16(1):1--118

  215. [223]

    and Lakshminarayanan, B

    Mohamed, S. and Lakshminarayanan, B. (2016). Learning in implicit generative models. arXiv preprint arXiv:1610.03483

  216. [224]

    and Abbeel, P

    Mordatch, I. and Abbeel, P. (2018). Emergence of grounded compositional language in multi-agent populations. In Thirty-Second AAAI Conference on Artificial Intelligence

  217. [225]

    Mosqueira-Rey, E., Hern \'a ndez-Pereira, E., Alonso-R \' os, D., Bobes-Bascar \'a n, J., and Fern \'a ndez-Leal, \'A . (2023). Human-in-the-loop machine learning: A state of the art. Artificial Intelligence Review , 56(4):3005--3054

  218. [226]

    M., Roijers, D

    Mossalam, H., Assael, Y. M., Roijers, D. M., and Whiteson, S. (2016). Multi-objective deep reinforcement learning. arXiv preprint arXiv:1610.02707

  219. [227]

    Murphy, K. P. (2012). Machine learning: a probabilistic perspective . MIT press

  220. [228]

    Nachum, O., Chow, Y., Dai, B., and Li, L. (2019). Dual DICE : Behavior-agnostic estimation of discounted stationary distribution corrections. In Advances in Neural Information Processing Systems (NeurIPS) , pages 2318--2328

  221. [229]

    Nachum, O., Norouzi, M., Xu, K., and Schuurmans, D. (2018). Trust- PCL : An off-policy trust region method for continuous control. In Proceedings of the 6th International Conference on Learning Representations (ICLR)

  222. [230]

    and Hinton, G

    Nair, V. and Hinton, G. E. (2010). Rectified linear units improve restricted boltzmann machines. In Proceedings of the 27th international conference on machine learning (ICML-10) , pages 807--814

  223. [231]

    Y., Harada, D., and Russell, S

    Ng, A. Y., Harada, D., and Russell, S. (1999). Policy invariance under reward transformations: Theory and application to reward shaping. In ICML , volume 99, pages 278--287

  224. [232]

    Y., Russell, S., et al

    Ng, A. Y., Russell, S., et al. (2000). Algorithms for inverse reinforcement learning. In Icml , volume 1, page 2

  225. [233]

    O'Donoghue, B., Munos, R., Kavukcuoglu, K., and Mnih, V. (2016). Combining policy gradient and q-learning. arXiv preprint arXiv:1611.01626

  226. [234]

    Oord, A. v. d., Dieleman, S., Zen, H., Simonyan, K., Vinyals, O., Graves, A., Kalchbrenner, N., Senior, A., and Kavukcuoglu, K. (2016). Wavenet: A generative model for raw audio. arXiv preprint arXiv:1609.03499

  227. [235]

    W., Medina, J

    Otter, D. W., Medina, J. R., and Kalita, J. K. (2020). A survey of the usages of deep learning for natural language processing. IEEE transactions on neural networks and learning systems , 32(2):604--624

  228. [236]

    Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022). Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems , 35:27730--27744

  229. [237]

    C., Shevchuk, G., and Sadigh, D

    Palan, M., Landolfi, N. C., Shevchuk, G., and Sadigh, D. (2019). Learning reward functions by integrating human demonstrations and preferences. arXiv preprint arXiv:1906.08928

  230. [238]

    Pan, A., Bhatia, K., and Steinhardt, J. (2022). The effects of reward misspecification: Mapping and mitigating misaligned models. arXiv preprint arXiv:2201.03544

  231. [239]

    Pan, L., Malkin, N., Zhang, D., and Bengio, Y. (2023). Better training of gflownets with local credit and incomplete trajectories. arXiv preprint arXiv:2302.01687

  232. [240]

    Papadopoulos, A. I. and Linke, P. (2006). Multiobjective molecular design for integrated process-solvent systems synthesis. AIChE Journal , 52(3):1057--1070

  233. [241]

    M., Z ilinskas, A., Z ilinskas, J., et al

    Pardalos, P. M., Z ilinskas, A., Z ilinskas, J., et al. (2017). Non-convex multi-objective optimization . Springer

  234. [242]

    Parisi, S., Pirotta, M., Smacchia, N., Bascetta, L., and Restelli, M. (2014). Policy gradient approaches for multi-objective sequential decision making. In 2014 International Joint Conference on Neural Networks (IJCNN) , pages 2323--2330. IEEE

  235. [243]

    Paszke, A., Gross, S., Massa, F., Lerer, A., Bradbury, J., Chanan, G., Killeen, T., Lin, Z., Gimelshein, N., Antiga, L., et al. (2019). Pytorch: An imperative style, high-performance deep learning library. Advances in neural information processing systems , 32

  236. [244]

    M., Dawson, M

    Pilarski, P. M., Dawson, M. R., Degris, T., Fahimi, F., Carey, J. P., and Sutton, R. S. (2011). Online human training of a myoelectric prosthesis controller via actor-critic reinforcement learning. In 2011 IEEE international conference on rehabilitation robotics , pages 1--7. IEEE

  237. [245]

    E., Wray, K

    Pineda, L. E., Wray, K. H., and Zilberstein, S. (2015). Revisiting multi-objective mdps with relaxed lexicographic preferences. In 2015 AAAI Fall Symposium Series

  238. [246]

    Polyak, B. (1970). Iterative methods using lagrange multipliers for solving extremal problems with constraints of the equation type. USSR Computational Mathematics and Mathematical Physics , 10(5):42--52

  239. [247]

    Pomerleau, D. A. (1991). Efficient training of artificial neural networks for autonomous navigation. Neural computation , 3(1):88--97

  240. [248]

    Popova, M., Isayev, O., and Tropsha, A. (2018). Deep reinforcement learning for de novo drug design. Science advances , 4(7):eaap7885

  241. [249]

    Puterman, M. L. (1990). Markov decision processes. Handbooks in operations research and management science , 2:331--434

  242. [250]

    and Amir, E

    Ramachandran, D. and Amir, E. (2007). Bayesian inverse reinforcement learning. In IJCAI , volume 7, pages 2586--2591

  243. [251]

    P., Luu, A

    Ramp \'a s ek, L., Galkin, M., Dwivedi, V. P., Luu, A. T., Wolf, G., and Beaini, D. (2022). Recipe for a general, powerful, scalable graph transformer. Advances in Neural Information Processing Systems , 35:14501--14515

  244. [252]

    and Alstr m, P

    Randl v, J. and Alstr m, P. (1998). Learning to drive a bicycle using reinforcement learning and shaping. In ICML , volume 98, pages 463--471

  245. [253]

    S., Farquhar, G., Foerster, J., and Whiteson, S

    Rashid, T., Samvelyan, M., Witt, C. S., Farquhar, G., Foerster, J., and Whiteson, S. (2018). Qmix: Monotonic value function factorisation for deep multi-agent reinforcement learning. In International Conference on Machine Learning , pages 4292--4301

  246. [254]

    D., Bagnell, J

    Ratliff, N. D., Bagnell, J. A., and Zinkevich, M. A. (2006). Maximum margin planning. In Proceedings of the 23rd international conference on Machine learning , pages 729--736

  247. [255]

    D., Silver, D., and Bagnell, J

    Ratliff, N. D., Silver, D., and Bagnell, J. A. (2009). Learning to search: Functional gradient techniques for imitation learning. Autonomous Robots , 27:25--53

  248. [256]

    Ratner, E., Hadfield-Menell, D., and Dragan, A. D. (2018). Simplifying reward design through divide-and-conquer. arXiv preprint arXiv:1806.02501

  249. [257]

    Ray, A., Achiam, J., and Amodei, D. (2019). Benchmarking safe exploration in deep reinforcement learning. arXiv preprint arXiv:1910.01708 , 7

  250. [258]

    D., and Levine, S

    Reddy, S., Dragan, A. D., and Levine, S. (2019). SQIL : Imitation learning via reinforcement learning with sparse rewards

  251. [259]

    Resnick, C., Eldridge, W., Ha, D., Britz, D., Foerster, J., Togelius, J., Cho, K., and Bruna, J. (2018). Pommerman: A multi-agent playground. arXiv preprint arXiv:1809.07124

  252. [260]

    F., Steckelmacher, D., Roijers, D

    Reymond, M., Hayes, C. F., Steckelmacher, D., Roijers, D. M., and Now \'e , A. (2023). Actor-critic multi-objective reinforcement learning for non-linear utility functions. Autonomous Agents and Multi-Agent Systems , 37(2):23

  253. [261]

    and Now \'e , A

    Reymond, M. and Now \'e , A. (2019). Pareto-dqn: Approximating the pareto front in complex multi-objective decision problems. In Proceedings of the adaptive and learning agents workshop (ALA-19) at AAMAS

  254. [262]

    and Mohamed, S

    Rezende, D. and Mohamed, S. (2015). Variational inference with normalizing flows. In International conference on machine learning , pages 1530--1538. PMLR

  255. [263]

    M., Vamplew, P., Whiteson, S., and Dazeley, R

    Roijers, D. M., Vamplew, P., Whiteson, S., and Dazeley, R. (2013). A survey of multi-objective sequential decision-making. Journal of Artificial Intelligence Research , 48:67--113

  256. [264]

    Rosenbaum, C., Klinger, T., and Riemer, M. (2017). Routing networks: Adaptive selection of non-linear functions for multi-task learning. arXiv preprint arXiv:1711.01239

  257. [265]

    Rosenblatt, F. (1958). The perceptron: a probabilistic model for information storage and organization in the brain. Psychological review , 65(6):386

  258. [266]

    and Bagnell, D

    Ross, S. and Bagnell, D. (2010). Efficient reductions for imitation learning. In Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS) , pages 661--668

  259. [267]

    Ross, S., Gordon, G., and Bagnell, D. (2011). A reduction of imitation learning and structured prediction to no-regret online learning. In Proceedings of the fourteenth international conference on artificial intelligence and statistics , pages 627--635. JMLR Workshop and Confe...

  260. [268]

    Roy, J., Girgis, R., Romoff, J., Bacon, P.-L., and Pal, C. (2021). Direct behavior specification via constrained reinforcement learning. arXiv preprint arXiv:2112.12228

  261. [269]

    E., Hinton, G

    Rumelhart, D. E., Hinton, G. E., and Williams, R. J. (1986). Learning representations by back-propagating errors. nature , 323(6088):533--536

  262. [270]

    Russell, S. (1998). Learning agents for uncertain environments. In Proceedings of the eleventh annual conference on Computational learning theory , pages 101--103

  263. [271]

    Russell, S. J. and Zimdars, A. (2003). Q-decomposition for reinforcement learning agents. In Proceedings of the 20th International Conference on Machine Learning (ICML-03) , pages 656--663

  264. [272]

    Rust, J. (2008). Dynamic programming. The new Palgrave dictionary of economics , 1:8

  265. [273]

    D., Sastry, S., and Seshia, S

    Sadigh, D., Dragan, A. D., Sastry, S., and Seshia, S. A. (2017). Active preference-based learning of reward functions

  266. [274]

    Sasaki, F., Yohira, T., and Kawaguchi, A. (2018). Sample efficient imitation learning for continuous control. In Proceedings of the 6th International Conference on Learning Representations (ICLR)

  267. [275]

    Saunders, W., Sastry, G., Stuhlmueller, A., and Evans, O. (2017). Trial without error: Towards safe reinforcement learning via human intervention. arXiv preprint arXiv:1707.05173

  268. [276]

    Schadd, F., Bakkes, S., and Spronck, P. (2007). Opponent modeling in real-time strategy games. In GAMEON , pages 61--70

  269. [277]

    Schaul, T., Horgan, D., Gregor, K., and Silver, D. (2015a). Universal value function approximators. In International conference on machine learning , pages 1312--1320

  270. [278]

    Schaul, T., Quan, J., Antonoglou, I., and Silver, D. (2015b). Prioritized experience replay. arXiv preprint arXiv:1511.05952

  271. [279]

    Schulman, J., Levine, S., Abbeel, P., Jordan, M., and Moritz, P. (2015). Trust region policy optimization. In Proceedings of the 32nd International Conference on Machine Learning (ICML) , pages 1889--1897

  272. [280]

    Schulman, J., Wolski, F., Dhariwal, P., Radford, A., and Klimov, O. (2017). Proximal policy optimization algorithms. arXiv preprint arXiv:1707.06347

  273. [281]

    and Amir, O

    Septon, Y. and Amir, O. (2022). Integrating policy summaries with reward decomposition explanations. In ICAPS 2022 Workshop on Explainable AI Planning

  274. [282]

    Settles, B. (2009). Active learning literature survey

  275. [283]

    and Ghobaei-Arani, M

    Shahidinejad, A. and Ghobaei-Arani, M. (2020). Joint computation offloading and resource provisioning for e dge-cloud computing environment: A machine learning-based approach. Software: Practice and Experience , 50(12):2212--2230

  276. [284]

    Shao, K., Tang, Z., Zhu, Y., Li, N., and Zhao, D. (2019). A survey of deep reinforcement learning in video games. arXiv preprint arXiv:1912.10944

  277. [285]

    Shelhamer, E., Mahmoudieh, P., Argus, M., and Darrell, T. (2016). Loss is its own reward: Self-supervision for reinforcement learning. arXiv preprint arXiv:1612.07307

  278. [286]

    Siddique, U., Weng, P., and Zimmer, M. (2020). Learning fair policies in multi-objective (deep) reinforcement learning with average and discounted rewards. In International Conference on Machine Learning , pages 8905--8915. PMLR

  279. [287]

    Silver, D., Bagnell, J., and Stentz, A. (2008). High performance outdoor navigation from overhead data using imitation learning. Robotics: Science and Systems IV, Zurich, Switzerland , 1

  280. [288]

    J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al

    Silver, D., Huang, A., Maddison, C. J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al. (2016). Mastering the game of go with deep neural networks and tree search. nature , 529(7587):484--489

  281. [289]

    Silver, D., Lever, G., Heess, N., Degris, T., Wierstra, D., and Riedmiller, M. (2014). Deterministic policy gradient algorithms. In International conference on machine learning , pages 387--395. Pmlr

  282. [290]

    Silver, D., Schrittwieser, J., Simonyan, K., Antonoglou, I., Huang, A., Guez, A., Hubert, T., Baker, L., Lai, M., Bolton, A., et al. (2017). Mastering the game of go without human knowledge. nature , 550(7676):354--359

  283. [291]

    Silver, D., Singh, S., Precup, D., and Sutton, R. S. (2021). Reward is enough. Artificial Intelligence , 299:103535

  284. [292]

    L., and Barto, A

    Singh, S., Lewis, R. L., and Barto, A. G. (2009). Where do rewards come from. In Proceedings of the annual conference of the cognitive science society , pages 2601--2606. Cognitive Science Society

  285. [293]

    L., Sorg, J., Barto, A

    Singh, S., Lewis, R. L., Sorg, J., Barto, A. G., and Helou, A. (2010). On separating agent designer goals from agent goals: Breaking the preferences--parameters confound

  286. [294]

    Skalse, J., Hammond, L., Griffin, C., and Abate, A. (2022). Lexicographic multi-objective reinforcement learning. arXiv preprint arXiv:2212.13769

  287. [295]

    Song, H., Li, A., Wang, T., and Wang, M. (2021). Multimodal deep reinforcement learning with auxiliary task for obstacle avoidance of indoor mobile robot. Sensors , 21(4):1363

  288. [296]

    Sorg, J. D. (2011). The optimal reward problem: Designing effective reward for bounded agents . PhD thesis, University of Michigan

  289. [297]

    Srivastava, N., Hinton, G., Krizhevsky, A., Sutskever, I., and Salakhutdinov, R. (2014). Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research , 15(1):1929--1958

  290. [298]

    St hl, N., Falkman, G., Karlsson, A., Mathiason, G., and Bostrom, J. (2019). Deep reinforcement learning for multiparameter optimization in de novo drug design. Journal of chemical information and modeling , 59(7):3166--3176

  291. [299]

    Stiennon, N., Ouyang, L., Wu, J., Ziegler, D., Lowe, R., Voss, C., Radford, A., Amodei, D., and Christiano, P. F. (2020). Learning to summarize with human feedback. Advances in Neural Information Processing Systems , 33:3008--3021

  292. [300]

    Stooke, A., Achiam, J., and Abbeel, P. (2020). Responsive safety in reinforcement learning by pid lagrangian methods. In International Conference on Machine Learning , pages 9133--9143. PMLR

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.