Pith. sign in

REVIEW 3 major objections 6 minor 178 references

A Research Agenda for Usability and Generalisation in Reinforcement Learning

T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash

Pith's one-line read This paper argues that reinforcement learning environments should be described in user-friendly domain-specific or natural languages, and that complete descriptions, supplied to agents as context, are the route to zero-shot generalization…

desk verdict A clear, honest position paper that usefully spells out a research agenda for DSL/natural-language environment descriptions in RL—but its load-bearing assumption about user-friendliness remains quantitatively untested. read the letter →

arxiv 2412.16970 v2 pith:3L2STC2L submitted 2024-12-22 cs.AI stat.ML

classification cs.AIstat.ML
keywords reinforcementlearningenvironmentdescriptionlanguagesdomain-specificnaturallanguageinterfaceszero-shotgeneralizationusabilitybenchmarkscontextualMDP
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Reinforcement learning today assumes that each new task is a bespoke simulator written by an engineer in a general-purpose programming language. This position paper argues that this assumption blocks two goals: letting people without programming expertise apply RL to their own problems, and letting agents generalize instantly to new problems. The proposed fix is to describe environments in user-friendly domain-specific languages or natural language, compile those descriptions into simulators, and hand the same descriptions to agents as context. If the agenda is right, benchmark practice shifts from hand-coded environments toward description-native ones, and non-engineers become first-class RL users. The paper states its central position as a call for more benchmarks with environments defined in user-friendly DSLs or natural languages.

What carries the argument

The mechanism is the environment description as a dual-use object: a user-friendly formulation in a DSL or natural language that a compiler or language model translates into a runnable simulator, and that is simultaneously provided to the agent as context conditioning its policy or value function. Formalized as a contextual (PO)MDP or Markov game, this context must be complete enough to disambiguate between environments; incompleteness turns the collection of possible environments into a partially observable problem. The shared vocabulary of the description language is what makes the loop work in both directions: environments can be generated from contexts and contexts from environments, enabling procedural generation of training tasks and, in principle, zero-shot transfer to unseen descriptions.

What would settle it

Run a controlled usability study in which people with no programming background describe the same dozen tasks (board games, simple control problems) in a user-friendly DSL, in natural language, and in a general-purpose language, then measure whether the DSL and natural-language descriptions are more accurate, complete, and faster to produce. If novices produce unusable or incomplete descriptions at comparable rates, the usability and generalisation arguments for the agenda collapse.

Watch

Extended reading notes

Core claim

The paper's central claim is that the customary workflow—an engineer implementing each environment directly in a general-purpose programming language or a hardware-acceleration framework—is itself an obstacle to RL adoption and to generalization. It proposes that environments should be described in a shared, user-friendly language, with complete descriptions that can be compiled to executable simulators; those same descriptions, supplied to an agent as context, are argued to be a prerequisite for unrestricted zero-shot generalisation across every task expressible in the language. The paper supports this with a concrete example of a game description language that lets a user write tic-tac-toe as a short high-level script, and notes a library of over 1400 game descriptions contributed in part by non-programmers. It also identifies a practical rupture: succinct description languages often make it impossible to infer the full action space in advance, which violates a common assumption in deep RL APIs. The conclusion is a position statement: the RL community should place greater focus on benchmarks with environments defined in user-friendly DSLs or natural languages.

Load-bearing premise

The entire agenda rests on the assumption that non-engineers can write complete, unambiguous environment descriptions in a DSL or natural language more easily than in code; the paper itself admits it has no quantitative evidence for this.

Editorial extensions

If this is right

  • Benchmark suites should be built around description languages rather than hand-written simulators, and evaluation should test agents on unseen descriptions in the same language.
  • Non-engineers—private individuals, small organisations, and domain experts—could specify their own tasks and receive a policy without writing code, provided the compiler and a sufficiently general agent exist.
  • Complete, compileable descriptions are claimed as a prerequisite for unrestricted zero-shot generalisation; partial contexts such as numeric goal coordinates or short instructions cannot do the same job.
  • Existing assumptions like knowing the full action space in advance may fail for user-friendly DSLs, so deep RL methods will need to handle action aliasing or variable action spaces.
  • Procedural generation of new descriptions in the same language can supply a curriculum, letting agents learn the semantics of the language and generalise across the whole describable space.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The agenda implicitly predicts a convergence between RL environment design and language-model-driven program synthesis: if natural-language descriptions become the interface, the reliability of translating language to simulators becomes a core RL benchmark question rather than a side concern.
  • A testable extension is to measure how much zero-shot transfer performance scales with the number and diversity of descriptions seen during pretraining; the paper does not make a scaling-law claim, but its argument suggests such a relationship.
  • The complete-description requirement may be too strong for physical-world tasks, where dynamics are not fully describable in language; the paper allows reward-only descriptions for such cases, which leaves a gap between virtual and physical generalization.
  • If the position is adopted, evaluation methodology must separate what an agent learned about the description language from what it learned about general RL competence, since performance on unseen descriptions could come from either.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. This position paper argues that the customary practice of implementing RL environments in general-purpose programming languages imposes a usability barrier on non-engineer users and also obstructs progress on generalization, because there is no shared formalism in which different problems are represented. The authors advocate for a research agenda centered on describing environments in user-friendly domain-specific languages (DSLs) or natural language, such that (i) users with little programming expertise can formally describe their problems and (ii) algorithms can use the resulting complete descriptions as context to generalize zero-shot across all tasks describable in the chosen language. The paper states two explicit assumptions (Section 3.4), discusses potential issues such as action-space inference and simulation speed, notes other barriers to RL adoption (Section 5), and lays out desiderata and research directions (Section 6). It is an extended version of an earlier position paper, with additions covering other barriers and a more detailed agenda.

Significance. If the central position is correct, it would motivate a substantial shift in benchmark practice away from bespoke simulator code toward description-language-native environments, potentially democratizing RL for non-engineers and opening a new axis for zero-shot generalization research. The paper is honest and internally consistent: the two core assumptions are stated explicitly, the generalization claims in Section 4.3 are hedged, and Section 6 openly concedes the lack of quantitative evidence for the user-friendliness trade-off. It also credibly connects the agenda to concrete existing artifacts (Ludii, GAVEL, Ludax), gives a concrete desiderata list, and distinguishes its proposal from related DSL/context work. Its main weakness is that the load-bearing Assumption 1 rests on informal evidence and anecdotal experience rather than a user study; nevertheless the agenda is constructed so that this assumption could, in principle, be tested, which is a strength worth acknowledging.

major comments (3)
  1. [Section 3.4 and Section 6] Assumption 1 (that DSLs or natural language are more user-friendly than general-purpose programming languages for defining environments) is load-bearing for both halves of the agenda: the usability motivation in Section 3 and the generalization thesis in Section 4.3 both require non-engineers to be able to author complete, unambiguous descriptions. The only support offered is Ludii's design goal, the count of over 1400 game descriptions, and third-party forum contributions, while Section 6 admits "we have no quantitative evidence at this point." Because the entire agenda collapses if this assumption fails for the intended end users, the manuscript should either include or cite a user study that measures whether non-programmers can successfully write and validate complete environment descriptions, or explicitly elevate this to the first, falsifiable item of the research agenda with pre-registered success criteria.
  2. [Section 3.3] The proposed natural-language workflow requires an LLM to translate the description into a DSL, after which "a user can inspect the generated description and make corrections if necessary before it is compiled into a simulator." This verification-and-correction step still demands the ability to read and edit DSL code, which is precisely the expertise the agenda aims to remove. The paper should address how much DSL/verification competence is assumed of the end user and whether the verification step is realistically feasible for the target population; otherwise the usability advantage of natural-language descriptions is substantially weakened.
  3. [Section 4.3] The statement that complete environment descriptions "are likely to be a prerequisite for unrestricted, zero-shot generalisation in RL" is supported only by an analogy to humans learning new board games from rules, and the paper itself gives a video-game fire example where humans generalize without a complete description of the environment. This is not an internally inconsistent claim, but it is underspecified: the term "unrestricted" is never defined, and no concrete evidence or formal argument is given for why completeness is necessary rather than merely helpful. The claim should be reframed as a falsifiable hypothesis with a precise scope (e.g., across the set of tasks describable in a given DSL), and the authors should specify what experimental comparisons (such as context completeness versus zero-shot transfer on a DSL benchmark suite) would support or refute it.
minor comments (6)
  1. [Section 3.4] The word "exectuable" should be "executable."
  2. [Section 2.1] The phrase "when action according to a policy" should be "when acting according to a policy."
  3. [Section 3.2] The Ludii example may confuse readers unfamiliar with the language because the comment says some rules are omitted as defaults without explaining what those defaults are; a short note on the default turn-taking and draw conditions would improve readability.
  4. [Section 6] The t-SNE figure (Fig. 2) is described only as "reduced from a larger feature space" with a citation; the caption could usefully state which features from [116] were used and how the embedding was computed.
  5. [References] Reference [76] contains "hum4n l4ngu4ge" which appears to be either a deliberate obfuscation or a transcription error; if deliberate, the authors should add a note, as it may confuse readers.
  6. [Section 4.1] The term "unrestricted, zero-shot generalisation" is used in Section 4.3 before being defined; a definition or at least a clarifying sentence in Section 4.1 would help.

Circularity Check

0 steps flagged · score 2.0 of 10

No circular derivation in this position paper; the only borderline issue is self-referential evidence for Assumption 1, which is not a load-bearing circular step.

full rationale

This is a position paper, not a derivation chain. It contains no fitted parameters, no equations whose outputs are also inputs, and no 'prediction' that is statistically forced. The central Position (Section 3.4) is a normative call for more DSL/natural-language benchmarks, supported by two explicitly stated assumptions. Section 4.3 argues that complete environment descriptions are 'likely to be a prerequisite' for unrestricted zero-shot generalisation, citing external theoretical work ([56]) and using an analogy to human board-game learning; it is not presented as a theorem and does not reduce to its inputs. The main self-reference appears in Assumption 1: the paper supports the claim that DSLs can be user-friendly by citing the authors' own Ludii design goal [115] and the >1400-game library on the authors' platform, and Section 6 concedes 'we have no quantitative evidence at this point' for the user-friendliness/generality trade-off. This is self-referential evidence rather than circular derivation: the conclusion is not equivalent to the citation by construction, and the library statistics are publicly checkable. Under the hard rule requiring a quoted reduction, no actual circular step can be exhibited, so the appropriate finding is no significant circularity.

Assumptions & free parameters 0 free parameters · 4 assumptions · 0 invented entities

The paper adds no fitted parameters and no invented entities. Its load-bearing inputs are domain assumptions, two of which (Assumptions 1 and 2) the authors state explicitly in Section 3.4. The remaining assumptions are the generalisation theses in Section 4, which the paper presents as arguments rather than theorems.

assumptions (4)
  • domain assumption Assumption 1: Defining environments in DSLs or natural languages can be more user-friendly than defining them in general-purpose programming languages.
    Stated explicitly in Section 3.4. Supported by Ludii's design goals and a 1400+ game library count (footnotes 3 and 4), plus the authors' own experience; the paper admits no quantitative evidence in Section 6.
  • domain assumption Assumption 2: Enabling environments to be defined in more user-friendly ways is desirable.
    Normative premise stated in Section 3.4. Supported by surveys of game-industry engineers [58] and AI adoption studies [141] that discuss AI friction generally, not RL-specific usability data.
  • domain assumption Complete, compilable environment descriptions are a prerequisite for unrestricted zero-shot generalisation in RL.
    Section 4.3 argues this by analogy with human learning and cites [56]; the paper hedges it as 'could be argued' and 'likely', and offers no proof.
  • domain assumption The full action space cannot be reliably inferred from user-friendly DSLs such as Ludii.
    Section 3.4 argues this from the procedural semantics of Ludii's keywords and cites practical challenges (action aliasing in [143]). It motivates why description-based benchmarks expose new RL problems.

how reviews work

0 comments
Cite this review

Pith. "Pith review of A Research Agenda for Usability and Generalisation in Reinforcement Learning." pith.science (2026). https://pith.science/paper/3L2STC2L

@misc{pith2026241216970,
  author       = {Pith},
  title        = {Pith review of: A Research Agenda for Usability and Generalisation in Reinforcement Learning},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/3L2STC2L}},
  note         = {Machine review of arXiv:2412.16970}
}
read the original abstract

It is common practice in reinforcement learning (RL) research to train and deploy agents in bespoke simulators, typically implemented by engineers directly in general-purpose programming languages or hardware acceleration frameworks such as CUDA or JAX. This means that programming and engineering expertise is not only required to develop RL algorithms, but is also required to use already developed algorithms for novel problems. The latter poses a problem in terms of the usability of RL, in particular for private individuals and small organisations without substantial engineering expertise. We also perceive this as a challenge for effective generalisation in RL, in the sense that is no standard, shared formalism in which different problems are represented. As we typically have no consistent representation through which to provide information about any novel problem to an agent, our agents also cannot instantly or rapidly generalise to novel problems. In this position paper, we advocate for a research agenda centred around the use of user-friendly description languages for describing problems, such that (i) users with little to no engineering expertise can formally describe the problems they would like to be tackled by RL algorithms, and (ii) algorithms can leverage problem descriptions to effectively generalise among all problems describable in the language of choice.

Figures

Figures reproduced from arXiv: 2412.16970 by the authors.

Figure 1
Figure 1. Orange boxes with dashed lines represent components that require sub [PITH_FULL_IMAGE:figures/full_fig_p003_1.png] view at source ↗
Figure 2
Figure 2. A set of 1059 different board games, described in Ludii’s game description [PITH_FULL_IMAGE:figures/full_fig_p017_2.png] view at source ↗

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

178 extracted references · 56 canonical work pages

  1. [1]

    In: 2019 ICML Workshop on Human in the Loop Learning (2019)

    Abid, A., Abdalla, A., Abid, A., Khan, D., Alfozan, A., Zou, J.: Gradio: Hassle- free sharing and testing of ML models in the wild. In: 2019 ICML Workshop on Human in the Loop Learning (2019)

  2. [2]

    In: AAAI 2024 Workshop on Synergy of Reinforcement Learning and Large Language Models (2024)

    Afshar, A., Li, W.: DeLF: Designing learning environments with foundation mod- els. In: AAAI 2024 Workshop on Synergy of Reinforcement Learning and Large Language Models (2024)

  3. [3]

    In: Ranzato, M., Beygelzimer,A.,Dauphin,Y.,Liang,P.,Vaughan, J.W.(eds.)Advances inNeural InformationProcessingSystems.vol.34,pp.29304–29320.CurranAssociates,Inc

    Agarwal, R., Schwarzer, M., Castro, P.S., Courville, A., Bellemare, M.G.: Deep reinforcement learning at the edge of the statistical precipice. In: Ranzato, M., Beygelzimer,A.,Dauphin,Y.,Liang,P.,Vaughan, J.W.(eds.)Advances inNeural InformationProcessingSystems.vol.34,pp.29304–29320.CurranAssociates,Inc. (2021)

  4. [4]

    In: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A

    Agarwal, R., Schwarzer, M., Castro, P.S., Courville, A.C., Bellemare, M.: Reincar- nating reinforcement learning: Reusing prior computation to accelerate progress. In: Koyejo, S., Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A. (eds.) Advances in Neural Information Processing Systems. vol. 35, pp. 28955–28971. Curran Associates, Inc. (2022)

  5. [5]

    In: IEEE ICRA 2024 Workshop on Vision-Language Models for Navigation and Ma- nipulation (2024) 20 D.J.N.J

    Ahn, M., Dwibedi, D., Finn, C., Gonzalez Arenas, M., Gopalakrishnan, K., Haus- man, K., Ichter, B., Irpan, A., Joshi, N., Julian, R., Kirmani, S., Leal, I., Lee, E., Levine, S., Lu, Y., Maddineni, S., Rao, K., Sadigh, D., Sanketi, P., Sermanet, P., Vuong, Q., Welker, S., Xia, F., Xiao, T., Xu, P., Xu, S., Xu, Z.: AutoRT: Embodied foundation models for lar...

  6. [6]

    In: Proceedings of the 34th International Conference on Machine Learning

    Andreas, J., Klein, D., Levine, S.: Modular multitask reinforcement learning with policy sketches. In: Proceedings of the 34th International Conference on Machine Learning. vol. 70, pp. 166–175. PMLR (2017)

  7. [7]

    In: 2021 International Conference on Learning Representations (2021)

    Andrychowicz, M., Raichuk, A., Stań’czyk, P., Orsini, M., Girgin, S., Marinier, R., Hussenot, L., Geist, M., Pietquin, O., Michalski, M., Gelly, S., Bachem, O.: What matters for on-policy deep actor-critic methods? a large-scale study. In: 2021 International Conference on Learning Representations (2021)

  8. [8]

    Journal of Internet Services and Applications6(13) (2015)

    Aram, M., Neumann, G.: Multilayered analysis of co-development of business information systems. Journal of Internet Services and Applications6(13) (2015)

Show all 178 references
  1. [9]

    In: AAAI-21 Workshop on Reinforcement Learning in Games (2021)

    Bamford, C., Huang, S., Lucas, S.: Griddly: A platform for AI research in games. In: AAAI-21 Workshop on Reinforcement Learning in Games (2021)

  2. [10]

    In: The 20th International Joint Conference on Artificial Intelligence

    Banerjee, B., Stone, P.: General game learning using knowledge transfer. In: The 20th International Joint Conference on Artificial Intelligence. pp. 672–677 (2007)

  3. [11]

    https://arxiv.org/abs/2301.08028 (2023)

    Beck, J., Vuorio, R., Liu, E.Z., Xiong, Z., Zintgraf, L., Finn, C., Whiteson, S.: A survey of meta-reinforcement learning. https://arxiv.org/abs/2301.08028 (2023)

  4. [12]

    Nature588, 77–82 (2020)

    Bellemare, M.G., Candido, S., Castro, P.S., Gong, J., Machado, M.C., Moitra, S., Ponda, S.S., Wang, Z.: Autonomous navigation of stratospheric balloons using reinforcement learning. Nature588, 77–82 (2020)

  5. [13]

    Journal of Artificial Intelli- gence Research 47(1), 253–279 (2013)

    Bellemare, M.G., Naddaf, Y., Veness, J., Bowling, M.: The arcade learning envi- ronment: An evaluation platform for general agents. Journal of Artificial Intelli- gence Research 47(1), 253–279 (2013)

  6. [14]

    Transactions on Machine Learning Research (2023)

    Benjamins, C., Eimer, T., Schubert, F., Mohan, A., Döhler, S., Biedenkapp, A., Rosenhahn, B., Hutter, F., Lindauer, M.: Contextualize me – the case for context in reinforcement learning. Transactions on Machine Learning Research (2023)

  7. [15]

    In: Proceedings of the 16th International Symposium on Distributed Autonomous Robotic Systems

    Bettini, M., Kortvelesy, R., Blumenkamp, J., Prorok, A.: VMAS: A vectorized multi-agent simulator for collective robot learning. In: Proceedings of the 16th International Symposium on Distributed Autonomous Robotic Systems. DARS ’22, Springer (2022)

  8. [16]

    Voleti, Z.E., Letts, A., Jampani, V., Rombach, R.: Stable video diffusion: Scaling latent video diffusion models to large datasets

    Blattmann, A., Dockhorn, T., Kulal, S., Mendelevitch, D., Kilian, M., Lorenz, D., Levi, Y., anda V. Voleti, Z.E., Letts, A., Jampani, V., Rombach, R.: Stable video diffusion: Scaling latent video diffusion models to large datasets. https: //arxiv.org/abs/2311.15127 (2023)

  9. [17]

    In: Proceedings of the 42nd International Conference on Machine Learning (2025), to appear

    Blili-Hamelin, B., Graziul, C., Hancox-Li, L., Hazan, H., El-Mhamdi, E.M., Ghosh, A., Heller, K., Metcalf, J., Murai, F., Salvaggio, E., Smart, A., Snider, T., Tighanimine, M., Ringer, T., Mitchell, M., Dori-Hacohen, S.: Position: Stop treating ‘AGI’ as the north-star goal of ...

  10. [18]

    In: Proceedings of the International Conference on Learning Represen- tations (2024)

    Bonnet, C., Luo, D., Byrne, D., Surana, S., Abramowitz, S., Duckworth, P., Coyette, V., Midgley, L.I., Tegegn, E., Kalloniatis, T., Mahjoub, O., Macfarlane, M., Smit, A.P., Grinsztajn, N., Bolge, R., Waters, C.N., Mimouni, M.A., Sob, U.A.M., de Kock, R., Singh, S., Furelos-Bla...

  11. [19]

    In: Xing, E.P., Jebara, T

    Bou Ammar, H., Eaton, E., Ruvolo, P., Taylor, M.E.: Online multi-task learn- ing for policy gradient methods. In: Xing, E.P., Jebara, T. (eds.) Proceedings of the 31st International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 32, pp. 1206–1214 (2014)

  12. [20]

    com/google/jax

    Bradbury, J., Frostig, R., Hawkins, P., Johnson, M.J., Leary, C., Maclaurin, D., Necula, G., Paszke, A., VanderPlas, J., Wanderman-Milne, S., Zhang, Q.: JAX: A Research Agenda for Usability and Generalisation in RL 21 composable transformations of Python+NumPy programs (2018),...

  13. [21]

    Journal of Artificial Intelligence Research43, 661–704 (2012)

    Branavan, S.R.K., Silver, D., Barzilay, R.: Learning to win by reading manuals in a Monte-Carlo framework. Journal of Artificial Intelligence Research43, 661–704 (2012)

  14. [22]

    https://arxiv.org/abs/1606.01540 (2016)

    Brockman, G., Cheung, V., Pettersson, L., Schneider, J., Schulman, J., Tang, J., Zaremba, W.: OpenAI gym. https://arxiv.org/abs/1606.01540 (2016)

  15. [23]

    ludii.games/downloads/LudiiLanguageReference.pdf (2020)

    Browne, C., Soemers, D.J.N.J., Piette, É., Stephenson, M., Crist, W.: Ludii lan- guage reference. ludii.games/downloads/LudiiLanguageReference.pdf (2020)

  16. [24]

    Phd thesis, Faculty of Information Technology, Queensland University of Tech- nology, Queensland, Australia (2009)

    Browne, C.B.: Automatic Generation and Evaluation of Recombination Games. Phd thesis, Faculty of Information Technology, Queensland University of Tech- nology, Queensland, Australia (2009)

  17. [25]

    In: Daumé III, H., Singh, A

    Cobbe, K., Hesse, C., Hilton, J., Schulman, J.: Leveraging procedural generation to benchmark reinforcement learning. In: Daumé III, H., Singh, A. (eds.) Pro- ceedings of the 37th International Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 119,...

  18. [26]

    In: Chaudhuri, K., Salakhutdinov, R

    Cobbe, K., Klimov, O., Hesse, C., Kim, T., Schulman, J.: Quantifying gener- alization in reinforcement learning. In: Chaudhuri, K., Salakhutdinov, R. (eds.) Proceedings of the 36th International Conference on Machine Learning. Proceed- ings of Machine Learning Research, vol. 9...

  19. [27]

    https://arxiv.org/abs/2109

    Cummins, C., Wasti, B., Guo, J., Cui, B., Ansel, J., Gomez, S., Jain, S., Liu, J., Teytaud, O., Steiner, B., Tian, Y., Leather, H.: Compilergym: Robust, performant compiler optimization environments for ai research. https://arxiv.org/abs/2109. 08267 (2021)

  20. [28]

    In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H

    Dalton, S., Frosio, I.: Accelerating reinforcement learning through GPU Atari emulation. In: Larochelle, H., Ranzato, M., Hadsell, R., Balcan, M., Lin, H. (eds.) Advances in Neural Information Processing Systems. vol. 33, pp. 19773–19782. Curran Associates, Inc. (2020)

  21. [29]

    In: Finding the Frame workshop @ Reinforcement Learning Con- ference (2024)

    Davidson, G., Gureckis, T.M.: Toward complex and structured goals in reinforce- ment learning. In: Finding the Frame workshop @ Reinforcement Learning Con- ference (2024)

  22. [30]

    https://arxiv.org/abs/2405.13242 (2024)

    Davidson,G.,Todd,G.,Togelius,J.,Gureckis,T.M.,Lake,B.M.:Goalsasreward- producing programs. https://arxiv.org/abs/2405.13242 (2024)

  23. [31]

    Nature602, 414–419 (2022)

    Degrave, J., Felici, F., Buchli, J., Neunert, M., Tracey, B., Carpanese, F., Ewalds, T., Hafner, R., Abdolmaleki, A., de las Casas, D., Donner, C., Fritz, L., Galperti, C., Huber, A., Keeling, J., Tsimpoukelli, M., Kay, J., Merle, A., Moret, J.M., Noury, S., Pesamosca, F., Pfa...

  24. [32]

    In: Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA)

    Deisenroth, M.P., Englert, P., Peters, J., Fox, D.: Multi-task policy search for robotics. In: Proceedings of the 2014 IEEE International Conference on Robotics and Automation (ICRA). pp. 3876–3881 (2014)

  25. [33]

    In: Advances in Neural Information Processing Systems

    Dennis, M., Jaques, N., Vinitsky, E., Bayen, A., Russell, S., Critch, A., Levine, S.: Emergent complexity and zero-shot transfer via unsupervised environment design. In: Advances in Neural Information Processing Systems. vol. 33, pp. 13049–13061 (2020)

  26. [34]

    In: Proceedings of the 38th International Conference on Software Engineering

    Desai, A., Gulwani, S., Hingorani, V., Jain, N., Karkare, A., Marron, M., R, S., Roy, S.: Program synthesis using natural language. In: Proceedings of the 38th International Conference on Software Engineering. p. 345–356. Association for Computing Machinery (2016) 22 D.J.N.J. ...

  27. [35]

    In: Pro- ceedings of the Thirtieth International Joint Conference on Artificial Intelligence

    Eimer, T., Biedenkapp, A., Reimer, M., Adriaensen, S., Hutter, F., Lindauer, M.: DACBench: A benchmark library for dynamic algorithm configuration. In: Pro- ceedings of the Thirtieth International Joint Conference on Artificial Intelligence. pp. 1668–1674 (2021)

  28. [36]

    In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J

    Eimer, T., Lindauer, M., Raileanu, R.: Hyperparameters in reinforcement learning and how to tune them. In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J. (eds.) Proceedings of the 40th International Conference on Machine Learning. Proceedings of M...

  29. [37]

    In: Advances in Neural Information Processing Systems (2023), accepted

    Ellis, B., Cook, J., Moalla, S., Samvelyan, M., Sun, M., Mahajan, A., Foerster, J.N.,Whiteson,S.:SMACv2:Animprovedbenchmarkforcooperativemulti-agent reinforcement learning. In: Advances in Neural Information Processing Systems (2023), accepted

  30. [38]

    https://arxiv.org/abs/1810.00123 (2018)

    Farebrother, J., Machado, M.C., Bowling, M.: Generalization and regularization in DQN. https://arxiv.org/abs/1810.00123 (2018)

  31. [39]

    Progress in AI2(1), 13–27 (2013)

    Fernandéz, F., Veloso, M.: Learning domain structure through probabilistic policy reuse in reinforcement learning. Progress in AI2(1), 13–27 (2013)

  32. [40]

    In: Precup, D., Teh, Y.W

    Finn, C., Abeel, P., Levine, S.: Model-agnostic meta-learning for fast adaptation of deep networks. In: Precup, D., Teh, Y.W. (eds.) Proceedings of the 34th In- ternational Conference on Machine Learning. Proceedings of Machine Learning Research, vol. 70, pp. 1126–1135 (2017)

  33. [41]

    Freeman, C.D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., Bachem, O.: Brax - a differentiable physics engine for large scale rigid body simulation (2021), http: //github.com/google/brax

  34. [42]

    In: ICAIF ’23: Proceedings of the Fourth ACM International Conference on AI in Finance

    Frey, S., Li, K., Nagy, P., Sapora, S., Lu, C., Zohren, S., Foerster, J., Calinescu, A.: JAX-LOB: A GPU-accelerated limit order book simulator to unlock large scale reinforcement learning for trading. In: ICAIF ’23: Proceedings of the Fourth ACM International Conference on AI ...

  35. [43]

    In: Proceedings of the 36th International Conference on Machine Learning

    Gamrian, S., Goldberg, Y.: Transfer learning for related reinforcement learning tasks via image-to-image translation. In: Proceedings of the 36th International Conference on Machine Learning. pp. 2063–2072 (2019)

  36. [44]

    Synthesis Lectures on Ar- tificial Intelligence and Machine Learning, Morgan & Claypool Publishers (2014)

    Genesereth, M., Thielscher, M.: General Game Playing. Synthesis Lectures on Ar- tificial Intelligence and Machine Learning, Morgan & Claypool Publishers (2014)

  37. [45]

    In: Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W

    Ghosh, D., Rahme, J., Kumar, A., Zhang, A., Adams, R.P., Levine, S.: Why gen- eralization in rl is difficult: Epistemic pomdps and implicit partial observability. In: Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W. (eds.) Advances in Neural Information Proc...

  38. [46]

    In: Brazilian Conference on Intelligent Systems (BRACIS)

    Glatt, R., da Silva, F.L., Costa, A.H.R.: Towards knowledge transfer in deep re- inforcement learning. In: Brazilian Conference on Intelligent Systems (BRACIS). pp. 91–96. IEEE (2016)

  39. [47]

    Ex- pert Systems with Applications156 (2020)

    Glatt, R., Silva, F.L.D., da Costa Bianchi, R.A., Costa, A.H.R.: DECAF: Deep case-based policy inference for knowledge transfer in reinforcement learning. Ex- pert Systems with Applications156 (2020)

  40. [48]

    Goldie, A.D., Lu, C., Jackson, M.T., Whiteson, S., Foerster, J.N.: Can learned optimization make reinforcement learning less difficult? In: AutoRL Workshop @ ICML 2024 (2024)

  41. [49]

    In: The Thirty-Fourth AAAI Conference on Artificial Intelligence

    Goldwaser, A., Thielscher, M.: Deep reinforcement learning for general game play- ing. In: The Thirty-Fourth AAAI Conference on Artificial Intelligence. pp. 1701–

  42. [50]

    In: Proceedings of the 2020 Conference on Robot Learning

    Goyal, P., Niekum, S., Mooney, R.J.: PixL2R: Guiding reinforcement learning using natural language by mapping pixels to rewards. In: Proceedings of the 2020 Conference on Robot Learning. PMLR, vol. 155, pp. 485–497 (2021)

  43. [51]

    In: Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks (2023)

    Gulino, C., Fu, J., Luo, W., Tucker, G., Bronstein, E., Lu, Y., Harb, J., Pan, X., Wang, Y., Chen, X., Co-Reyes, J.D., Agarwal, R., Roelofs, R., Lu, Y., Montali, N., Mougin, P., Yang, Z., B, W., Faust, A., McAllister, R., Anguelov, D., Sapp, B.: Waymax: An accelerated, data-dr...

  44. [52]

    In: Proceedings of the 42nd International Conference on Machine Learning (2025), to appear

    Hartman, S., Ong, C.S., Powles, J., Kuhnert, P.: Position: We need responsible, application-driven (RAD) AI research. In: Proceedings of the 42nd International Conference on Machine Learning (2025), to appear

  45. [53]

    In: Proceedings of the 32nd AAAI Confer- ence on Artificial Intelligence

    Henderson, P., Islam, R., Bachman, P., Pineau, J., Precup, D., Meger, D.: Deep reinforcement learning that matters. In: Proceedings of the 32nd AAAI Confer- ence on Artificial Intelligence. pp. 3207–3214. AAAI (2018)

  46. [54]

    Communications of the ACM64(12), 58–65 (2021)

    Hooker, S.: The hardware lottery. Communications of the ACM64(12), 58–65 (2021)

  47. [55]

    The MIT Press, Cambridge, Massachusetts (1960)

    Howard, R.A.: Dynamic Programming and Markov Processes. The MIT Press, Cambridge, Massachusetts (1960)

  48. [56]

    In: ICML 2019 Workshop on Understanding and Improving Generalization in Deep Learning (2019)

    Irpan, A., Song, X.: The principle of unchanged optimality in reinforcement learn- ing generalization. In: ICML 2019 Workshop on Understanding and Improving Generalization in Deep Learning (2019)

  49. [57]

    In: Advances in Neural Information Processing Systems

    Jackson, M.T., Jiang, M., Parker-Holder, J., Vuorio, R., Lu, C., Farquhar, G., Whiteson, S., Foerster, J.N.: Discovering general reinforcement learning algo- rithms with adversarial environment design. In: Advances in Neural Information Processing Systems. vol. 36, pp. 79980–7...

  50. [58]

    it’s unwieldy and it takes a lot of time

    Jacob, M., Devlin, S., Hofmann, K.: “it’s unwieldy and it takes a lot of time” — challengesandopportunitiesforcreatingagentsincommercialgames.In:Proceed- ings of the Sixteenth AAAI Conference on Artificial Intelligence and Interactive Digital Entertainment. pp. 88–94 (2020)

  51. [59]

    In: Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F

    Jordan, S.M., White, A., da Silva, B.C., White, M., Thomas, P.S.: Position: Benchmarking is limited in reinforcement learning research. In: Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F. (eds.) Proceedings of the 41st Internatio...

  52. [60]

    https://nips.cc/ virtual/2022/63891 (2022), opinion talk contributed to the Deep Reinforcement Learning Workshop at NeurIPS 2022

    Jordan, S.M.: Scientific experiments in reinforcement learning. https://nips.cc/ virtual/2022/63891 (2022), opinion talk contributed to the Deep Reinforcement Learning Workshop at NeurIPS 2022

  53. [61]

    In: Advances in Neural Information Processing Systems

    Jothimurugan, K., Alur, R., Bastani, O.: A composable specification language for reinforcement learning tasks. In: Advances in Neural Information Processing Systems. vol. 32, pp. 13041–13051 (2019)

  54. [62]

    In: NeurIPS 2018 Workshop on Deep Reinforcement Learning (2018)

    Justesen, N., Torrado, R.R., Bontrager, P., Khalifa, A., Togelius, J., Risi, S.: Illuminating generalization in deep reinforcement learning through procedural level generation. In: NeurIPS 2018 Workshop on Deep Reinforcement Learning (2018)

  55. [63]

    In: Proceedings of the AAAI Conference on Artificial Intelligence

    Kaiser, Ł., Stafiniak, Ł.: First-order logic with counting for general game playing. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 25, pp. 791–796 (2011) 24 D.J.N.J. Soemers et al

  56. [64]

    Transactions on Machine Learning Research (2025)

    Kaufmann, T., Weng, P., Bengs, V., Hüllermeier, E.: A survey of reinforce- ment learning from human feedback. Transactions on Machine Learning Research (2025)

  57. [65]

    https: //arxiv.org/abs/2401.02991 (2024)

    Kharyal, C., Krishna Gottipati, S., Kumar Sinha, T., Das, S., Taylor, M.E.: GLIDE-RL: Grounded language instruction through DEmonstration in RL. https: //arxiv.org/abs/2401.02991 (2024)

  58. [66]

    Journal of Artificial Intelligence Research 76, 201–264 (2023)

    Kirk, R., Zhang, A., Grefenstette, E., Rocktäschel, T.: A survey of zero-shot generalisation in deep reinforcement learning. Journal of Artificial Intelligence Research 76, 201–264 (2023)

  59. [67]

    In: Proceedings of the 2020 IEEE Conference on Games

    Kowalksi, J., Miernik, R., Mika, M., Pawlik, W., Sutowicz, J., Szykuła, M., Tkaczyk, A.: Efficient reasoning in regular boardgames. In: Proceedings of the 2020 IEEE Conference on Games. pp. 455–462. IEEE (2020)

  60. [68]

    In: Proceedings of the 33rd AAAI Conference on Artificial Intelligence

    Kowalski, J., Maksymilian, M., Sutowicz, J., Szykuła, M.: Regular boardgames. In: Proceedings of the 33rd AAAI Conference on Artificial Intelligence. vol. 33, pp. 1699–1706. AAAI Press (2019)

  61. [69]

    In: Advances in Neural Information Processing Systems (2023)

    Koyamada, S., Okano, S., Nishimori, S., Murata, Y., Habara, K., Kita, H., Ishii, S.:Pgx:Hardware-acceleratedparallelgamesimulatorsforreinforcementlearning. In: Advances in Neural Information Processing Systems (2023)

  62. [70]

    In: Kok, J., Koronacki, J., Mantaras, R., Matwin, S., Mladenič, D., Skowron, A

    Kuhlmann, G., Stone, P.: Graph-based domain mapping for transfer learning in general games. In: Kok, J., Koronacki, J., Mantaras, R., Matwin, S., Mladenič, D., Skowron, A. (eds.) Machine Learning: ECML 2007. Lecture Notes in Computer Science, vol. 4071, pp. 188–200. Springer, ...

  63. [71]

    Lange, R.T.: gymnax: A JAX-based reinforcement learning environment library (2022), http://github.com/RobertTLange/gymnax

  64. [72]

    In: Neural Networks

    Lange, S., Riedmiller, M.: Deep auto-encoder neural networks in reinforcement learning. In: Neural Networks. International Joint Conference. 2010. (IJCNN 2010). pp. 1623–1630. IEEE (2010)

  65. [73]

    In: Wiering, M., van Otterlo, M

    Lazaric, A.: Transfer in reinforcement learning: a framework and a survey. In: Wiering, M., van Otterlo, M. (eds.) Reinforcement Learning. Adaptation, Learn- ing, and Optimization, vol. 12, pp. 143–173. Springer, Berlin, Heidelberg (2012)

  66. [74]

    Nature521(7553), 436–444 (2015)

    LeCun, Y., Bengio, Y., Hinton, G.: Deep learning. Nature521(7553), 436–444 (2015)

  67. [75]

    https:// arxiv.org/abs/2306.14892 (2023)

    Lee, J.N., Xie, A., Pacchiano, A., Chandak, Y., Finn, C., Nachum, O., Brunskill, E.: Supervised pretraining can learn in-context reinforcement learning. https:// arxiv.org/abs/2306.14892 (2023)

  68. [76]

    Leivada, E., Marcus, G., Günther, F., Murphy, E.: A sentence is worth a thousand pictures: Can large language models understand hum4n l4ngu4ge and the w0rld behind w0rds? https://arxiv.org/abs/2308.00109 (2024)

  69. [77]

    Journal of Machine Learning Research10(40), 1131–1186 (2009)

    Li, H., Liao, X., Carin, L.: Multi-task reinforcement learning in partially ob- servable stochastic environments. Journal of Machine Learning Research10(40), 1131–1186 (2009)

  70. [78]

    In: Agmon, N., Taylor, M.E., Veloso, E.E.M

    Li, X., Zhang, J., Bian, J., Tong, Y., Liu, T.Y.: A cooperative multi-agent rein- forcement learning framework for resource balancing in complex logistics network. In: Agmon, N., Taylor, M.E., Veloso, E.E.M. (eds.) Proceedings of the 18th Inter- national Conference on Autonomo...

  71. [79]

    https://arxiv.org/abs/2306.00937 (2023)

    Lifschitz, S., Paster, K., Chan, H., Ba, J., McIlraith, S.: Steve-1: A generative model for text-to-behavior in Minecraft. https://arxiv.org/abs/2306.00937 (2023)

  72. [80]

    In: Proceedings of the International Conference on Automated Planning and Scheduling

    Lin, S., Bercher, P.: On the expressive power of planning formalisms in conjunc- tion with LTL. In: Proceedings of the International Conference on Automated Planning and Scheduling. vol. 32, pp. 231–240 (2022) A Research Agenda for Usability and Generalisation in RL 25

  73. [81]

    In: Proceedings of the 41st International Conference on Machine Learning

    Lindauer, M., Karl, F., Klier, A., Moosbauer, J., Tornede, A., Mueller, A., Hutter, F., Feurer, M., Bischl, B.: Position: A call to action for a human-centered Au- toML paradigm. In: Proceedings of the 41st International Conference on Machine Learning. PMLR, vol. 235, pp. 3056...

  74. [82]

    In: Proceedings of the Eleventh International Conference on Machine Learn- ing

    Littman, M.L.: Markov games as a framework for multi-agent reinforcement learn- ing. In: Proceedings of the Eleventh International Conference on Machine Learn- ing. pp. 157–163 (1994)

  75. [83]

    Love, N., Hinrichs, T., Haley, D., Schkufza, E., Genesereth, M.: General game playing: Game description language specification. Tech. Rep. LG-2006-01, Stan- ford Logic Group (2008)

  76. [84]

    In: Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligence, IJCAI-19

    Luketina, J., Nardelli, N., Farquhar, G., Foerster, J., Andreas, J., Grefenstette, E., Whiteson, S., Rocktä"schel, T.: A survey of reinforcement learning informed by natural language. In: Proceedings of the Twenty-Eighth International Joint Conference on Artificial Intelligenc...

  77. [85]

    In: Towards Generalist Robots: Learning Paradigms for Scalable Skill Acquisition @ CoRL2023 (2023)

    Luo, J., Hu, Z., Xu, C., Gadipudi, S., Sharma, A., Ahmad, R., Schaal, S., Finn, C., Gupta, A., Levine, S.: SERL: A software suite for sample-efficient robotic reinforcement learning. In: Towards Generalist Robots: Learning Paradigms for Scalable Skill Acquisition @ CoRL2023 (2023)

  78. [86]

    Journal of Machine Learning Research 9(86), 2579–2605 (2008)

    van der Maaten, L., Hinton, G.: Visualizing data using t-sne. Journal of Machine Learning Research 9(86), 2579–2605 (2008)

  79. [87]

    Journal of Artificial Intelligence Research61, 523–562 (2018)

    Machado, M.C., Bellemare, M.G., Talvitie, E., Veness, J., Hausknecht, M., Bowl- ing, M.: Revisiting the arcade learning environment: Evaluation protocols and open problems for general agents. Journal of Artificial Intelligence Research61, 523–562 (2018)

  80. [88]

    Machine Learning 22, 251–281 (1996)

    Maclin, R., Shavlik, J.W.: Creating advice-taking reinforcement learners. Machine Learning 22, 251–281 (1996)

  81. [89]

    In: Vanschoren, J., Yeung, S

    Makoviychuk, V., Wawrzyniak, L., Guo, Y., Lu, M., Storey, K., Macklin, M., Hoeller, D., Rudin, N., Allshire, A., Handa, A., State, G.: Isaac gym: High per- formance GPU based physics simulation for robot learning. In: Vanschoren, J., Yeung, S. (eds.) Proceedings of the Neural ...

  82. [90]

    Malik, D., Li, Y., Ravikumar, P.: When is generalizable reinforcement learning tractable? In: Ranzato, M., Beygelzimer, A., Dauphin, Y., Liang, P., Vaughan, J.W.(eds.)AdvancesinNeuralInformationProcessingSystems.vol.34,pp.8032–

  83. [91]

    https://arxiv.org/abs/2301.01320 (2023)

    Mannor, S., Tamar, A.: Towards deployable RL – what’s broken with RL research and a potential fix. https://arxiv.org/abs/2301.01320 (2023)

  84. [92]

    In: Proceedings of the AAAI Conference on Artificial Intelligence

    Maras, M., Kępa, M., Kowalski, J., Szykuła, M.: Fast and knowledge-free deep learning for general game playing (student abstract). In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 38, pp. 23576–23578 (2024)

  85. [93]

    McDermott, D., Ghallab, M., Howe, A., Knoblock, C., Ram, A., Veloso, M., Weld, D., Wilkins, D.: PDDL—the planning domain definition language. Tech. Rep. CVC TR98003/DCS TR1165, New Haven, CT: Yale Center for Computational Vision and Control (1998)

  86. [94]

    In: 2024 International Conference on Learning Represen- tations (2024)

    Mediratta, I., You, Q., Jiang, M., Raileanu, R.: A study of generalization in offline reinforcement learning. In: 2024 International Conference on Learning Represen- tations (2024)

  87. [95]

    ACM Computing Surveys37(4), 316–344 (2005) 26 D.J.N.J

    Mernik, M., Heering, J., Sloane, A.M.: When and how to develop domain-specific languages. ACM Computing Surveys37(4), 316–344 (2005) 26 D.J.N.J. Soemers et al

  88. [96]

    Nature594, 207–212 (2021)

    Mirhoseini, A., Goldie, A., Yazgan, M., Jiang, J.W., Songhori, E., Wang, S., Lee, Y.J., Johnson, E., Pathak, O., Nazi, A., Pak, J., Tong, A., Srinivasa, K., Hang, W., Tuncer, E., Le, Q.V., Laudon, J., Ho, R., Carpenter, R., Dean, J.: A graph placement methodology for fast chip...

  89. [97]

    In: Palmer, M., Hwa, R., Riedel, S

    Misra, D., Langford, J., Artzi, Y.: Mapping instructions and visual observations to actions with reinforcement learning. In: Palmer, M., Hwa, R., Riedel, S. (eds.) Proceedings of the 2017 Conference on Empirical Methods in Natural Language Processing. pp. 1004–1015 (2017)

  90. [98]

    In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW)

    Mittel, A., Munukutla, P.S.: Visual transfer between Atari games using competi- tive reinforcement learning. In: 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW). pp. 499–501 (2019)

  91. [99]

    https://arxiv

    Mnih, V., Kavukcuoglu, K., Silver, D., Graves, A., Antonoglou, I., Wierstra, D., Riedmiller, M.: Playing Atari with deep reinforcement learning. https://arxiv. org/abs/1312.5602 (2013)

  92. [100]

    In: Faust, A., Garnett, R., White, C., Hutter, F., Gardner, J.R

    Mohan, A., Benjamins, C., Wienecke, K., Dockhorn, A., Lindauer, M.: AutoRL hyperparameter landscapes. In: Faust, A., Garnett, R., White, C., Hutter, F., Gardner, J.R. (eds.) International Conference on Automated Machine Learning. Proceedings of Machine Learning Research, vol. ...

  93. [101]

    Journal of Artificial Intelligence Research79, 1167– 1236 (2024)

    Mohan, A., Zhang, A., Lindauer, M.: Structure in deep reinforcement learning: A survey and open problems. Journal of Artificial Intelligence Research79, 1167– 1236 (2024)

  94. [102]

    Müller-Brockhausen, M., Preuss, M., Plaat, A.: Procedural content generation: Betterbenchmarksfortransferreinforcementlearning.In:Proceedingsofthe2021 IEEE Conference on Games. pp. 924–931 (2021)

  95. [103]

    In: Proceedings of the 2019 International Conference on Learning Representations (2019)

    Nagabandi, A., Clavera, I., Liu, S., Fearing, R.S., Abbeel, P., Levine, S., Finn, C.: Learning to adapt in dynamic, real-world environments through meta- reinforcement learning. In: Proceedings of the 2019 International Conference on Learning Representations (2019)

  96. [104]

    https://arxiv.org/abs/1804.03720 (2018)

    Nichol, A., Pfau, V., Hesse, C., Klimov, O., Schulman, J.: Gotta learn fast: A new benchmark for generalization in RL. https://arxiv.org/abs/1804.03720 (2018)

  97. [105]

    In: Reinforcement Learning Conference (2024), accepted

    Obando-Ceron, J., Araú’jo, J.G.M., Courville, A., Castro, P.S.: On the consis- tency of hyper-parameter selection in value-based deep reinforcement learning. In: Reinforcement Learning Conference (2024), accepted

  98. [106]

    In: Meila, M., Zhang, T

    Obando-Ceron, J.S., Castro, P.S.: Revisiting rainbow: Promoting more insightful and inclusive deep reinforcement learning research. In: Meila, M., Zhang, T. (eds.) Proceedings of the 38th International Conference on Machine Learning. pp. 1373–

  99. [107]

    In: Proceedings of the 34th International Conference on Machine Learning

    Oh, J., Singh, S., Lee, H., Kohli, P.: Zero-shot task generalization with multi-task deep reinforcement learning. In: Proceedings of the 34th International Conference on Machine Learning. pp. 2661–2670. PMLR (2017)

  100. [108]

    https://openai.com/blog/chatgpt (2022), ac- cessed: 2024-01-02

    OpenAI: Introducing ChatGPT. https://openai.com/blog/chatgpt (2022), ac- cessed: 2024-01-02

  101. [109]

    In: Proceedings of the International Con- ference on Automated Planning and Scheduling

    Oswald, J., Srinivas, K., Kokel, H., Lee, J., Katz, M., Sohrabi, S.: Large language models as planning domain generators. In: Proceedings of the International Con- ference on Automated Planning and Scheduling. vol. 34, pp. 423–431 (2024)

  102. [110]

    In: Koyejo, S., A Research Agenda for Usability and Generalisation in RL 27 Mohamed, S., Agarwal, A., Belgrave, D., Cho, K., Oh, A

    Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., Schulman, J., Hilton, J., Kelton, F., Miller, L., Simens, M., Askell, A., Welinder, P., Christiano, P., Leike, J., Lowe, R.: Train- ing language models to follo...

  103. [111]

    In: Proceedings of the 39th International Conference on Machine Learning

    Parker-Holder, J., Jiang, M., Dennis, M., Samvelyan, M., Foerster, J., Grefen- stette, E., Rocktäschel, T.: Evolving curricula with regret-based environment de- sign. In: Proceedings of the 39th International Conference on Machine Learning. PMLR, vol. 162, pp. 17473–17498 (2022)

  104. [112]

    Journal of Artificial Intelligence Research74, 517–568 (2022)

    Parker-Holder, J., Rajan, R., Song, X., Biedenkapp, A., Miao, Y., Eimer, T., Zhang, B., Nguyen, V., Calandra, R., Faust, A., Hutter, F., Lindauer, M.: Auto- mated reinforcement learning (autoRL): A survey and open problems. Journal of Artificial Intelligence Research74, 517–568 (2022)

  105. [113]

    https://arxiv.org/abs/2304.01315 (2023)

    Patterson, A., Neumann, S., White, M., White, A.: Empirical design in reinforce- ment learning. https://arxiv.org/abs/2304.01315 (2023)

  106. [114]

    Stratega

    Perez-Liebana, D., Dockhorn, A., Grueso, J.H., Jeurissen, D.: The design of "Stratega": A general strategy games framework. In: Osborn, J.C. (ed.) Joint Pro- ceedings of the AIIDE 2020 Workshops co-located with 16th AAAI Conference on Artificial Intelligence and Interactive Di...

  107. [115]

    In: Giacomo, G.D., Catala, A., Dilkina, B., Milano, M., Barro, S., Bugarín, A., Lang, J

    Piette, É., Soemers, D.J.N.J., Stephenson, M., Sironi, C.F., Winands, M.H.M., Browne, C.: Ludii – the ludemic general game system. In: Giacomo, G.D., Catala, A., Dilkina, B., Milano, M., Barro, S., Bugarín, A., Lang, J. (eds.) Proceedings of the 24th European Conference on Art...

  108. [116]

    In: Proceedings of the 2021 IEEE Conference on Games (CoG)

    Piette, É., Stephenson, M., Soemers, D.J.N.J., Browne, C.: General board game concepts. In: Proceedings of the 2021 IEEE Conference on Games (CoG). pp. 932–939. IEEE (2021)

  109. [117]

    https://arxiv.org/abs/2307.01952 (2023)

    Podell, D., English, Z., Lacey, K.,Blattmann, A., Dockhorn,T., Müller, J., Penna, J., Rombach, R.: SDXL: Improving latent diffusion models for high-resolution image synthesis. https://arxiv.org/abs/2307.01952 (2023)

  110. [118]

    In: Reinforcement Learning Conference (2025), to appear

    Ponse, K., Kleuker, J.F., Moerland, T.M., Plaat, A.: Chargax: A JAX accelerated EV charging simulator. In: Reinforcement Learning Conference (2025), to appear

  111. [119]

    In: Chaudhuri, K., Salakhutdinov, R

    Rakelly, K., Zhou, A., Quillen, D., Finn, C., Levine, S.: Efficient off-policy meta- reinforcement learning via probabilistic context variables. In: Chaudhuri, K., Salakhutdinov, R. (eds.) Proceedings of the 36th International Conference on Machine Learning. Proceedings of Mac...

  112. [120]

    https://arxiv.org/abs/2204

    Ramesh, A., Dhariwal, P., Nichol, A., Chu, C., Chen, M.: Hierarchical text- conditional image generation with CLIP latents. https://arxiv.org/abs/2204. 06125 (2022)

  113. [121]

    rewards: A comparative study of ob- jective specification mechanisms

    Rani, S., Booth, S., Sreedharan, S.: Goals vs. rewards: A comparative study of ob- jective specification mechanisms. In: Reinforcement Learning Conference (2025), to appear

  114. [122]

    https://arxiv

    Raparthy, S.C., Hambro, E., Kirk, R., Henaff, M., Raileanu, R.: Generalization to new sequential decision making tasks with in-context learning. https://arxiv. org/abs/2312.03801 (2023)

  115. [123]

    Transactions on Machine Learning Research (2023) 28 D.J.N.J

    Reed, S., Żołna, K., Parisotto, E., Colmenarejo, S.G., Novikov, A., Barth-Maron, G., Giménez, M., Sulsky, Y., Kay, J., Springenberg, J.T., Eccles, T., Bruce, J., Razavi, A., Edwards, A., Heess, N., Chen, Y., Hadsell, R., Vinyals, O., Bordbar, M., de Freitas, N.: A generalist a...

  116. [124]

    In: Machine Learning: ECML 2005

    Riedmiller, M.: Neural fitted Q iteration - first experiences with a data efficient neural reinforcement learning method. In: Machine Learning: ECML 2005. Lec- ture Notes in Computer Science, vol. 3720, pp. 317–328. Springer (2005)

  117. [125]

    In: International Conference on Learning Representations (2024)

    Rigter, M., Jiang, M., Posner, I.: Reward-free curricula for training robust world models. In: International Conference on Learning Representations (2024)

  118. [126]

    In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J

    Rodriguez-Sanchez, R., Spiegel, B.A., Wang, J., Patel, R., Tellex, S., Konidaris, G.: RLang: A declarative language for describing partial world knowledge to rein- forcement learning agents. In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J. (eds....

  119. [127]

    In: Proceedings of the 41st International Conference on Machine Learning

    Rolnick, D., Aspuru-Guzik, A., Beery, S., Dilkina, B., Donti, P.L., Ghassemi, M., Kerner, H., Monteleoni, C., Rolf, E., Tambe, M., White, A.: Position: Application- driven innovation in machine learning. In: Proceedings of the 41st International Conference on Machine Learning....

  120. [128]

    Journal of Artificial Intelligence Research 67, 673–703 (2020)

    Rostami, M., Isele, D., Eaton, E.: Using task descriptions in lifelong machine learning for improved performance and zero-shot transfer. Journal of Artificial Intelligence Research 67, 673–703 (2020)

  121. [129]

    https: //arxiv.org/abs/1606.04671 (2016)

    Rusu, A.A., Rabinowitz, N.C., Desjardins, G., Soyer, H., Kirkpatrick, J., Kavukcuoglu, K., Pascanu, R., Hadsell, R.: Progressive neural networks. https: //arxiv.org/abs/1606.04671 (2016)

  122. [130]

    https://arxiv.org/abs/2311.10090 (2023)

    Rutherford, A., Ellis, B., Gallici, M., Cook, J., Lupu, A., Ingvarsson, G., Willi, T., Khan, A., de Witt, C.S., Souly, A., Bandyopadhyay, S., Samvelyan, M., Jiang, M., Lange, R.T., Whiteson, S., Lacerda, B., Hawes, N., Rocktäschel, T., Lu, C., Foerster, J.N.: JaxMARL: Multi-ag...

  123. [131]

    In: Proceedings of IEEE International Conference on Robotics and Automation (ICRA) (2025), to appear

    Sakçak, B., Shell, D.A., O’Kane, J.M.: Limits of specifiability for sensor-based robotic planning tasks. In: Proceedings of IEEE International Conference on Robotics and Automation (ICRA) (2025), to appear

  124. [132]

    In: International Conference on Learning Representations (2023)

    Samvelyan, M., Khan, A., Dennis, M., Jiang, M., Parker-Holder, J., Foerster, J., Raileanu, R., Rocktäschel, T.: MAESTRO: Open-ended environment design for multi-agent reinforcement learning. In: International Conference on Learning Representations (2023)

  125. [133]

    In: Advances in Neural Information Processing Systems (2021)

    Samvelyan, M., Kirk, R., Kurin, V., Parker-Holder, J., Jiang, M., Hambro, E., Petroni, F., Küttler, H., Grefenstette, E., Rocktäschel, T.: Minihack the planet: A sandbox for open-ended reinforcement learning research. In: Advances in Neural Information Processing Systems (2021)

  126. [134]

    In: Proceedings of the IEEE Conference on Computational Intelligence in Games

    Schaul, T.: A video game description language for model-based or interactive learning. In: Proceedings of the IEEE Conference on Computational Intelligence in Games. pp. 193–200. IEEE (2013)

  127. [135]

    In: Proceedings of the 32nd International Conference on Machine Learning

    Schaul,T.,Horgan,D.,Gregor,K.,Silver,D.:Universalvaluefunctionapproxima- tors. In: Proceedings of the 32nd International Conference on Machine Learning. JLMR: W&CP, vol. 37, pp. 1312–1320 (2015)

  128. [136]

    IEEE Transac- tions on Computational Intelligence and AI in Games6(4), 325–331 (Dec 2014)

    Schaul, T.: An extensible description language for video games. IEEE Transac- tions on Computational Intelligence and AI in Games6(4), 325–331 (Dec 2014). https://doi.org/10.1109/TCIAIG.2014.2352795

  129. [137]

    Schmidhuber, J.: On learning how to learn learning strategies. Tech. Rep. FKI- 198-94, Institut für Informatik, Technische Universität München (1994)

  130. [138]

    In: Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society

    Seger, E., Ovadya, A., Siddarth, D., Garfinkel, B., Dafoe, A.: Democratising AI: Multiple meanings, goals, and methods. In: Proceedings of the 2023 AAAI/ACM Conference on AI, Ethics, and Society. pp. 715–722 (2023) A Research Agenda for Usability and Generalisation in RL 29

  131. [139]

    In: International Conference on Learning Representations (2018)

    Shu, T., Xiong, C., Socher, R.: Hierarchical and interpretable skill acquisition in multi-task reinforcement learning. In: International Conference on Learning Representations (2018)

  132. [140]

    In: ICAPS Workshop on Bridging the Gap Between AI Planning and Reinforcement Learning (PRL) (2020)

    Silver, T., Chitnis, R.: PDDLGym: Gym environments from PDDL problems. In: ICAPS Workshop on Bridging the Gap Between AI Planning and Reinforcement Learning (PRL) (2020)

  133. [141]

    it is there, and you need it, so why do you not use it?

    Simkute, A., Luger, E., Evans, M., Jones, R.: “it is there, and you need it, so why do you not use it?” achieving better adoption of AI systems by domain experts, in the case study of natural science research. https://arxiv.org/abs/2403.16895 (2024)

  134. [142]

    https://arxiv.org/abs/1807.11074 (2018)

    Sobol, D., Wolf, L., Taigman, Y.: Visual analogies between atari games for study- ing transfer learning in rl. https://arxiv.org/abs/1807.11074 (2018)

  135. [143]

    ICGA Journal43(3), 146–161 (2022)

    Soemers, D.J.N.J., Mella, V., Browne, C., Teytaud, O.: Deep learning for general game playing with Ludii and Polygames. ICGA Journal43(3), 146–161 (2022)

  136. [144]

    Transactions on Machine Learning Research (2023)

    Soemers, D.J.N.J., Mella, V., Piette, É., Stephenson, M., Browne, C., Teytaud, O.: Towards a general transfer approach for policy-value networks. Transactions on Machine Learning Research (2023)

  137. [145]

    In: Proceedings of the 2024 IEEE Conference on Games

    Soemers, D.J.N.J., Piette, É., Stephenson, M., Browne, C.: The Ludii game de- scription language is universal. In: Proceedings of the 2024 IEEE Conference on Games. pp. 1–8 (2024)

  138. [146]

    In: Rocha, A.P., Steels, L., van den Herik, H.J

    Soemers, D.J.N.J., Samothrakis, S., Driessens, K., Winands, M.H.M.: Environ- ment descriptions for usability and generalisation in reinforcement learning. In: Rocha, A.P., Steels, L., van den Herik, H.J. (eds.) Proceedings of the 17th In- ternational Conference on Agents and A...

  139. [147]

    https://stability.ai/research/stable-audio-efficient-timing-latent-diffusion (2023), accessed: 2024-1-4

    Stability AI: Stable audio: Fast timing-conditioned latent audio diffusion. https://stability.ai/research/stable-audio-efficient-timing-latent-diffusion (2023), accessed: 2024-1-4

  140. [148]

    In: Browne, C., Kishimoto, A., Schaeffer, J

    Stephenson, M., Soemers, D.J.N.J., Piette, É., Browne, C.: Measuring board game distance. In: Browne, C., Kishimoto, A., Schaeffer, J. (eds.) Computers and Games. CG 2022. Lecture Notes in Computer Science, vol. 13865, pp. 121–130. Springer, Cham (2023)

  141. [149]

    https: //arxiv.org/abs/2101.02722 (2021)

    Stone, A., Ramirez, O., Konolige, K., Jonschkowski, R.: The distracting control suite – a challenging benchmark for reinforcement learning from pixels. https: //arxiv.org/abs/2101.02722 (2021)

  142. [150]

    In: International Confer- ence on Learning Representations (2020)

    Sun, S.H., Wu, T.L., Lim, J.J.: Program guided agent. In: International Confer- ence on Learning Representations (2020)

  143. [151]

    MIT Press, Cambridge, MA, 2 edn

    Sutton, R.S., Barto, A.G.: Reinforcement Learning: An Introduction. MIT Press, Cambridge, MA, 2 edn. (2018)

  144. [152]

    Tassa, Y., Doron, Y., Muldal, A., Erez, T., Li, Y., de Las Casas, D., Budden, D., Abdolmaleki, A., Merel, J., Lefrancq, A., Lillicrap, T., Riedmiller, M.: DeepMind control suite (2018)

  145. [153]

    In: Mahadevan, S

    Taylor, M.E., Stone, P.: Transfer learning for reinforcement learning domains: A survey. In: Mahadevan, S. (ed.) Journal of Machine Learning Research. vol. 10, pp. 1633–1685 (2009)

  146. [154]

    In: Advances in Neural Information Processing Systems

    Terry, J.K., Black, B., Grammel, N., Jayakumar, M., Hari, A., Sullivan, R., San- tos,L.,Perez,R.,Horsch,C.,Dieffendahl,C.,Williams,N.L.,Lokesh,Y.,Ravi,P.: Pettingzoo: A standard API for multi-agent reinforcement learning. In: Advances in Neural Information Processing Systems. ...

  147. [155]

    In: Proceedings of the AAAI Conference on Artificial Intelligence

    Tessler, C., Givony, S., Zahavy, T., Mankowitz, D., Mannor, S.: A deep hierar- chical approach to lifelong learning in minecraft. In: Proceedings of the AAAI Conference on Artificial Intelligence. pp. 1553–1561. AAAI (2017)

  148. [156]

    In: Proceedings of the Twenty-second International Joint Conference on Artificial Intelligence, IJCAI-11

    Thielscher, M.: The general game playing description language is universal. In: Proceedings of the Twenty-second International Joint Conference on Artificial Intelligence, IJCAI-11. pp. 1107–1112 (2011)

  149. [157]

    In: Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C

    Todd, G., Padula, A., Stephenson, M., Piette, É., Soemers, D.J.N.J., Togelius, J.: GAVEL: Generating games via evolution and language models. In: Globerson, A., Mackey, L., Belgrave, D., Fan, A., Paquet, U., Tomczak, J., Zhang, C. (eds.) Advances in Neural Information Processi...

  150. [158]

    https://arxiv.org/abs/ 2506.22609 (2025)

    Todd, G., Padula, A.G., Soemers, D.J.N.J., Togelius, J.: Ludax: A GPU- accelerated domain specific language for board games. https://arxiv.org/abs/ 2506.22609 (2025)

  151. [159]

    https://arxiv.org/ abs/2101.04808 (2021)

    Trofin, M., Qian, Y., Brevdo, E., Lin, Z., Choromanski, K., Li, D.: MLGO: a machine learning guided compiler optimizations framework. https://arxiv.org/ abs/2101.04808 (2021)

  152. [160]

    In: Finding the Frame workshop @ Reinforcement Learning Conference (2024)

    Voelcker, C., Hussing, M., Eaton, E.: Can we hop in general? A discussion of benchmark selection and design using the Hopper environment. In: Finding the Frame workshop @ Reinforcement Learning Conference (2024)

  153. [161]

    van der Wal, J.: Stochastic Dynamic Programming. No. 139 in Mathematical Centre tracts, Morgan Kaufmann, Amsterdam (1981)

  154. [162]

    In: 2011 IEEE Symposium on Adaptive Dynamic Programming and Reinforcement Learning (ADPRL)

    Whiteson, S., Tanner, B., Taylor, M.E., Stone, P.: Protecting against evalua- tion overfitting in empirical reinforcement learning. In: 2011 IEEE Symposium on Adaptive Dynamic Programming and Reinforcement Learning (ADPRL). pp. 120–127 (2011)

  155. [163]

    In: 2018 IEEE International Conference on Robotics and Automation

    Williams, E.C., Gopalan, N., Rhee, M., Tellex, S.: Learning to parse natural language to gronuded reward functions with weak supervision. In: 2018 IEEE International Conference on Robotics and Automation. pp. 4430–4436 (2018)

  156. [164]

    In: Proceedings of the 24th International Confer- ence on Machine Learning

    Wilson, A., Fern, A., Ray, S., Tadepalli, P.: Multi-task reinforcement learning: A hierarchical Bayesian approach. In: Proceedings of the 24th International Confer- ence on Machine Learning. pp. 1015–1022 (2007)

  157. [165]

    https:// arxiv.org/abs/2101.02230 (2021)

    Yang, K.: Learn dynamic-aware state embedding for transfer learning. https:// arxiv.org/abs/2101.02230 (2021)

  158. [166]

    In: International Confer- ence on Learning Representations (2024)

    Yang, M., Du, Y., Ghasemipour, K., Tompson, J., Kaelbling, L., Schuurmans, D., Abbeel, P.: Learning interactive real-world simulators. In: International Confer- ence on Learning Representations (2024)

  159. [167]

    In: Proceedings of the International Conference on Learning Representations (2021)

    Yoon, D., Hong, S., Lee, B.J., Kim, K.E.: Winning the L2RPN challenge: Power grid management via semi-Markov afterstate actor critic. In: Proceedings of the International Conference on Learning Representations (2021)

  160. [168]

    https://arxiv.org/abs/1806.07937 (2020)

    Zhang, A., Ballas, N., Pineau, J.: A dissection of overfitting and generalization in continuous reinforcement learning. https://arxiv.org/abs/1806.07937 (2020)

  161. [169]

    https://arxiv.org/abs/1804.06893 (2018)

    Zhang, C., Vinyals, O., Munos, R., Bengio, S.: A study on overfitting in deep reinforcement learning. https://arxiv.org/abs/1804.06893 (2018)

  162. [170]

    In: 2020 IEEE Symposium Series on Com- putational Intelligence (SSCI)

    Zhao, W., Queralta, J.P., Westerlund, T.: Sim-to-real transfer in deep reinforce- ment learning for robotics: A survey. In: 2020 IEEE Symposium Series on Com- putational Intelligence (SSCI). pp. 737–744 (2020)

  163. [171]

    https://arxiv.org/abs/2303.18223 (2023), accessed: 25-01-2024 A Research Agenda for Usability and Generalisation in RL 31

    Zhao, W.X., Zhou, K., Li, J., Tang, T., Wang, X., Hou, Y., Min, Y., Zhang, B., Zhang, J., Dong, Z., Du, Y., Yang, C., Chen, Y., Chen, Z., Jiang, J., Ren, R., Li, Y., Tang, X., Liu, Z., Liu, P., Nie, J.Y., Wen, J.R.: A survey of large language models. https://arxiv.org/abs/2303...

  164. [172]

    In: International Conference on Learning Repre- sentations (2020)

    Zhong, V., Rocktäschel, T., Grefenstette, E.: RTFM: Generalising to novel envi- ronment dynamics via reading. In: International Conference on Learning Repre- sentations (2020)

  165. [173]

    https://arxiv

    Zhu, Z., de Salvo Braz, R., Bhandari, J., Jiang, D., Wan, Y., Efroni, Y., Wang, L., Xu, R., Guo, H., Nikulkov, A., Korenkevych, D., Dogan, U., Cheng, F., Wu, Z., Xu, W.: Pearl: A production-ready reinforcement learning agent. https://arxiv. org/abs/2312.03814 (2023)

  166. [174]

    IEEE Transactions on Pattern Analysis and Machine Intelli- gence 45(11), 13344–13362 (2023)

    Zhu, Z., Lin, K., Jain, A.K., Zhou, J.: Transfer learning in deep reinforcement learning: A survey. IEEE Transactions on Pattern Analysis and Machine Intelli- gence 45(11), 13344–13362 (2023)

  167. [175]

    Zintgraf, L.: Fast Adaptation via Meta Reinforcement Learning. Ph.D. thesis, University of Oxford, Oxford, United Kingdom (2022)

  168. [176]

    https://arxiv

    Zuo, M., Velez, F.P., Li, X., Littman, M.L., Bach, S.H.: Planetarium: A rigorous benchmark for translating text to structured planning languages. https://arxiv. org/abs/2407.03321 (2024)

  169. [1708]

    AAAI Press (2020) A Research Agenda for Usability and Generalisation in RL 23

  170. [8045]

    Curran Associates, Inc. (2021)

Pith tools

Reviewed August 11, 2026 · model on record in the stance chip above.