Pith. sign in

Paper Citation Record · LEDGER

How Should We Meta-Learn Reinforcement Learning Algorithms?

As of 17 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 1 inbound Pith citation observation for arXiv:2507.17668.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17668 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:48:52.758471Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:39:32.117730Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact4
  • verified fuzzy26
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a3d7ccb0-92bc-4646-b308-e5a76539886e · outbound

This paper cites Loss of Plasticity in Continual Deep Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Loss of Plasticity in Continual Deep Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.616239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.616239Z digest=sha256:381f052046be441f19eabaf7754b888008d81fdf3dc7f245686795755b952c8f

Observation 24da469e-b54b-4146-bb26-ec474352ed58 · outbound

This paper cites Towards Characterizing Divergence in Deep Q-Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Towards Characterizing Divergence in Deep Q-Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.681243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.681243Z digest=sha256:c7f866845d678ff9b7dbfc4b03a19d9341d2ba2f1eba93439b52eb58d27b7e0d

Observation f48faf5f-8f10-4095-b923-e3fe5bf3f8f4 · outbound

This paper cites A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.755232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.755232Z digest=sha256:a1a58ac511d4bead80d2ffba8104c8da8c2e436fc670f3d902b494ffc746f131

Observation b19f019a-93e9-4b0d-b221-c531836e6370 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

How Should We Meta-Learn Reinforcement Learning Algorithms? Deep reinforcement learning at the edge of the statistical precipice

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.809562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.809562Z digest=sha256:4c9a6ffda2195fc70830b5a08c971c76bc0cc53a2f0d53f1d79861a5c1c788e4

Observation 44b1374f-aa2b-4bff-a13f-84907f107a90 · outbound

This paper cites A Generalizable Approach to Learning Optimizers.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Generalizable Approach to Learning Optimizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.948249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.948249Z digest=sha256:5021a1ea8d11abb855f828f1d8b35ac5462b9cf1e68e1aeeac869b0702013b17

Observation 34b95b01-ac5e-4eaf-afc5-9dd2e849b3f3 · outbound

This paper cites Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando de Freitas.

How Should We Meta-Learn Reinforcement Learning Algorithms? Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando de Freitas

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.089106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.089106Z digest=sha256:d48f078188f5697fd9f364a17039c1672e6ee78d88d86db288e3d488295d7e1c

Observation 2fafa30b-8079-487a-bd81-65d3d1db55d3 · outbound

This paper cites An information-theoretic perspective on intrinsic motivation in reinforcement learning: A survey.

How Should We Meta-Learn Reinforcement Learning Algorithms? An information-theoretic perspective on intrinsic motivation in reinforcement learning: A survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.201906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.201906Z digest=sha256:2dc1dcd497b1b373a0cb11ff787efc6d3fc9f7e6dcad7ce6b65ded7541a3194b

Observation 59369bea-951d-4879-922e-ac071fdf0ea3 · outbound

This paper cites A Tutorial on Meta-Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Tutorial on Meta-Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.339004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.339004Z digest=sha256:8b77d58efb7f4c1ef639c26ff27e28ea9a0a5c50c73298224b343e1d51c01a7c

Observation 8fd95d5d-5e2d-46b6-84db-12a0aa2f778f · outbound

This paper cites OpenAI gym, 2016.

How Should We Meta-Learn Reinforcement Learning Algorithms? OpenAI gym, 2016

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.485426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.485426Z digest=sha256:a0de8ba73f42bbda2cc1ad02b55e5f19275f7b987a3a0cb105f3fd0c4e6ab0ac

Observation 379f9e2e-23a4-4b86-8f11-d59bb3c045d6 · outbound

This paper cites Exploration by Random Network Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Exploration by Random Network Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.624540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.624540Z digest=sha256:2bb2341da124166a95aa92cd58910c7407906df91106a5094940ee16156bdea7

Observation 5dc19526-a0f3-4c51-bbe0-c15aaf14490d · outbound

This paper cites Boltzmann exploration done right.

How Should We Meta-Learn Reinforcement Learning Algorithms? Boltzmann exploration done right

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.793806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.793806Z digest=sha256:1244cbb892ce35e85551a61cbce774fea5ec7603b67e6668e993754391c51412

Observation 2049f298-757a-4752-a563-7d78d809c499 · outbound

This paper cites Symbolic Discovery of Optimization Algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Symbolic Discovery of Optimization Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.939586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.939586Z digest=sha256:ec4a256933b9f39c642448d2890c1c87ca4786495a5fd2f5cdd744f4ab9bf03a

Observation 1e634b23-c977-4fb1-adee-0abef0b42e18 · outbound

This paper cites Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl.

How Should We Meta-Learn Reinforcement Learning Algorithms? Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.030210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.030210Z digest=sha256:db614840e4dd9f759f223cd1b91234f223f76fba4e524069d1c59b5e514b3110

Observation 1aea3f56-9912-422e-8722-f2befb70c15f · outbound

This paper cites Discovering Symbolic Models from Deep Learning with Inductive Biases.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering Symbolic Models from Deep Learning with Inductive Biases

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.105628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.105628Z digest=sha256:b91b128f29c2451ce97ac5ed9a210397c6aa1a95974642b5dcf2503a939c4179

Observation 707eb909-41c4-4c9a-bf29-55fb711e05c8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.158711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.158711Z digest=sha256:2378c0b06365874a66156ef69f2fc6a67ea5ac77732240113bf75a0f1ad21ef0

Observation 96506966-4351-4c0c-80b6-4511d5ffe9e5 · outbound

This paper cites Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design.

How Should We Meta-Learn Reinforcement Learning Algorithms? Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.239937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.239937Z digest=sha256:9ce2884af77c554ce65bdaefdc536d68c71e026790cc64d2d1cb7ded6393ab9c

Observation 82b12f95-c64e-4c2d-849b-84e96f4da4d9 · outbound

This paper cites Loss of plasticity in deep continual learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Loss of plasticity in deep continual learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.309720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.309720Z digest=sha256:24aea2926db1b1918019f78822ce983fc5c4a35ddb1925754176d0729941597f

Observation bad298e5-b312-469a-9370-21d35555a317 · outbound

This paper cites RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.395863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.395863Z digest=sha256:ef1422c5d0c64fe933cabceb1975e21729791a83e74b006373ffa9f48c46f45d

Observation f70732c3-e78c-447d-b9f7-793c88ad0058 · outbound

This paper cites Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps.

How Should We Meta-Learn Reinforcement Learning Algorithms? Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:48:54.113207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:46.513473Z digest=sha256:51da57c3819aa444cd50b62f0c888d8b8480cadbba800a82a65fc21fa0db26ed

Observation e97cd32c-805c-4709-bced-cce321896e67 · outbound

This paper cites OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code.

How Should We Meta-Learn Reinforcement Learning Algorithms? OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.585185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.585185Z digest=sha256:be04e0f4c8a10bbc9b5d86b564a3a678c23a7630b07b3561ef87cf7e5ad0e2e2

Observation 6454b23b-eafe-4b95-bb70-59f7f5a77502 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

How Should We Meta-Learn Reinforcement Learning Algorithms? Model-agnostic meta-learning for fast adaptation of deep networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.675019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.675019Z digest=sha256:2ccb1567dc7de8bf253609054b1fb9d631c7ea00ef307e254cab86bb48f1b926

Observation 23895565-0168-4277-9fd2-a917f64ddf2e · outbound

This paper cites Noisy Networks for Exploration.

How Should We Meta-Learn Reinforcement Learning Algorithms? Noisy Networks for Exploration

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.738150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.738150Z digest=sha256:8ed96c62c780ab0c50bb9ba98a2cce077d86447ddcc64a587bdce27f32389b84

Observation faaafd20-4489-409b-b34e-4e5bd93cbcb1 · outbound

This paper cites Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem.

How Should We Meta-Learn Reinforcement Learning Algorithms? Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.822940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.822940Z digest=sha256:710f15cc81242c21ce65bf7ec3b4cd8014ed479e6aff172e8ef64fedb49b2834

Observation 0648e23b-f653-4be6-b3c1-e4613c3559fc · outbound

This paper cites Born Again Neural Networks.

How Should We Meta-Learn Reinforcement Learning Algorithms? Born Again Neural Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.904574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.904574Z digest=sha256:3fcf3d7750205866c6eaf7e4c11e2c0d6b0d9106beb633593e011813bcb83553

Observation c4e7164c-4f0c-4ab6-a9c1-9520d6a6f49a · outbound

This paper cites Goldie, Chris Lu, Matthew T.

How Should We Meta-Learn Reinforcement Learning Algorithms? Goldie, Chris Lu, Matthew T

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:59.178524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.000738Z digest=sha256:d615ff7cb532826d31e5bbf284c098eb29e1d673075d900faa270d7cab72a6b3

Observation 0d14a9d1-f1ad-4f76-9b59-997042574c7c · outbound

This paper cites Benchmarking the Spectrum of Agent Capabilities.

How Should We Meta-Learn Reinforcement Learning Algorithms? Benchmarking the Spectrum of Agent Capabilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.078571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.078571Z digest=sha256:a4b2deba830f35c1b7a1c1f7bdc1b207f6e1ce15d728968f6175e53cab865228

Observation 3d14ae1c-d8f4-47b8-a380-a881f6b7f07d · outbound

This paper cites Distilling the Knowledge in a Neural Network.

How Should We Meta-Learn Reinforcement Learning Algorithms? Distilling the Knowledge in a Neural Network

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.149052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.149052Z digest=sha256:e80ce55981da50c8e72b6dea5029765995c1a0417df9c5fb38aa054f26c74cf9

Observation 4080903f-56ca-481e-9526-3d647585a01f · outbound

This paper cites Long short-term memory.

How Should We Meta-Learn Reinforcement Learning Algorithms? Long short-term memory

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.222901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.222901Z digest=sha256:6a59ac37b99a9658db73008a38cb5d670b0c6d9b499921be14a555021b380e16

Observation dc9c16ba-60c7-4c79-90e9-7a662088ee40 · outbound

This paper cites Automated Design of Agentic Systems.

How Should We Meta-Learn Reinforcement Learning Algorithms? Automated Design of Agentic Systems

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.306436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.306436Z digest=sha256:693df00e401bfcf8a39321615c9cbdb7553512b81c33dddc5d7fdbf67d9efff6

Observation 41b311f6-d302-40b8-a5ea-f13008d342b1 · outbound

This paper cites Transient non-stationarity and generalisation in deep reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Transient non-stationarity and generalisation in deep reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.888824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.383576Z digest=sha256:fd654703e08ff261d4384d409ca3339e06c8240017501a56e914475fd3497e77

Observation 0edfac8e-8273-47a5-aa23-a12663df7d34 · outbound

This paper cites Transient non-stationarity and generalisation in deep reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Transient non-stationarity and generalisation in deep reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.610653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.471532Z digest=sha256:678fb5140dcaf563c09f7da113638ef626623bdfbcdcc183850c242959fb0b58

Observation 7d42ccf3-7005-4fd5-8841-a55667aa143c · outbound

This paper cites Discovering General Reinforcement Learning Algorithms with Adversarial Environment Design.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering General Reinforcement Learning Algorithms with Adversarial Environment Design

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:53.861384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.539950Z digest=sha256:fe10f26eceec38ebe614be3c4d6fcdcf96fc131f2caaf78d10011aaa4dbf2240

Observation 820da51f-8024-4178-abbc-91a7e9405067 · outbound

This paper cites Discovering temporally-aware reinforcement learning algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering temporally-aware reinforcement learning algorithms

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.350351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.617075Z digest=sha256:f8e4f801c9a8d66d81c189656e4ba4f8cd5a8b970d586bbc6dd5ca899cf518dc

Observation 6bf5b200-36ad-4e13-9410-c95c9baf1e10 · outbound

This paper cites Improving policy optimization with generalist-specialist learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Improving policy optimization with generalist-specialist learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.129243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.682731Z digest=sha256:84f6de6d6831e2394fb4846c3b1ab8426aca21baa5b13661cff683588974109e

Observation 3a1a5cb6-507c-49c1-adf6-d89d5bf2585e · outbound

This paper cites Meta Learning Backpropagation And Improving It.

How Should We Meta-Learn Reinforcement Learning Algorithms? Meta Learning Backpropagation And Improving It

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:53.690763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.763840Z digest=sha256:c5ff0fd65eacd5c92fbdf7f1f1feb9e15ed6cff70249b122a96a2506a9bd0977

Observation 267ff5a4-13c3-4ac9-98bb-2b92691851ef · outbound

This paper cites Improving generalization in meta reinforcement learning using learned objectives.

How Should We Meta-Learn Reinforcement Learning Algorithms? Improving generalization in meta reinforcement learning using learned objectives

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.828958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.817050Z digest=sha256:b5deee32fe30e959c28a0ad35090bb995ed613e0f73233308b3b9c67ef0a77af

Observation 2517e4d0-afda-4d88-927b-a6d3a685f536 · outbound

This paper cites Mirror Learning: A Unifying Framework of Policy Optimisation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Mirror Learning: A Unifying Framework of Policy Optimisation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.878994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.878994Z digest=sha256:574c1c57cf1d5d491466bc21a38ff5ab1ce269750ea587e79ece77e460d5f7f0

Observation 853ffeb4-20c3-4278-8903-a0cc889d3f0f · outbound

This paper cites Learning to Optimize for Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Learning to Optimize for Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.922233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.922233Z digest=sha256:7fdb1416d02965e23dee63ec526172fbed3ae5d1eb3f634ac04be72b7e1651c1

Observation 670b210f-af41-461e-9f14-69132d7b1c52 · outbound

This paper cites gymnax: A JAX -based reinforcement learning environment library, 2022 a.

How Should We Meta-Learn Reinforcement Learning Algorithms? gymnax: A JAX -based reinforcement learning environment library, 2022 a

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.659181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.972077Z digest=sha256:14f926f2911270f749e374a33f23bc7e8cd0fa5283e9c58c9301fb4b41d456fe

Observation 1c5b6137-edd7-4d39-98f1-7d410a52472f · outbound

This paper cites evosax: JAX-based Evolution Strategies.

How Should We Meta-Learn Reinforcement Learning Algorithms? evosax: JAX-based Evolution Strategies

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:48.110073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:48.110073Z digest=sha256:3524ddb3a6265c30fb43ab7f38dea398ba5ce5c1ecca166e57c5a512cd4c0b95

Observation 2a2eb53a-4375-4fbb-9164-1db3ebb27504 · outbound

This paper cites In-context reinforcement learning with algorithm distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? In-context reinforcement learning with algorithm distillation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.496541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.255410Z digest=sha256:a3034bfdcc9de300a36b5b53d005704d1b8941b5e30b1fe29c8870375792fb4b

Observation a65675a6-b2a8-42da-91e0-7c76d1806eac · outbound

This paper cites Evolution through Large Models.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolution through Large Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:48.383996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:48.383996Z digest=sha256:46b79b92b454c358f0cc5ece0758fff0ff9c7d644f72b905ff591c6fd8fc2306

Observation ea817af4-650b-4692-a29a-e8901cb12399 · outbound

This paper cites Rediscovering orbital mechanics with machine learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Rediscovering orbital mechanics with machine learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.338948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.486735Z digest=sha256:b647a325722ecaa9614d09d607044d84e97a764e6ac2c9dc7dda2abb0a8a0600

Observation 35ca94b3-032a-4957-86d1-36418967adeb · outbound

This paper cites Discovered policy optimisation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovered policy optimisation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.186310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.575011Z digest=sha256:e0f1e827ef988159cad7fee8c98245101b4f851d8870e02fb52d13cb2880ede6

Observation 4d9a5bb5-e31a-4381-89b2-6e54d990ab38 · outbound

This paper cites Discovering Preference Optimization Algorithms with and for Large Language Models.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering Preference Optimization Algorithms with and for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:48.754244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:48.754244Z digest=sha256:3140680ebb9ca24daf3b1989e2621672b7cd695cd85481a62609bbaff6146b4a

Observation af015423-f7fa-41e0-9775-a370f6364081 · outbound

This paper cites Behaviour Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Behaviour Distillation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.015958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.891797Z digest=sha256:fdef280910a61f10fd9755b445179db5c03395315af1c4493032bf618175d94a

Observation 0c5f61b8-a66c-4923-a626-748f3e6bdea9 · outbound

This paper cites Understanding plasticity in neural networks.

How Should We Meta-Learn Reinforcement Learning Algorithms? Understanding plasticity in neural networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.023177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.023177Z digest=sha256:6846c9e7a324640777fa75bd2a505288b30f6e4973687b66d1276a1f59f7e12f

Observation 0336ceb2-4e37-44af-9cc1-59d24cf54dfd · outbound

This paper cites Craftax: a lightning-fast benchmark for open-ended reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Craftax: a lightning-fast benchmark for open-ended reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.824630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:49.134832Z digest=sha256:28e44f99a2275df1b76c1faf5c4607514987d0fa660b7f29453c352b258addff

Observation 759aebf1-ca1f-4df1-8d37-a0074a1264b0 · outbound

This paper cites Interpretable machine learning methods applied to jet background subtraction in heavy-ion collisions.

How Should We Meta-Learn Reinforcement Learning Algorithms? Interpretable machine learning methods applied to jet background subtraction in heavy-ion collisions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.652401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:49.301257Z digest=sha256:50ec05bc44d2e66551581ed62f437986c5592d23fc7956432598587c40edf82e

Observation 6ce3b857-d15e-46fb-98c2-40ec67b58731 · outbound

This paper cites Meta-Learning Update Rules for Unsupervised Representation Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Meta-Learning Update Rules for Unsupervised Representation Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.432448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.432448Z digest=sha256:32f6b2f913d2a7af9f54bd974ea21686a2e050ce225e020704ec6fe27e79a7fc

Observation 0df60976-ee7b-48c8-8d27-81fe624486ab · outbound

This paper cites Understanding and correcting pathologies in the training of learned optimizers.

How Should We Meta-Learn Reinforcement Learning Algorithms? Understanding and correcting pathologies in the training of learned optimizers

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:48:53.486491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:49.527570Z digest=sha256:dff5bd194d370b68506e9a0d42c9f1b5d327222f7d629f0a81796dac880c4132

Observation 44f62159-29b7-427f-9199-f6b3536a3d0c · outbound

This paper cites Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves.

How Should We Meta-Learn Reinforcement Learning Algorithms? Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.605196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.605196Z digest=sha256:87592580b49c79dc8c74c43a70301fb4caa0d4b2c3774e282b855c2a7191cb2e

Observation 0610a14e-4f3a-47f2-9159-606f1641d390 · outbound

This paper cites Gradients are Not All You Need.

How Should We Meta-Learn Reinforcement Learning Algorithms? Gradients are Not All You Need

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.755270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.755270Z digest=sha256:1d7ca191ac98fdb64dad349bf85fc5a755c95a9b791586b1687ea7b2ab27e788

Observation 76f8e25e-1c33-4720-a7f4-075bb6db140e · outbound

This paper cites VeLO: Training Versatile Learned Optimizers by Scaling Up.

How Should We Meta-Learn Reinforcement Learning Algorithms? VeLO: Training Versatile Learned Optimizers by Scaling Up

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.872182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.872182Z digest=sha256:5145c483f5df862d8297dcf33f0b0f6a6a8672a88ebd3f48d4bb93553b769166

Observation a99fa079-00a9-49d9-b6e0-52fbad2f45cb · outbound

This paper cites Self-distillation amplifies regularization in hilbert space.

How Should We Meta-Learn Reinforcement Learning Algorithms? Self-distillation amplifies regularization in hilbert space

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.459946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.019447Z digest=sha256:99f56161fb8e9d188e9c548b11393080a42ff6785cdaf87505b323e47477a65a

Observation 72858513-525e-4cb8-8404-30205921394e · outbound

This paper cites Small batch deep reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Small batch deep reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.295508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.139633Z digest=sha256:19d791df45f24c4e09d566ec8a8fa8684b2f840657c89015825b95b055141a49

Observation afcfcb2f-aebb-41cd-89c1-e2a2d7896257 · outbound

This paper cites Discovering reinforcement learning algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering reinforcement learning algorithms

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.134460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.254660Z digest=sha256:9747631ea56644e91717aa1d746b8bb0380255d20b2096bd8c629ffd1eec74b9

Observation 8075000c-14fd-47d8-9aed-dfdac0ce7168 · outbound

This paper cites Openai o3-mini, January 2025.

How Should We Meta-Learn Reinforcement Learning Algorithms? Openai o3-mini, January 2025

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.965617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.382465Z digest=sha256:4a543a1052882c545c5d1333513bf7fa51b4ca835e5b6ecb9425b4653c2523d9

Observation 836da2ff-865c-4390-b4c4-0b61b1be5c16 · outbound

This paper cites Stabilizing transformers for reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Stabilizing transformers for reinforcement learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.789331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.452623Z digest=sha256:6fe280bcad9ea56484f2c67752bcc3f3db634a4a648d8977c141190e7a688879

Observation b3da679a-906b-4c9d-810a-8580885f52c3 · outbound

This paper cites Evolving Curricula with Regret - Based Environment Design.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolving Curricula with Regret - Based Environment Design

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.589660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.495021Z digest=sha256:5fd9f0f7c3d9114edd4ce3a0d41bf02c657855c7097374d007b69ea75000d05e

Observation 8f69eaaf-88c3-44ef-a1dd-9c73bda00a80 · outbound

This paper cites Parameter Space Noise for Exploration.

How Should We Meta-Learn Reinforcement Learning Algorithms? Parameter Space Noise for Exploration

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:50.593230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:50.593230Z digest=sha256:72b05b1662370dfc89675a1afaa0628a908f58b1e184647634f10e7712a64543

Observation d359abff-4acd-473c-8099-de350efa7649 · outbound

This paper cites Tunability: Importance of hyperparameters of machine learning algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Tunability: Importance of hyperparameters of machine learning algorithms

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.346301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.706376Z digest=sha256:c69866e079f9bb3a56c574ff19e7e645174a17b40b7e9a3dbb02b4869887eaae

Observation a2e2d967-75c0-47e9-8a37-fb87210c6462 · outbound

This paper cites Evolutionsstrategie : Optimierung technischer systeme nach prinzipien der biologischen evolution.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolutionsstrategie : Optimierung technischer systeme nach prinzipien der biologischen evolution

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.153556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.800901Z digest=sha256:d1fdc48ab34176cde939c4a72396019d82025862510b3ca44db8433c933cc595

Observation 9bb36e11-6709-4c14-9b4c-ea544ef8c666 · outbound

This paper cites Pawan Kumar, Emilien Dupont, Francisco J.

How Should We Meta-Learn Reinforcement Learning Algorithms? Pawan Kumar, Emilien Dupont, Francisco J

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:50.860051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:50.860051Z digest=sha256:b3f53c82dead98a963ff43e8b3d2a245047c79bc366fe54402c5d023097ebfb3

Observation f77d5312-72e7-4aeb-93d5-b885a82980a1 · outbound

This paper cites Policy Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Policy Distillation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:50.969246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:50.969246Z digest=sha256:0c0fd2b0de04769c7a91306e5feb06269ac944d3b07859fde4c3df6bffa00d57

Observation 1dfd04f3-40b7-4c24-ac86-7f19ddad9f85 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.080292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.080292Z digest=sha256:8e6e26b1bf267f05644dcb16fe7b34b32b795512ee7dcd07111d026e57f7aa4f

Observation 21e3588a-9d4c-44ad-b16e-341ea7d14213 · outbound

This paper cites Proximal Policy Optimization Algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Proximal Policy Optimization Algorithms

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.130615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.130615Z digest=sha256:427c05865e8c10d3b0882fd2ad6ecfa30dae04ca9b65657346e744cffffeb8a2

Observation 30db874b-70dd-4448-bbc1-1cb7d6c5b9ef · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

How Should We Meta-Learn Reinforcement Learning Algorithms? High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.229277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.229277Z digest=sha256:0ffb803bdaa97bc1d560d5f1b26f2f5f6a56b6bd85e1d334f5ea8093771cead1

Observation 945e19ab-25ed-45ef-90d3-58f813481d85 · outbound

This paper cites The Dormant Neuron Phenomenon in Deep Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? The Dormant Neuron Phenomenon in Deep Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.326750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.326750Z digest=sha256:df787f6bf2e6ce9b687350b6f1e2e39612aa3620486357b0e2f88c94fe02cc6a

Observation d765ea3d-991b-4127-946d-725e108653ad · outbound

This paper cites Distilling Reinforcement Learning Algorithms for In-Context Model-Based Planning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Distilling Reinforcement Learning Algorithms for In-Context Model-Based Planning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.431138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.431138Z digest=sha256:b747e3da6bf4fa780b54caa559a11a18fb6bef09d988cc9da4e08d7fa0eda345

Observation 3f177879-b4e0-4a29-adab-9a732bd859f2 · outbound

This paper cites Generalizable Symbolic Optimizer Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Generalizable Symbolic Optimizer Learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.021535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:51.499138Z digest=sha256:264b36634326e1b6f6c6604692ec24f83ae217a5ea92d8481b9819e31f08a637

Observation cf324fbd-f556-48f8-ba71-a16468de3b0e · outbound

This paper cites Position: Leverage Foundational Models for Black-Box Optimization.

How Should We Meta-Learn Reinforcement Learning Algorithms? Position: Leverage Foundational Models for Black-Box Optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.550816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.550816Z digest=sha256:dd055a8c007a69670b82f3c7e152a1f6ccaf245e1608253395137736e313301a

Observation 4d92d76f-80a5-4dc5-9d43-0693b7e3567e · outbound

This paper cites Maxinfo RL : Boosting exploration in reinforcement learning through information gain maximization.

How Should We Meta-Learn Reinforcement Learning Algorithms? Maxinfo RL : Boosting exploration in reinforcement learning through information gain maximization

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.840199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:51.635219Z digest=sha256:c8573f3ecddcdf777087fa39293a688c0be6c167419f2acee554141aba3c6ec5

Observation 9f2931ef-91ed-4c19-9777-88dfd6d5e8e6 · outbound

This paper cites Sutton and Andrew Barto.

How Should We Meta-Learn Reinforcement Learning Algorithms? Sutton and Andrew Barto

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.668222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:51.716178Z digest=sha256:bac3989d4b2267ff13210fad5c39462de8e40dbf879206a677042a2f7bbdd137

Observation 3acf0046-0b42-41dd-8def-cac57e2a55b3 · outbound

This paper cites Improving deep reinforcement learning by reducing the chain effect of value and policy churn.

How Should We Meta-Learn Reinforcement Learning Algorithms? Improving deep reinforcement learning by reducing the chain effect of value and policy churn

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.798344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.798344Z digest=sha256:1a5fd932fe43600e3e01b038ba2d7f4db792150060bc1418f6017cfaa3ba1bf4

Observation 7658bc7e-8099-4ad4-8afb-4a365f10451b · outbound

This paper cites MuJoCo : A physics engine for model-based control.

How Should We Meta-Learn Reinforcement Learning Algorithms? MuJoCo : A physics engine for model-based control

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.925909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.925909Z digest=sha256:6cd2101e98b27e320380fd18ff9d7cf08b3a28b87fe24d45a2bacce109aaf1e3

Observation 9d4d48f1-3bb7-4255-a172-039e321357a0 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

How Should We Meta-Learn Reinforcement Learning Algorithms? Deep Reinforcement Learning and the Deadly Triad

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.022023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.022023Z digest=sha256:9da9b9adf7a31ee19a57e0108d414f0a9ef76d8d04c0461161b9c65d9855e7be

Observation b9178950-d006-48d7-84c0-46e344cce11d · outbound

This paper cites Attention Is All You Need.

How Should We Meta-Learn Reinforcement Learning Algorithms? Attention Is All You Need

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.137791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.137791Z digest=sha256:229a5bdb41ab66d9432dca2e4dd2ef73bc1f75898255dab8879c5698e767acaf

Observation 8f07cb6a-ce7b-45e4-8917-29771434a0ae · outbound

This paper cites Dataset Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Dataset Distillation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.204307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.204307Z digest=sha256:80333ed39e7f93401e6f8238d5299a380603928b3b892babd7f2df60f8efbfb8

Observation 59339402-f08a-4723-bd2c-7434f8185b54 · outbound

This paper cites Natural Evolution Strategies.

How Should We Meta-Learn Reinforcement Learning Algorithms? Natural Evolution Strategies

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:53.127951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.287410Z digest=sha256:d2b3296e4bfa229240e704a5ea6bddaebb8a598408633adb0788c044e7e857ed

Observation 985ed171-fb21-495a-b2b0-84d329f3aa4b · outbound

This paper cites Understanding short-horizon bias in stochastic meta-optimization.

How Should We Meta-Learn Reinforcement Learning Algorithms? Understanding short-horizon bias in stochastic meta-optimization

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.523597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.381492Z digest=sha256:ce310288fbbe2af0b5a550b7e64e24a12e6848971e4bdb0e94a21bb909413b59

Observation 1e7ab7c4-6b5d-4d2a-8ff5-7a5056aac1b8 · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

How Should We Meta-Learn Reinforcement Learning Algorithms? MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.488734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.488734Z digest=sha256:d3ef497db8b8e069cd218f93aae68441cdf901c5122e933b942d1f6214df544b

Observation cc47c1d0-9afe-4395-bec3-813cad431c96 · outbound

This paper cites Self-distillation as instance-specific label smoothing.

How Should We Meta-Learn Reinforcement Learning Algorithms? Self-distillation as instance-specific label smoothing

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.343900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.573001Z digest=sha256:8c80c8e527286f72b99da087dfd61331b3e897ecd01b16d28176e1f74648878a

Observation 44945f70-98c5-4abf-b02c-65bc7fd38f80 · outbound

This paper cites Symbolic Learning to Optimize: Towards Interpretability and Scalability.

How Should We Meta-Learn Reinforcement Learning Algorithms? Symbolic Learning to Optimize: Towards Interpretability and Scalability

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:52.957731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.681440Z digest=sha256:a940e436579bffabde708b0e8c61978ab7f9163a9ff4ea8d6b61f0a33ae9b449

Observation 41439ad1-985a-4ebb-b5bf-0348a9646f81 · outbound

This paper cites write newline.

How Should We Meta-Learn Reinforcement Learning Algorithms? write newline

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.758471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.758471Z digest=sha256:134a74e1561908053697a9624ebc9d627d6a44011f7f236fd66d7de9875abd7f

Pith citing papers

Observation c32431f8-ff8d-4845-9333-b867b3e88878 · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback How Should We Meta-Learn Reinforcement Learning Algorithms?

Reference 243

Resolution
verified exact
local_arxiv, observed 2026-08-03T04:44:19.150566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-03T04:39:32.117730Z digest=sha256:263767aa68c48f8bb4ff388dee3e744512e38ac81a3f54d41eab2b267af98413