Pith. sign in

Paper Citation Record · LEDGER

How Should We Meta-Learn Reinforcement Learning Algorithms?

As of 8 August 2026, this Paper Citation Record lists 85 of 85 outbound references and 1 inbound Pith citation observation for arXiv:2507.17668.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.17668 v2

Coverage vector

measured 85 of 85 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:48:52.758471Z

measured 86 of 86 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T04:39:32.117730Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

85 of 85 outbound references displayed

  • verified exact4
  • verified fuzzy26
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

0
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation a3d7ccb0-92bc-4646-b308-e5a76539886e · outbound

This paper cites Loss of Plasticity in Continual Deep Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Loss of Plasticity in Continual Deep Reinforcement Learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.616239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.616239Z digest=sha256:3e5f9d3425885e6301186969bab8816753ec7e4999c6128be9f456f26b9f7838

Observation 24da469e-b54b-4146-bb26-ec474352ed58 · outbound

This paper cites Towards Characterizing Divergence in Deep Q-Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Towards Characterizing Divergence in Deep Q-Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.681243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.681243Z digest=sha256:9061fc253369e38b7297a7533cd1f62879796fd1643e7a177cb5b85d737e7ef5

Observation f48faf5f-8f10-4095-b923-e3fe5bf3f8f4 · outbound

This paper cites A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Method for Evaluating Hyperparameter Sensitivity in Reinforcement Learning

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.755232Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.755232Z digest=sha256:73cdb38c8888e01623433a4eb4091ace111dbcb704170259ef272625e1c31600

Observation b19f019a-93e9-4b0d-b221-c531836e6370 · outbound

This paper cites Deep reinforcement learning at the edge of the statistical precipice.

How Should We Meta-Learn Reinforcement Learning Algorithms? Deep reinforcement learning at the edge of the statistical precipice

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.809562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.809562Z digest=sha256:25b3f6bbf33bd1c72aca40720abb99b7269183d40853b93bb00d670d77e2a3f9

Observation 44b1374f-aa2b-4bff-a13f-84907f107a90 · outbound

This paper cites A Generalizable Approach to Learning Optimizers.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Generalizable Approach to Learning Optimizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:44.948249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:44.948249Z digest=sha256:3412da6022775b00ca7cfe1751ac40fc9e672a3b3d4deb352f42650259d0c0a0

Observation 34b95b01-ac5e-4eaf-afc5-9dd2e849b3f3 · outbound

This paper cites Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando de Freitas.

How Should We Meta-Learn Reinforcement Learning Algorithms? Hoffman, David Pfau, Tom Schaul, Brendan Shillingford, and Nando de Freitas

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.089106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.089106Z digest=sha256:bf7f78ce4677a2619ee8a640e53fa5ace44dee4bec2e7d6f50c51b7080769fc4

Observation 2fafa30b-8079-487a-bd81-65d3d1db55d3 · outbound

This paper cites An information-theoretic perspective on intrinsic motivation in reinforcement learning: A survey.

How Should We Meta-Learn Reinforcement Learning Algorithms? An information-theoretic perspective on intrinsic motivation in reinforcement learning: A survey

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.201906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.201906Z digest=sha256:824de9a1cc0bf36d0b3bfa4bfdfa8447d4094fda38324892ecd554a5d565a2b9

Observation 59369bea-951d-4879-922e-ac071fdf0ea3 · outbound

This paper cites A Tutorial on Meta-Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? A Tutorial on Meta-Reinforcement Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.339004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.339004Z digest=sha256:c62edc5dd53d1cacc528ea0d12af7043fd51df44883246d50c4457f138cde9e2

Observation 8fd95d5d-5e2d-46b6-84db-12a0aa2f778f · outbound

This paper cites OpenAI gym, 2016.

How Should We Meta-Learn Reinforcement Learning Algorithms? OpenAI gym, 2016

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.485426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.485426Z digest=sha256:24f5c737d6307502f569fe9219cb4f9cac48ad1d0b3c6cf7da91932e6d4e6f43

Observation 379f9e2e-23a4-4b86-8f11-d59bb3c045d6 · outbound

This paper cites Exploration by Random Network Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Exploration by Random Network Distillation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.624540Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.624540Z digest=sha256:91389a982ca66c01cb83a1e2b61fa29bebea4994d047e03796360ff6770d919e

Observation 5dc19526-a0f3-4c51-bbe0-c15aaf14490d · outbound

This paper cites Boltzmann exploration done right.

How Should We Meta-Learn Reinforcement Learning Algorithms? Boltzmann exploration done right

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.793806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.793806Z digest=sha256:bb992b9fb2fd62cb13c36c8b657187e0de2f6fac325779c671dae581ae8bfe38

Observation 2049f298-757a-4752-a563-7d78d809c499 · outbound

This paper cites Symbolic Discovery of Optimization Algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Symbolic Discovery of Optimization Algorithms

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:45.939586Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:45.939586Z digest=sha256:dc32c82f5d8ea1714d7ea9dc5549593ecf22ac90d0e9716bdc84914bcbbbef4d

Observation 1e634b23-c977-4fb1-adee-0abef0b42e18 · outbound

This paper cites Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl.

How Should We Meta-Learn Reinforcement Learning Algorithms? Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.030210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.030210Z digest=sha256:0b9ff246d7c97eac69363ca4bb6e40bfe6937f1cebffe89747910a63c9ede918

Observation 1aea3f56-9912-422e-8722-f2befb70c15f · outbound

This paper cites Discovering Symbolic Models from Deep Learning with Inductive Biases.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering Symbolic Models from Deep Learning with Inductive Biases

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.105628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.105628Z digest=sha256:bb75643c18f7d88ebf44c282109fbc19e921ef842a39f68cd01eadc66063225e

Observation 707eb909-41c4-4c9a-bf29-55fb711e05c8 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.158711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.158711Z digest=sha256:759adfae2a4a740f9e4781cafb435dc94f9362959950a106aa7f2917f055bba2

Observation 96506966-4351-4c0c-80b6-4511d5ffe9e5 · outbound

This paper cites Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design.

How Should We Meta-Learn Reinforcement Learning Algorithms? Emergent Complexity and Zero-shot Transfer via Unsupervised Environment Design

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.239937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.239937Z digest=sha256:b1dea7b4b1c3206cfc92d08529b10a4e70ec615492648dd8fb799cecd65cd2ff

Observation 82b12f95-c64e-4c2d-849b-84e96f4da4d9 · outbound

This paper cites Loss of plasticity in deep continual learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Loss of plasticity in deep continual learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.309720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.309720Z digest=sha256:d4de5a0b9249ab4e38a830f4b7a151c2e372d818cba690c2df8241dbcd440743

Observation bad298e5-b312-469a-9370-21d35555a317 · outbound

This paper cites RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? RL$^2$: Fast Reinforcement Learning via Slow Reinforcement Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.395863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.395863Z digest=sha256:c983f80d2edd0f22b81b5235afef853be384ee1482b5d038de57f97f04a30d34

Observation f70732c3-e78c-447d-b9f7-793c88ad0058 · outbound

This paper cites Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps.

How Should We Meta-Learn Reinforcement Learning Algorithms? Adam on Local Time: Addressing Nonstationarity in RL with Relative Adam Timesteps

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:48:54.113207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:46.513473Z digest=sha256:6f56d482c359b626734d9efe0fb3ce33c5f441bf129b2e1053a03fa43385a46c

Observation e97cd32c-805c-4709-bced-cce321896e67 · outbound

This paper cites OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code.

How Should We Meta-Learn Reinforcement Learning Algorithms? OMNI-EPIC: Open-endedness via Models of human Notions of Interestingness with Environments Programmed in Code

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.585185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.585185Z digest=sha256:286416b45037dd4811fd7de2f266fb8baf6947d523abb2d38e189377ae07175d

Observation 6454b23b-eafe-4b95-bb70-59f7f5a77502 · outbound

This paper cites Model-agnostic meta-learning for fast adaptation of deep networks.

How Should We Meta-Learn Reinforcement Learning Algorithms? Model-agnostic meta-learning for fast adaptation of deep networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.675019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.675019Z digest=sha256:9e2d03c8da0c65e3138e6a24142060eec82c501197ad48ba58559f7ffc66d0f4

Observation 23895565-0168-4277-9fd2-a917f64ddf2e · outbound

This paper cites Noisy Networks for Exploration.

How Should We Meta-Learn Reinforcement Learning Algorithms? Noisy Networks for Exploration

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.738150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.738150Z digest=sha256:27f741486226bb9057b3079b90d4e30d8ac9f3c4e6d103c5c80095360304e2c0

Observation faaafd20-4489-409b-b34e-4e5bd93cbcb1 · outbound

This paper cites Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem.

How Should We Meta-Learn Reinforcement Learning Algorithms? Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.822940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.822940Z digest=sha256:c00eafc18a1800310e607786e40d40b2b614a6dfa2febf217a76337e03e6ca8d

Observation 0648e23b-f653-4be6-b3c1-e4613c3559fc · outbound

This paper cites Born Again Neural Networks.

How Should We Meta-Learn Reinforcement Learning Algorithms? Born Again Neural Networks

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:46.904574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:46.904574Z digest=sha256:b71cf9711853499a3951fbc86101fd59d5885caaaf3e6d148524d3612d94b806

Observation c4e7164c-4f0c-4ab6-a9c1-9520d6a6f49a · outbound

This paper cites Goldie, Chris Lu, Matthew T.

How Should We Meta-Learn Reinforcement Learning Algorithms? Goldie, Chris Lu, Matthew T

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:59.178524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.000738Z digest=sha256:7807c9001c0e02a345f0df418b670b4052dfc5b500d069f5258def042ece2aed

Observation 0d14a9d1-f1ad-4f76-9b59-997042574c7c · outbound

This paper cites Benchmarking the Spectrum of Agent Capabilities.

How Should We Meta-Learn Reinforcement Learning Algorithms? Benchmarking the Spectrum of Agent Capabilities

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.078571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.078571Z digest=sha256:256f2abc94f22dd07d562c1de63926736a58b94be50746f5012a08c166b5eabb

Observation 3d14ae1c-d8f4-47b8-a380-a881f6b7f07d · outbound

This paper cites Distilling the Knowledge in a Neural Network.

How Should We Meta-Learn Reinforcement Learning Algorithms? Distilling the Knowledge in a Neural Network

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.149052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.149052Z digest=sha256:9131add5f3ef80964ea173ceafa2ad7c166cec7f6620fc5cc7a4bfe4c2473f53

Observation 4080903f-56ca-481e-9526-3d647585a01f · outbound

This paper cites Long short-term memory.

How Should We Meta-Learn Reinforcement Learning Algorithms? Long short-term memory

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.222901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.222901Z digest=sha256:35a8bc555b1fe238738ff4a867c2f7db0dbbc82955c4912f3bab298c5a3358ba

Observation dc9c16ba-60c7-4c79-90e9-7a662088ee40 · outbound

This paper cites Automated Design of Agentic Systems.

How Should We Meta-Learn Reinforcement Learning Algorithms? Automated Design of Agentic Systems

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.306436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.306436Z digest=sha256:50a8f76a943c2e880704d7aed6b186a0caa045115e0900f119f05a35adc385b7

Observation 41b311f6-d302-40b8-a5ea-f13008d342b1 · outbound

This paper cites Transient non-stationarity and generalisation in deep reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Transient non-stationarity and generalisation in deep reinforcement learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.888824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.383576Z digest=sha256:4cb82587b7dd7d17176c0c7b027a9547b0e36428e0a76d58cfec13e804c85b35

Observation 0edfac8e-8273-47a5-aa23-a12663df7d34 · outbound

This paper cites Transient non-stationarity and generalisation in deep reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Transient non-stationarity and generalisation in deep reinforcement learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.610653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.471532Z digest=sha256:9d88c865f7350eaab2f4cda4db80416dc872094ea5af08c3ae6746297b6c1c2c

Observation 7d42ccf3-7005-4fd5-8841-a55667aa143c · outbound

This paper cites Discovering General Reinforcement Learning Algorithms with Adversarial Environment Design.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering General Reinforcement Learning Algorithms with Adversarial Environment Design

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:53.861384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.539950Z digest=sha256:e805665dddd30093179b2eb6ac276cbce9e39324b74443be88614e4b6164700a

Observation 820da51f-8024-4178-abbc-91a7e9405067 · outbound

This paper cites Discovering temporally-aware reinforcement learning algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering temporally-aware reinforcement learning algorithms

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.350351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.617075Z digest=sha256:8bcea80acb0e3fb3bef34586c61c39846033748c088b2c145a4a100b69538fbd

Observation 6bf5b200-36ad-4e13-9410-c95c9baf1e10 · outbound

This paper cites Improving policy optimization with generalist-specialist learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Improving policy optimization with generalist-specialist learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:58.129243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.682731Z digest=sha256:656a5d0857189193f04ff53f1b5f661fe074240f933430d0e450fc0cea7fce13

Observation 3a1a5cb6-507c-49c1-adf6-d89d5bf2585e · outbound

This paper cites Meta Learning Backpropagation And Improving It.

How Should We Meta-Learn Reinforcement Learning Algorithms? Meta Learning Backpropagation And Improving It

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:53.690763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.763840Z digest=sha256:50d62adab58e8d9f7978039a364603254d9206e8a34eb26efd52480aec28af1b

Observation 267ff5a4-13c3-4ac9-98bb-2b92691851ef · outbound

This paper cites Improving generalization in meta reinforcement learning using learned objectives.

How Should We Meta-Learn Reinforcement Learning Algorithms? Improving generalization in meta reinforcement learning using learned objectives

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.828958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.817050Z digest=sha256:073b6db4592622689fc022bc368c590ee3fd08887b602bfd9151ddbac0b5dfcb

Observation 2517e4d0-afda-4d88-927b-a6d3a685f536 · outbound

This paper cites Mirror Learning: A Unifying Framework of Policy Optimisation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Mirror Learning: A Unifying Framework of Policy Optimisation

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.878994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.878994Z digest=sha256:b29da412b2a4547ac987f760fbf681b7fbe377130b91b2a221af263a7921ef43

Observation 853ffeb4-20c3-4278-8903-a0cc889d3f0f · outbound

This paper cites Learning to Optimize for Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Learning to Optimize for Reinforcement Learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:47.922233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:47.922233Z digest=sha256:c74186acd7d5d707d74cce75d59edb9e83ad0bf7c2f8c0c325b2d1bd9a585077

Observation 670b210f-af41-461e-9f14-69132d7b1c52 · outbound

This paper cites gymnax: A JAX -based reinforcement learning environment library, 2022 a.

How Should We Meta-Learn Reinforcement Learning Algorithms? gymnax: A JAX -based reinforcement learning environment library, 2022 a

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.659181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:47.972077Z digest=sha256:36d9a373c5c1aed58ab94fdeea9cb14d64cae477b0ddcc6b5a7cca8d92ac47b4

Observation 1c5b6137-edd7-4d39-98f1-7d410a52472f · outbound

This paper cites evosax: JAX-based Evolution Strategies.

How Should We Meta-Learn Reinforcement Learning Algorithms? evosax: JAX-based Evolution Strategies

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:48.110073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:48.110073Z digest=sha256:6fbd07ad6fe7fd66a6c024f223259671c08b9d38d86cab2080aa9da33238e008

Observation 2a2eb53a-4375-4fbb-9164-1db3ebb27504 · outbound

This paper cites In-context reinforcement learning with algorithm distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? In-context reinforcement learning with algorithm distillation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.496541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.255410Z digest=sha256:d7476ab3d79c4cd28309e7e7c812a0a27bbe3b1ab84e0fd6d4b21318b70e2611

Observation a65675a6-b2a8-42da-91e0-7c76d1806eac · outbound

This paper cites Evolution through Large Models.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolution through Large Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:48.383996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:48.383996Z digest=sha256:ec5f3ce3e710e2f67a3ffa9ccc882ac11bca2f4d58b3dd2b0275545eeb912f3f

Observation ea817af4-650b-4692-a29a-e8901cb12399 · outbound

This paper cites Rediscovering orbital mechanics with machine learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Rediscovering orbital mechanics with machine learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.338948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.486735Z digest=sha256:218c644ea4abad7299ea332db0ef89278fb3f058d72f57ed812c9dc70900c251

Observation 35ca94b3-032a-4957-86d1-36418967adeb · outbound

This paper cites Discovered policy optimisation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovered policy optimisation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.186310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.575011Z digest=sha256:595396afe5626092c9a65c389f740b6c2102ac4cf5952dacee06b45bc000c7bb

Observation 4d9a5bb5-e31a-4381-89b2-6e54d990ab38 · outbound

This paper cites Discovering Preference Optimization Algorithms with and for Large Language Models.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering Preference Optimization Algorithms with and for Large Language Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:48.754244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:48.754244Z digest=sha256:a4575608aa07537c2b5208f0ed9a17d13c15b895feed6ae630122bb5ec6e8501

Observation af015423-f7fa-41e0-9775-a370f6364081 · outbound

This paper cites Behaviour Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Behaviour Distillation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:57.015958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:48.891797Z digest=sha256:279ba1a42cc2f7317e863e53e43b4e4daa946e62efd80e9a58aa51ac87946c42

Observation 0c5f61b8-a66c-4923-a626-748f3e6bdea9 · outbound

This paper cites Understanding plasticity in neural networks.

How Should We Meta-Learn Reinforcement Learning Algorithms? Understanding plasticity in neural networks

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.023177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.023177Z digest=sha256:5e8cea4266b5ced0aab48ea2a1adf66f8bd13138b6ea60214a9daab51f07eeb0

Observation 0336ceb2-4e37-44af-9cc1-59d24cf54dfd · outbound

This paper cites Craftax: a lightning-fast benchmark for open-ended reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Craftax: a lightning-fast benchmark for open-ended reinforcement learning

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.824630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:49.134832Z digest=sha256:7f544f22d9fcfb1fe991901ef6676151b523428fe3df6ae3cea0f9d503a5ea7e

Observation 759aebf1-ca1f-4df1-8d37-a0074a1264b0 · outbound

This paper cites Interpretable machine learning methods applied to jet background subtraction in heavy-ion collisions.

How Should We Meta-Learn Reinforcement Learning Algorithms? Interpretable machine learning methods applied to jet background subtraction in heavy-ion collisions

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.652401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:49.301257Z digest=sha256:416c762c73b8d29f79cd6757932abf068cd59c7903d856580cace576d77ab02a

Observation 6ce3b857-d15e-46fb-98c2-40ec67b58731 · outbound

This paper cites Meta-Learning Update Rules for Unsupervised Representation Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Meta-Learning Update Rules for Unsupervised Representation Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.432448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.432448Z digest=sha256:b37d82051e8b5258aca975a8b4f03ef27c99d9a307ec457bd5db74294627f14d

Observation 0df60976-ee7b-48c8-8d27-81fe624486ab · outbound

This paper cites Understanding and correcting pathologies in the training of learned optimizers.

How Should We Meta-Learn Reinforcement Learning Algorithms? Understanding and correcting pathologies in the training of learned optimizers

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T14:48:53.486491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:49.527570Z digest=sha256:f82e82dc5aa755d259903b8b5a2e9445d2aeb29490391778bd6e2dac1f0de44a

Observation 44f62159-29b7-427f-9199-f6b3536a3d0c · outbound

This paper cites Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves.

How Should We Meta-Learn Reinforcement Learning Algorithms? Tasks, stability, architecture, and compute: Training more effective learned optimizers, and using them to train themselves

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.605196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.605196Z digest=sha256:76b2358cffab4818d6d2d1c9f81d42077c7bff0ebad81ea53f71c96053518ad9

Observation 0610a14e-4f3a-47f2-9159-606f1641d390 · outbound

This paper cites Gradients are Not All You Need.

How Should We Meta-Learn Reinforcement Learning Algorithms? Gradients are Not All You Need

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.755270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.755270Z digest=sha256:f069eee77e54775756c9c393436fd152568cfaba133574e146a87b4cda0f60f5

Observation 76f8e25e-1c33-4720-a7f4-075bb6db140e · outbound

This paper cites VeLO: Training Versatile Learned Optimizers by Scaling Up.

How Should We Meta-Learn Reinforcement Learning Algorithms? VeLO: Training Versatile Learned Optimizers by Scaling Up

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:49.872182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:49.872182Z digest=sha256:b70463ef5a701cdd2f08b73808e207a1fb32995f50f1c4d949d2e97e44722555

Observation a99fa079-00a9-49d9-b6e0-52fbad2f45cb · outbound

This paper cites Self-distillation amplifies regularization in hilbert space.

How Should We Meta-Learn Reinforcement Learning Algorithms? Self-distillation amplifies regularization in hilbert space

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.459946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.019447Z digest=sha256:d159b9180a088aa9f67f5c624a33a0538e644333295aa56dc73ffe40c30d94a0

Observation 72858513-525e-4cb8-8404-30205921394e · outbound

This paper cites Small batch deep reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Small batch deep reinforcement learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.295508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.139633Z digest=sha256:6a42ea2accda11b5278cbbe24ddb36cd7c862cf7339834cbd8eef7dfe2a8b244

Observation afcfcb2f-aebb-41cd-89c1-e2a2d7896257 · outbound

This paper cites Discovering reinforcement learning algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Discovering reinforcement learning algorithms

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:56.134460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.254660Z digest=sha256:542747e5d35c3f8607ed087745be8b657107e764d57f815e4804d65cdf384929

Observation 8075000c-14fd-47d8-9aed-dfdac0ce7168 · outbound

This paper cites Openai o3-mini, January 2025.

How Should We Meta-Learn Reinforcement Learning Algorithms? Openai o3-mini, January 2025

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.965617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.382465Z digest=sha256:5f7ba62ac5244dad99945b44afe0654ccbd164c6062d3796c8ede83846f618a9

Observation 836da2ff-865c-4390-b4c4-0b61b1be5c16 · outbound

This paper cites Stabilizing transformers for reinforcement learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Stabilizing transformers for reinforcement learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.789331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.452623Z digest=sha256:5f9dea2b71564faf5fc7b9d5852d294334368ee7d2a82d0d6db1c46855e595c0

Observation b3da679a-906b-4c9d-810a-8580885f52c3 · outbound

This paper cites Evolving Curricula with Regret - Based Environment Design.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolving Curricula with Regret - Based Environment Design

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.589660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.495021Z digest=sha256:a61432cc90115d835746642df32c50670b983f76c444bbb5bbd7441e447fbdcd

Observation 8f69eaaf-88c3-44ef-a1dd-9c73bda00a80 · outbound

This paper cites Parameter Space Noise for Exploration.

How Should We Meta-Learn Reinforcement Learning Algorithms? Parameter Space Noise for Exploration

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:50.593230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:50.593230Z digest=sha256:1744c4fde520f36db087b7f29da564b26ba4823ab7e4f925faafb3dddc277b86

Observation d359abff-4acd-473c-8099-de350efa7649 · outbound

This paper cites Tunability: Importance of hyperparameters of machine learning algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Tunability: Importance of hyperparameters of machine learning algorithms

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.346301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.706376Z digest=sha256:571b3327a2384ce7a1dab2c185d52617e186c5a396253a4847a2d64248af7b42

Observation a2e2d967-75c0-47e9-8a37-fb87210c6462 · outbound

This paper cites Evolutionsstrategie : Optimierung technischer systeme nach prinzipien der biologischen evolution.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolutionsstrategie : Optimierung technischer systeme nach prinzipien der biologischen evolution

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.153556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:50.800901Z digest=sha256:14ebca08f2eef5595479d4dfc5343c71c895f937ed12022801ee53beba616f15

Observation 9bb36e11-6709-4c14-9b4c-ea544ef8c666 · outbound

This paper cites Pawan Kumar, Emilien Dupont, Francisco J.

How Should We Meta-Learn Reinforcement Learning Algorithms? Pawan Kumar, Emilien Dupont, Francisco J

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:50.860051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:50.860051Z digest=sha256:1b1291b2ae62ec60392c3028a89b321b47e9fcf886a7841c999591b99bf79758

Observation f77d5312-72e7-4aeb-93d5-b885a82980a1 · outbound

This paper cites Policy Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Policy Distillation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:50.969246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:50.969246Z digest=sha256:85cb7f028a912278cc524c4a1af10ca247bf94d48cb9c4780a7e737934767641

Observation 1dfd04f3-40b7-4c24-ac86-7f19ddad9f85 · outbound

This paper cites Evolution Strategies as a Scalable Alternative to Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Evolution Strategies as a Scalable Alternative to Reinforcement Learning

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.080292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.080292Z digest=sha256:a5eacb28422ee70d709442746a4bbc644e125f97704edbbaf25cf2bd2900a0a5

Observation 21e3588a-9d4c-44ad-b16e-341ea7d14213 · outbound

This paper cites Proximal Policy Optimization Algorithms.

How Should We Meta-Learn Reinforcement Learning Algorithms? Proximal Policy Optimization Algorithms

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.130615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.130615Z digest=sha256:3ced2976c7a95ac4f25ce69afe5a40bf6d6e9eddfd7701d64a3471ffd30ff39a

Observation 30db874b-70dd-4448-bbc1-1cb7d6c5b9ef · outbound

This paper cites High-Dimensional Continuous Control Using Generalized Advantage Estimation.

How Should We Meta-Learn Reinforcement Learning Algorithms? High-Dimensional Continuous Control Using Generalized Advantage Estimation

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.229277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.229277Z digest=sha256:10ce1b86c1b94745ea1ef39cb156bf59ab8864a6632423a35c07ec31dd782ad8

Observation 945e19ab-25ed-45ef-90d3-58f813481d85 · outbound

This paper cites The Dormant Neuron Phenomenon in Deep Reinforcement Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? The Dormant Neuron Phenomenon in Deep Reinforcement Learning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.326750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.326750Z digest=sha256:539d34e4c53bb02212cfa3968e7c4c94721a9aa3fac81262cd384871ce6bb3da

Observation d765ea3d-991b-4127-946d-725e108653ad · outbound

This paper cites Distilling Reinforcement Learning Algorithms for In-Context Model-Based Planning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Distilling Reinforcement Learning Algorithms for In-Context Model-Based Planning

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.431138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.431138Z digest=sha256:e4da2d1a608ab87e12bd8f5f9bf4851b5c38fecee4356296d7f341047cbe8ac4

Observation 3f177879-b4e0-4a29-adab-9a732bd859f2 · outbound

This paper cites Generalizable Symbolic Optimizer Learning.

How Should We Meta-Learn Reinforcement Learning Algorithms? Generalizable Symbolic Optimizer Learning

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:55.021535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:51.499138Z digest=sha256:350411a03810aa0a4188a78893018ba3b19c25729944cbe6412d675a87a9dc84

Observation cf324fbd-f556-48f8-ba71-a16468de3b0e · outbound

This paper cites Position: Leverage Foundational Models for Black-Box Optimization.

How Should We Meta-Learn Reinforcement Learning Algorithms? Position: Leverage Foundational Models for Black-Box Optimization

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.550816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.550816Z digest=sha256:6855b05f0d3b0dec2e9c77bc607ecae1e4a93ab1435c4f29c53fe2a42fb2d1bc

Observation 4d92d76f-80a5-4dc5-9d43-0693b7e3567e · outbound

This paper cites Maxinfo RL : Boosting exploration in reinforcement learning through information gain maximization.

How Should We Meta-Learn Reinforcement Learning Algorithms? Maxinfo RL : Boosting exploration in reinforcement learning through information gain maximization

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.840199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:51.635219Z digest=sha256:1e2a1f48c408b64d50b29b1a9f8834d31335537783a37dbc00dacb4e44a54e3a

Observation 9f2931ef-91ed-4c19-9777-88dfd6d5e8e6 · outbound

This paper cites Sutton and Andrew Barto.

How Should We Meta-Learn Reinforcement Learning Algorithms? Sutton and Andrew Barto

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.668222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:51.716178Z digest=sha256:cf83e2a022e4b9a2560bb6b1598d97cef3889bd0e6419cff3cdbf828810b8046

Observation 3acf0046-0b42-41dd-8def-cac57e2a55b3 · outbound

This paper cites Improving deep reinforcement learning by reducing the chain effect of value and policy churn.

How Should We Meta-Learn Reinforcement Learning Algorithms? Improving deep reinforcement learning by reducing the chain effect of value and policy churn

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.798344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.798344Z digest=sha256:2b789b1ab1c20000a09eaa6c2db6f13cc72efb2e9b2787aca1c01855db4bd82e

Observation 7658bc7e-8099-4ad4-8afb-4a365f10451b · outbound

This paper cites MuJoCo : A physics engine for model-based control.

How Should We Meta-Learn Reinforcement Learning Algorithms? MuJoCo : A physics engine for model-based control

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:51.925909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:51.925909Z digest=sha256:26de816a77e584d2ae712b765a16a9c4442636b16db4e6d69a3f53bd412bbdfd

Observation 9d4d48f1-3bb7-4255-a172-039e321357a0 · outbound

This paper cites Deep Reinforcement Learning and the Deadly Triad.

How Should We Meta-Learn Reinforcement Learning Algorithms? Deep Reinforcement Learning and the Deadly Triad

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.022023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.022023Z digest=sha256:4e318287b3816687d6b55411fc870db0f11f6ec64e294213465324acfcadac14

Observation b9178950-d006-48d7-84c0-46e344cce11d · outbound

This paper cites Attention Is All You Need.

How Should We Meta-Learn Reinforcement Learning Algorithms? Attention Is All You Need

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.137791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.137791Z digest=sha256:9c51ad4aff80488b19eedfd8771199efca2ba0a521ee872bca9cd1cd0e55b0bd

Observation 8f07cb6a-ce7b-45e4-8917-29771434a0ae · outbound

This paper cites Dataset Distillation.

How Should We Meta-Learn Reinforcement Learning Algorithms? Dataset Distillation

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.204307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.204307Z digest=sha256:addb1113ab8053a380746e251b4c4b67dcd9281d05c01bfd2b9e8178e4c127d6

Observation 59339402-f08a-4723-bd2c-7434f8185b54 · outbound

This paper cites Natural Evolution Strategies.

How Should We Meta-Learn Reinforcement Learning Algorithms? Natural Evolution Strategies

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:53.127951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.287410Z digest=sha256:c5bcfa7d59da3dc60f2128492adf85750769492e37f6997a4272ef2d08f7c364

Observation 985ed171-fb21-495a-b2b0-84d329f3aa4b · outbound

This paper cites Understanding short-horizon bias in stochastic meta-optimization.

How Should We Meta-Learn Reinforcement Learning Algorithms? Understanding short-horizon bias in stochastic meta-optimization

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.523597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.381492Z digest=sha256:384f2136424d874da39d813f25eb439ed4ef6e2641c831f7a44c2c715fe5d357

Observation 1e7ab7c4-6b5d-4d2a-8ff5-7a5056aac1b8 · outbound

This paper cites MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments.

How Should We Meta-Learn Reinforcement Learning Algorithms? MinAtar: An Atari-Inspired Testbed for Thorough and Reproducible Reinforcement Learning Experiments

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.488734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.488734Z digest=sha256:5289e12bc34536225008816ec178d42e0abc45ff82dfb8a82e8630a3a3128b14

Observation cc47c1d0-9afe-4395-bec3-813cad431c96 · outbound

This paper cites Self-distillation as instance-specific label smoothing.

How Should We Meta-Learn Reinforcement Learning Algorithms? Self-distillation as instance-specific label smoothing

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T14:48:54.343900Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.573001Z digest=sha256:d5adc1aac9a4eec8a51de1e46008362f775a7a598c24f8324ef62550a3191780

Observation 44945f70-98c5-4abf-b02c-65bc7fd38f80 · outbound

This paper cites Symbolic Learning to Optimize: Towards Interpretability and Scalability.

How Should We Meta-Learn Reinforcement Learning Algorithms? Symbolic Learning to Optimize: Towards Interpretability and Scalability

Reference 84

Resolution
verified exact
local_arxiv, observed 2026-08-06T14:48:52.957731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-06T14:48:52.681440Z digest=sha256:4d35f679bf3ca013fa85d04956a00389f34b83df9055e464428a0ee81bb017bc

Observation 41439ad1-985a-4ebb-b5bf-0348a9646f81 · outbound

This paper cites write newline.

How Should We Meta-Learn Reinforcement Learning Algorithms? write newline

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-06T14:48:52.758471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:48:52.758471Z digest=sha256:ef6e55e4641d35ff2d5813953ef9f69b324cc260e75fc791de289ed082599528

Pith citing papers

Observation c32431f8-ff8d-4845-9333-b867b3e88878 · inbound

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback cites this paper.

LEMUR: Learning to Align with Multi-Objective Reinforcement Learning from Preference Feedback How Should We Meta-Learn Reinforcement Learning Algorithms?

Reference 243

Resolution
verified exact
local_arxiv, observed 2026-08-03T04:44:19.150566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-08-03T04:39:32.117730Z digest=sha256:58419cbfa41241453630d878799d9bbbfa37040eb89f6c6519420a5d0bceb8ce