Pith. sign in

Paper Citation Record · LEDGER

What Matters for Simulation to Online Reinforcement Learning on Real Robots

As of 8 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 5 inbound Pith citation observations for arXiv:2602.20220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.20220 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T21:36:20.594100Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T00:54:12.099045Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T19:30:07.865484Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved64
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0f41b2c-084d-4af3-abd4-5bde4a65733e · outbound

This paper cites Sutton and Andrew G.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Sutton and Andrew G

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:14.774918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:14.774918Z digest=sha256:db2441e1674e1f86b7d8e1df18e3c72a34d96beed157e647d89cdccd26ada58d

Observation 199c9738-505b-4318-b676-c2f9038dd93e · outbound

This paper cites Chapman & Hall/CRC Artificial Intelligence and Robotics Series.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Chapman & Hall/CRC Artificial Intelligence and Robotics Series

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:14.912650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:14.912650Z digest=sha256:36d5e4d0f2b94c96f103a47a05fdaa9cfd626a25fbf427ff9029085f806911c6

Observation ea16bc4e-da9f-45d8-94a4-48aa0f8cc747 · outbound

This paper cites Learning agile and dynamic motor skills for legged robots.Science Robotics, 2019.(Cited on pages 1, 4, and 16).

What Matters for Simulation to Online Reinforcement Learning on Real Robots Learning agile and dynamic motor skills for legged robots.Science Robotics, 2019.(Cited on pages 1, 4, and 16)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:15.038547Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:15.038547Z digest=sha256:bed679b52631c62e0da3424e744ce985034fc965e57c5b4681ba0ee657b83531

Observation c4db4aa6-7aa6-472e-8232-0ea10948f984 · outbound

This paper cites IndustReal: Transferring Contact-Rich Assembly Tasks from Simulation to Reality.

What Matters for Simulation to Online Reinforcement Learning on Real Robots IndustReal: Transferring Contact-Rich Assembly Tasks from Simulation to Reality

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:15.185043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:15.185043Z digest=sha256:655edd5974cd69dcf0f10d383abd8e2d1a4b9b6bcf7bbc60dd9c0766ed558856

Observation 5b539d2b-dbe1-4d68-b143-319871fdcb19 · outbound

This paper cites Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 2025.(Cited on page 1).

What Matters for Simulation to Online Reinforcement Learning on Real Robots Diffusion policy: Visuomotor policy learning via action diffusion.The International Journal of Robotics Research, 2025.(Cited on page 1)

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:15.320459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:15.320459Z digest=sha256:de0435febf5eea3c36fb9656d832aaa75ddb528c69be8b337b372bfdc85ec576

Observation 3b468079-d591-46a5-b6b9-a3df30d7e2e5 · outbound

This paper cites BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion.

What Matters for Simulation to Online Reinforcement Learning on Real Robots BeyondMimic: From Motion Tracking to Versatile Humanoid Control via Guided Diffusion

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:15.439847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:15.439847Z digest=sha256:602b580f8a1988b433cb7152ce151fe7d1235ac76f4369384b7c61334adf7fcf

Observation a8bc4987-e5a0-4568-9286-8af018904a9e · outbound

This paper cites an unresolved cited work.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:15.538983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:15.538983Z digest=sha256:0d5ed659c7709673cbd91a28bc085a3a688bbde04ab1c1b9f6da584ba4d79fc0

Observation 5045fc0f-ad30-4576-9aa4-c0617cf410e8 · outbound

This paper cites Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:15.643230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:15.643230Z digest=sha256:aa84059a0872a819e3d4f82be4ef9c4b6ebb66515dfd5bcd2d1f518204f1cb68

Observation e0fe0836-7a91-45e0-b36e-41a57adab712 · outbound

This paper cites MuJoCo Playground.

What Matters for Simulation to Online Reinforcement Learning on Real Robots MuJoCo Playground

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:15.812654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:15.812654Z digest=sha256:7ea81bf35435a4b1e2743fb0a007f5f94d3bd351b7328e18bec57ec10a9fde89

Observation 18d1731a-ce5b-4d04-be6f-fa854c6de330 · outbound

This paper cites Addressing function approximation error in actor-critic methods.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Addressing function approximation error in actor-critic methods

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:15.940499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:15.940499Z digest=sha256:9338ab87e104211969d2d46e159452757bccb52423af9de29e1d92c583de0b0a

Observation 2c5b7039-aa41-41da-8867-96c6d4ea6998 · outbound

This paper cites Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 2013.(Cited on page 2).

What Matters for Simulation to Online Reinforcement Learning on Real Robots Reinforcement learning in robotics: A survey.The International Journal of Robotics Research, 2013.(Cited on page 2)

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.060341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.060341Z digest=sha256:677afff396e1fd2de2d5904f1c38b5eab38d84f27a86f26f4cb30626de66df7f

Observation e992df3b-93cf-4d08-bb18-e23315fe83f2 · outbound

This paper cites Learning to Walk in the Real World with Minimal Human Effort.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Learning to Walk in the Real World with Minimal Human Effort

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.146523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.146523Z digest=sha256:9da03b7bde9c0d5cae0c9bbb523b1eb66c3f36d100012152dddfae87b078db85

Observation 35b970f1-1a5d-4fb4-8434-4fb20438dc4d · outbound

This paper cites COG: Connecting New Skills to Past Experience with Offline Reinforcement Learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots COG: Connecting New Skills to Past Experience with Offline Reinforcement Learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.288114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.288114Z digest=sha256:dc2f4de093b05d1475fbf17e7ead285b1b7639dd40ee3696f148638ceeb4e2ea

Observation e663363f-be65-4594-be85-639a72cf995d · outbound

This paper cites AWAC: Accelerating Online Reinforcement Learning with Offline Datasets.

What Matters for Simulation to Online Reinforcement Learning on Real Robots AWAC: Accelerating Online Reinforcement Learning with Offline Datasets

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.362061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.362061Z digest=sha256:84306188028e78dc92c10ae90b04f7a56746fc23aa72b063eda6b06e3bf6effd

Observation 49603da0-a967-47cc-b000-7cf3056edd41 · outbound

This paper cites Daydreamer: World models for physical robot learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Daydreamer: World models for physical robot learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.427429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.427429Z digest=sha256:ee0867f1d67bee275fb71999caaf5b6a7e78bea5334abcf477e1518a56789b06

Observation 49bdb634-e974-4521-9030-eea7149544a2 · outbound

This paper cites Finetuning offline world models in the real world.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Finetuning offline world models in the real world

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.486641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.486641Z digest=sha256:ecdf38ac3f63f0ec424b0d32048500ff4614cf753e4dcebbd6ede253b4c2c204

Observation 0afc3af8-ccc3-437b-9367-d565cf119dc2 · outbound

This paper cites Efficient online reinforcement learning fine-tuning need not retain offline data.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Efficient online reinforcement learning fine-tuning need not retain offline data

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.569698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.569698Z digest=sha256:f85e84e173e99771a4c7bae6a700847d16c89374c2e19c02554e89ec429ac025

Observation 4a5941a7-1329-4d9c-825f-63ffa8271895 · outbound

This paper cites Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning.Science Robotics, 2025.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Precise and dexterous robotic manipulation via human-in-the-loop reinforcement learning.Science Robotics, 2025

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.655141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.655141Z digest=sha256:7a8c570563eead51254dff76310bd47796d3f787f484b09422bad1ebddee415b

Observation aeb0df54-3849-4454-b1d4-7b27d093f396 · outbound

This paper cites How to train your robot with deep reinforcement learning: lessons we have learned.

What Matters for Simulation to Online Reinforcement Learning on Real Robots How to train your robot with deep reinforcement learning: lessons we have learned

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.710341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.710341Z digest=sha256:a9e7fa100e9eb147742f953c264f2903edbb8c9019851838d8874c44d2f15e5f

Observation 35779e2c-c688-4e04-a37c-181e12850100 · outbound

This paper cites Replay across experiments: A natural extension of off-policy RL.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Replay across experiments: A natural extension of off-policy RL

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.769044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.769044Z digest=sha256:de4e9ab7172ca8f720197d244785c8b6f62e00054570b3e6fc701f62bf6e4a28

Observation fcb15989-78d9-47cc-8d20-4a9197518654 · outbound

This paper cites Rapidly adapting policies to the real-world via simulation- guided fine-tuning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Rapidly adapting policies to the real-world via simulation- guided fine-tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.822116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.822116Z digest=sha256:2d674d31510ec6445ea5030a8c0e24dec1c3c9ee6ca497bc95e946815386b5e5

Observation eae6870c-b5d4-46d7-9f59-7c5d484b0d63 · outbound

This paper cites Legged robots that keep on learning: Fine-tuning locomotion policies in the real world.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Legged robots that keep on learning: Fine-tuning locomotion policies in the real world

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.872595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.872595Z digest=sha256:601189ee6da12f9ec7f76ad0421a4fb5af592e2da78b6e9a1d87f14861a22169

Observation ec06c701-36de-4cb1-b092-0e80c8cf97d6 · outbound

This paper cites A Walk in the Park: Learning to Walk in 20 Minutes With Model-Free Reinforcement Learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots A Walk in the Park: Learning to Walk in 20 Minutes With Model-Free Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:16.982549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:16.982549Z digest=sha256:5782b535f0f11fe5cdfd12cc9936bc4a0cdd6dd78337c9db8c99db163c53405d

Observation 8c8dacd8-67b1-44ed-a07e-b3560b586e1c · outbound

This paper cites John Wiley & Sons, 2014.(Cited on page 3).

What Matters for Simulation to Online Reinforcement Learning on Real Robots John Wiley & Sons, 2014.(Cited on page 3)

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.066546Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.066546Z digest=sha256:4bef6079b58dedf0ef0d73e5d39a28c30bf20af051dbdb525421c6aeaf69a9ca

Observation d2ec8e4b-ec90-42e5-8169-8eeb7f2886a8 · outbound

This paper cites Bertsekas.Dynamic Programming and Optimal Control.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Bertsekas.Dynamic Programming and Optimal Control

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.149872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.149872Z digest=sha256:a8b0d80319880e700cac717e85b425f329b6eab5a5357fc11c38ca33496c423d

Observation 95118829-be51-4362-8de9-a2b444fb34b8 · outbound

This paper cites Self-improving reactive agents based on reinforcement learning, planning and teaching.Machine Learning, 1992.(Cited on page 3).

What Matters for Simulation to Online Reinforcement Learning on Real Robots Self-improving reactive agents based on reinforcement learning, planning and teaching.Machine Learning, 1992.(Cited on page 3)

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.212167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.212167Z digest=sha256:b796206ca714a3f14d4b75926399e190a3cb23fd1f49cfba5d068a23ca1719ab

Observation 8a6cbda2-8f21-4002-a97d-9f3f62d7fdef · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Playing Atari with Deep Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.273492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.273492Z digest=sha256:b990afbfe0d2bc48cfe46f3ddcb80783b023d40304163c7aac0c8140d87ee991

Observation d9bbb50d-629c-4143-ad84-c34ed3a9a806 · outbound

This paper cites Leave no trace: Learning to reset for safe and autonomous reinforcement learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Leave no trace: Learning to reset for safe and autonomous reinforcement learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.343917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.343917Z digest=sha256:25e4114554eeba2914d68bc1d4079a4a0515e346998a0d7cd05b544b30923149

Observation ae849b38-2b39-446c-952f-017029ec61da · outbound

This paper cites Autonomous reinforcement learning via subgoal curricula.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Autonomous reinforcement learning via subgoal curricula

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.434804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.434804Z digest=sha256:9bcb00a0522d2361b659da096fb47e966595ba74ac834ba5b7c667d1026f7442

Observation 9ea8fff8-18b2-415e-b6f4-9aab26499aa2 · outbound

This paper cites Autonomous reinforcement learning: Formalism and bench- marking.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Autonomous reinforcement learning: Formalism and bench- marking

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.498275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.498275Z digest=sha256:ad18fe3b8b3d3b7ca6edce1f81ade4ad92d8204e2ed4d09fd6eef13f5f2bdbd3

Observation 5ef68a0d-8401-45e4-b983-2aca514542f6 · outbound

This paper cites A state-distribution matching ap- proach to non-episodic reinforcement learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots A state-distribution matching ap- proach to non-episodic reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.537624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.537624Z digest=sha256:3376d09d2bbc4a9d0a13cc33c8a7e9db5d00e68bddfb459be24344ed3e131664

Observation 6543067f-033e-4fc1-bad9-9fa9fb4a43d1 · outbound

This paper cites Conservative q-learning for offline reinforcement learning, 2020.(Cited on page 3).

What Matters for Simulation to Online Reinforcement Learning on Real Robots Conservative q-learning for offline reinforcement learning, 2020.(Cited on page 3)

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.613572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.613572Z digest=sha256:e16671f9dafeaefae9db3787880f41cc476906c2c81e752cbf2a4df0c1f7f5c6

Observation 728fc3ef-3fa6-41dc-a5c6-8930da6f9094 · outbound

This paper cites Mopo: Model-based offline policy optimization.Interna- tional Conference on Neural Information Processing Systems, 2020.(Cited on page 3).

What Matters for Simulation to Online Reinforcement Learning on Real Robots Mopo: Model-based offline policy optimization.Interna- tional Conference on Neural Information Processing Systems, 2020.(Cited on page 3)

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.697358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.697358Z digest=sha256:677a8a77780f495ee39a69c86d8a1aa6defd832602878ac50f9f96bdb77adb4b

Observation e7485183-7f90-4f72-8341-4a8b32ff8772 · outbound

This paper cites Offline robotic world model: Learning robotic policies without a physics simulator.arXiv preprint arXiv:2504.16680, 2025.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Offline robotic world model: Learning robotic policies without a physics simulator.arXiv preprint arXiv:2504.16680, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.782001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.782001Z digest=sha256:3cc7758502d6a240cd4e66535a928254e84ce6a27a1656b85428638a13c21a37

Observation 5bdaafc0-11e2-4038-b166-0913b520896b · outbound

This paper cites Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Isaac Lab: A GPU-Accelerated Simulation Framework for Multi-Modal Robot Learning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.872048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.872048Z digest=sha256:28e115814d50bbbcaae37167a1bf82c5e6bf2d185e9eff1126288f250b43f3b6

Observation 2b4a616a-9301-4e32-a7e2-17f787cf13d0 · outbound

This paper cites Domain randomization for transferring deep neural networks from simulation to the real world, 2017.(Cited on page 4).

What Matters for Simulation to Online Reinforcement Learning on Real Robots Domain randomization for transferring deep neural networks from simulation to the real world, 2017.(Cited on page 4)

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:17.936062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:17.936062Z digest=sha256:485ad8719e206444b43e8e7e8a660ab82263f277ca2e8719a5b07ac7369dd121

Observation 09b1d3c5-c5b2-487b-abda-34881eb05015 · outbound

This paper cites Proximal Policy Optimization Algorithms.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Proximal Policy Optimization Algorithms

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:18.048453Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:18.048453Z digest=sha256:1b408eeb0377890998e5c4abc6701ace1bd63f397f74a980946a2b1bfcf9f86b

Observation a7f03011-c7d5-40d4-8012-4ae88726a5be · outbound

This paper cites Continuous control with deep reinforcement learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Continuous control with deep reinforcement learning

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:18.148009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:18.148009Z digest=sha256:4a712f032f203d22397bb360f1456fc2e447f7383d29dbe0b60dfb173840f015

Observation 12607e4a-8e25-4eb0-b54c-f04c69513eaa · outbound

This paper cites Maximum a Posteriori Policy Optimisation.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Maximum a Posteriori Policy Optimisation

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:18.267758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:18.267758Z digest=sha256:c6eb0c516c7a718222a7cc645779ffc599d53d96ea8d88a9191f0b5e8f0413a0

Observation ec0e7c28-e502-4ccc-b5c7-65ea28f899d9 · outbound

This paper cites Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:18.389537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:18.389537Z digest=sha256:fe0431eb05dc5d83039f6b4750ed7ae4771f61434476e5f5962e611c2433dc90

Observation 8b9944c4-d5cd-40ba-b0da-54d8bf9e1c1b · outbound

This paper cites Getting sac to work on a massive parallel simulator: An rl journey with off-policy algorithms.araffin.github.io, Feb 2025.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Getting sac to work on a massive parallel simulator: An rl journey with off-policy algorithms.araffin.github.io, Feb 2025

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:18.524919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:18.524919Z digest=sha256:daf234d060fe2c07dad252683ac278786b539d63d98219609774aa01b7fdce94

Observation f91f8002-d474-4132-b42d-28742e810476 · outbound

This paper cites Deterministic policy gradient algorithms.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Deterministic policy gradient algorithms

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:18.649530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:18.649530Z digest=sha256:714a0390a3bf179523bc7249a6238e74b358e5c6dc174fa57a2aa1b711b3aa05

Observation e38426cc-261d-4af9-a24b-a1b26f067720 · outbound

This paper cites Equivalence Between Policy Gradients and Soft Q-Learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Equivalence Between Policy Gradients and Soft Q-Learning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:18.708908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:18.708908Z digest=sha256:5e8a1ab4a27bc0b6030fdd441ac95750a31b0d7ff19e8744d9ac0c46abb9d584

Observation f8038379-071f-4908-8ce9-0964fd7587bb · outbound

This paper cites Approximately optimal approximate reinforcement learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Approximately optimal approximate reinforcement learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:18.776319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:18.776319Z digest=sha256:8a512c6986c35ffc91e79a4ad8bf44613d725707d0fb464b53a4bc37ec0cddf8

Observation 99aa29d9-ae53-43fa-a639-6ff2d9fda713 · outbound

This paper cites Efficient online reinforce- ment learning with offline data.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Efficient online reinforce- ment learning with offline data

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:18.847505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:18.847505Z digest=sha256:93e3618dbd360c744eba682824ec472297a55c17899e4452172d789e6915f9d2

Observation 0046c12a-d261-42b2-8c9d-f213ce832d7c · outbound

This paper cites Hybrid RL: Using both offline and online data can make RL efficient.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Hybrid RL: Using both offline and online data can make RL efficient

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:18.979143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:18.979143Z digest=sha256:950ee45aa70599714c57ba07648850b4f6bc79cffcf9105acba3cb978a827ae0

Observation e3c04ac4-a1e0-4474-bb0e-76dad9d3473f · outbound

This paper cites Steering your generalists: Improving robotic foundation models via value guidance.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Steering your generalists: Improving robotic foundation models via value guidance

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.076082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.076082Z digest=sha256:6bf8c61d396fb83d8d7fa1ae34ac830d746739fd6df7ce7323f40430010763ad

Observation 2755953c-b859-48c9-a2d3-eb29f534a7b7 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.144684Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.144684Z digest=sha256:ed5898b7e5fab579df159c00ffcbb79ff800c4b0a243b2e82b7a2449651e9f40

Observation 17a6ed49-e9d1-4538-9dd9-4301de990671 · outbound

This paper cites When to trust your model: Model-based policy optimization.Advances in Neural Information Processing Systems, 2019.(Cited on page 6).

What Matters for Simulation to Online Reinforcement Learning on Real Robots When to trust your model: Model-based policy optimization.Advances in Neural Information Processing Systems, 2019.(Cited on page 6)

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.267132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.267132Z digest=sha256:eb517b4b3193c42fe3b50ca9cb40a03eebe22675412bafb635c0be0fe0244665

Observation 98f63cc1-5242-4820-86e2-1075fe6e3200 · outbound

This paper cites Randomized Ensembled Double Q-Learning: Learning Fast Without a Model.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Randomized Ensembled Double Q-Learning: Learning Fast Without a Model

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.350061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.350061Z digest=sha256:5da969ae7b95a004454526ae862e6754baa7e75442673834124dff0d5d506d67

Observation 0883ae8d-da4f-4058-987b-f6103e43e0ee · outbound

This paper cites Overestimation, overfitting, and plasticity in actor-critic: the bitter lesson of reinforcement learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Overestimation, overfitting, and plasticity in actor-critic: the bitter lesson of reinforcement learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.444034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.444034Z digest=sha256:d3a7975f32cdf1c356dea93563cc12be028ca52e4ed7b74b98c9de7c5d1f7866

Observation ae73a42f-fc73-455f-8f40-635d2680f5a4 · outbound

This paper cites On actor-critic algorithms.SIAM journal on Control and Optimization, 2003.(Cited on page 6).

What Matters for Simulation to Online Reinforcement Learning on Real Robots On actor-critic algorithms.SIAM journal on Control and Optimization, 2003.(Cited on page 6)

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.522344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.522344Z digest=sha256:f5f6262d7f46df0a3a58a503d5236d36c37b4f4534259680aa2c32ace28937f9

Observation f97fd0ea-306a-40e3-b65a-d84a07841582 · outbound

This paper cites Convergence rate of linear two-time-scale stochastic approximation.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Convergence rate of linear two-time-scale stochastic approximation

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.606910Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.606910Z digest=sha256:3f2bfdb9a50259c529a41a41ab99368758bad8c92eb44ccc7867d3996492e3db

Observation 8803372a-5837-4c18-81e8-79baf7ef0c02 · outbound

This paper cites Springer.(Cited on page 6).

What Matters for Simulation to Online Reinforcement Learning on Real Robots Springer.(Cited on page 6)

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.714180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.714180Z digest=sha256:ceb8efc2f0780481356b687c4a72d6695d7d6d18605eab12bf0d654648b1976d

Observation 694a49a7-0eed-41b5-b595-255e0af53ddb · outbound

This paper cites Amz driverless: The full autonomous racing system.Journal of Field Robotics, 2020.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Amz driverless: The full autonomous racing system.Journal of Field Robotics, 2020

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.793441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.793441Z digest=sha256:72dc9433c0342412fff6de175b7af7b1fdfea08d63febe2c573297eacafdfb3d

Observation b18bc99e-e26c-4b4f-afcb-bc9e5ec90bd2 · outbound

This paper cites Bigger, regularized, optimistic: scaling for compute and sample efficient continuous control.Advances in neural information processing systems, 2024.(Cited on page 6).

What Matters for Simulation to Online Reinforcement Learning on Real Robots Bigger, regularized, optimistic: scaling for compute and sample efficient continuous control.Advances in neural information processing systems, 2024.(Cited on page 6)

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.884990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.884990Z digest=sha256:f4daa07d731cfcd97efc50532354c5cb90a8d740eb1c36bb5c68a191b86cf5db

Observation 80659e87-a4d6-4f0c-8f44-c83961372503 · outbound

This paper cites Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Mastering Visual Continuous Control: Improved Data-Augmented Reinforcement Learning

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:19.987578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:19.987578Z digest=sha256:584c522034c095d8a0c25443366358bc6e9a0ec8fb73174980981344a363b037

Observation a6637cae-b873-4075-88ed-f1a08803f236 · outbound

This paper cites Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Daniel Freeman, Erik Frey, Anton Raichuk, Sertan Girgin, Igor Mordatch, and Olivier Bachem

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:20.045335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:20.045335Z digest=sha256:167d83a5b8f5ee7ec5885b75d8a7f727131277300cca7b3168301016bdeec3f5

Observation e7321c3d-81bd-4178-9007-74804c64cf8d · outbound

This paper cites an unresolved cited work.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:20.218776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:20.218776Z digest=sha256:cad03d1df2ffaf72c5a6901fdc237274a70cc94b3ea01c703f8ae5ef30ed70fe

Observation 9afc4bba-b672-4573-97ea-b3b0ba22dc8a · outbound

This paper cites Parallel𝑞- learning: Scaling off-policy reinforcement learning under massively parallel simulation.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Parallel𝑞- learning: Scaling off-policy reinforcement learning under massively parallel simulation

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:20.299420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:20.299420Z digest=sha256:25aa017eae8cae42a6f26d957284cfb9032410d56b70517e50953efae6e52eea

Observation 9db694ae-f907-40ab-903c-7895f5901824 · outbound

This paper cites FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control.

What Matters for Simulation to Online Reinforcement Learning on Real Robots FastTD3: Simple, Fast, and Capable Reinforcement Learning for Humanoid Control

Reference 61

Resolution
malformed identifier
no resolver link, observed 2026-08-02T21:36:20.383255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:20.383255Z digest=sha256:290755fd95c4a6d48d05e21913b4d245c14c8b75a3727c7ec817b0af1bf937fc

Observation 6999d73e-3b90-4573-9448-2d9827b80b98 · outbound

This paper cites an unresolved cited work.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Unresolved cited work

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:20.445601Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:20.445601Z digest=sha256:90bf959eca4a6c298fc351f63ae0ca9b8dbd26438e61e542a47f092cb7907e46

Observation 5446a11e-446e-4b75-a361-2cd84843bc7c · outbound

This paper cites an unresolved cited work.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:20.529448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:20.529448Z digest=sha256:4204e9948cefc851577699a84de236904de3687e8fe9b62fbd10c5f8da983bdc

Observation 339d6e9c-c26e-4d28-a6e1-2bda3738025a · outbound

This paper cites These issues are easy to overlook but they change learning dynamics and final performance.

What Matters for Simulation to Online Reinforcement Learning on Real Robots These issues are easy to overlook but they change learning dynamics and final performance

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:20.594100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:20.594100Z digest=sha256:fdf00dddd121c0b1464ea51b95d9d96c8bcc4fb9602d4e4128b288d54a8de842

Observation fe94c184-2dd8-4ba4-ae7f-302b94df8581 · outbound

This paper cites an unresolved cited work.

What Matters for Simulation to Online Reinforcement Learning on Real Robots Unresolved cited work

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T21:36:20.132246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T21:36:20.132246Z digest=sha256:81156372622dc15ba4f18ee02aa795824fb53916e61c3a1838340ac0b9653166

Pith citing papers

Observation 50d12e37-728d-474c-9076-3b96a36dfde5 · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? What Matters for Simulation to Online Reinforcement Learning on Real Robots

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-24T01:23:06.202989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T05:33:19.038889Z digest=sha256:7e0ae308212032f45ce9c5aac1286199828d8f107aa5069b8c6cb111ddb3bdce

Observation 4164f62f-e63c-4b79-acf8-92dde9862f29 · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? What Matters for Simulation to Online Reinforcement Learning on Real Robots

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-24T01:23:06.202989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T06:27:38.643667Z digest=sha256:542205b6dfb34dfd85063bb933e6a22384c8d0276f7a6bb2417988128c85a3c8

Observation 4f4131a2-19b8-4be1-b740-ba5a28bc8d2f · inbound

When Does Non-Uniform Replay Matter in Reinforcement Learning? cites this paper.

When Does Non-Uniform Replay Matter in Reinforcement Learning? What Matters for Simulation to Online Reinforcement Learning on Real Robots

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-24T01:23:06.202989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T23:04:12.943222Z digest=sha256:1af9e8a30f6a68f131b5b45a9fffc6e64603b4ad36a0282efe840e1338c167d7

Observation c5dd250e-99a1-4000-9ec4-69b9c9d8cff6 · inbound

Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion cites this paper.

Bridging the Gap: Enabling Soft Actor Critic for High Performance Legged Locomotion What Matters for Simulation to Online Reinforcement Learning on Real Robots

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-24T01:23:06.202989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T00:54:12.099045Z digest=sha256:68a1f7e0094a5f4a1a844b117d742994a9988bcd6da28eb379245323c16df2c3

Observation a8e4ee06-e705-480f-8093-511a101ddd32 · inbound

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors cites this paper.

Beyond One-Size-Fits-All: Diagnosis-Driven Online Reinforcement Learning with Offline Priors What Matters for Simulation to Online Reinforcement Learning on Real Robots

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-24T01:23:06.202989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T21:15:07.735336Z digest=sha256:bcc36ad56bc48a4f1a97256db3b4dd76ecb97bcb67a51c0b43154dacac3f0468