Pith. sign in

Paper Citation Record · LEDGER

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions

As of 10 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 0 inbound Pith citation observations for arXiv:2607.08925.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.08925 v1

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T05:46:03.704125Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

34 of 34 outbound references displayed

  • verified exact4
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3218436a-3f43-4a88-a9ef-215fee392dff · outbound

This paper cites Ames, Samuel Coogan, Magnus Egerstedt, Gennaro Notomista, Koushil Sreenath, and Paulo Tabuada.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Ames, Samuel Coogan, Magnus Egerstedt, Gennaro Notomista, Koushil Sreenath, and Paulo Tabuada

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:8702115b62bc2241801c97c0f4369d6706d23fd441b83ccfd1825bcc2400f456

Observation 374f5067-d0b1-4270-8333-5a612d899c26 · outbound

This paper cites ISBN 9781605585161.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions ISBN 9781605585161

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:7d8ccde9dc20550539e189c60efe39a96fff7bfc11d0dc53f421a29bae8066bc

Observation 1ef3ff0c-df5d-4e33-82bc-da561075d2e6 · outbound

This paper cites Safe Exploration in Continuous Action Spaces.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Safe Exploration in Continuous Action Spaces

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:12372f530d20db5c0b8383fb3bdf673834cf1527d848936a9b44b96a612b1945

Observation 6ff456a5-5f81-491e-a958-8b54531a3919 · outbound

This paper cites Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Mohammadhosein Hasanbeig, Alessandro Abate, and Daniel Kroening

Reference 4

Resolution
verified exact
doi, observed 2026-07-13T05:49:23.740269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:9119fde155c7119efbcc30a91000d440d0fb3fca85909dd306af2428815de977

Observation f32a5ce2-e4b7-44ed-9bd1-c20276b28357 · outbound

This paper cites Shengyi Huang, Rousslan Fernand Julien Dossa, Chang Ye, Jeff Braga, Dipam Chakraborty, Kinal Mehta, and Jo ao G.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Shengyi Huang, Rousslan Fernand Julien Dossa, Chang Ye, Jeff Braga, Dipam Chakraborty, Kinal Mehta, and Jo ao G

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:dcc1a0d464c1261420a7d2519b3268363f46971d9b80af2164f35e6c2966679a

Observation 23c698ac-9b57-4aad-8d25-47142d5a8b12 · outbound

This paper cites Choi, Michael Janner, Claire J.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Choi, Michael Janner, Claire J

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:dd2e8ff2206e439064a0eb04a4f550b39ad8e99b978de75828a2b6785df68494

Observation 561d899a-c361-4a7d-a6b1-30e1e74cffbe · outbound

This paper cites Ashish Kumar, Zipeng Fu, Deepak Pathak, and Jitendra Malik.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Ashish Kumar, Zipeng Fu, Deepak Pathak, and Jitendra Malik

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:56037ce4fc81745ae224a9f711c1b0cd66df51d73c5ee808ed0c0ee5859a3025

Observation 66ea9a09-f858-46da-b009-5dc576c9b0d3 · outbound

This paper cites Robust Recovery Controller for a Quadrupedal Robot using Deep Reinforcement Learning.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Robust Recovery Controller for a Quadrupedal Robot using Deep Reinforcement Learning

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-13T05:49:23.735751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:77ac73e03a8470b908facc4ecf06214fffb0c2e5d6a1127188e6b461a2590cc2

Observation 358108df-7596-4570-8fd8-9900570da405 · outbound

This paper cites Sanmit Narvekar, Bei Peng, Matteo Leonetti, Jivko Sinapov, Matthew E.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Sanmit Narvekar, Bei Peng, Matteo Leonetti, Jivko Sinapov, Matthew E

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:3624774619e0b9848bbb789f4b4b018bb0581d4834558586896e82383f328e5d

Observation dccd144d-696a-4a93-96b5-cab7972b1249 · outbound

This paper cites Xue Bin Peng, Erwin Coumans, Tingnan Zhang, Tsang-Wei Edward Lee, Jie Tan, and Sergey Levine.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Xue Bin Peng, Erwin Coumans, Tingnan Zhang, Tsang-Wei Edward Lee, Jie Tan, and Sergey Levine

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:9c198e90162264bc6822f3831eda3d84a8db24d0731771c43fc7b68eaf0efc45

Observation 2f8ac172-9061-47f4-a0bf-f0aaad2a9e55 · outbound

This paper cites Xun Pua and Majid Khadiv.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Xun Pua and Majid Khadiv

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:27563516f3caf9c9ff772ec27c937c963ea529fe1813b787c938040217ae96c1

Observation 074112a7-aa09-4956-887b-1a17c04d10ff · outbound

This paper cites 2024.10769799.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions 2024.10769799

Reference 12

Resolution
verified exact
doi, observed 2026-07-13T05:49:23.732396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:9efd8f75029a95047f11f504d54e597060879f6c380116e3350efd5e5eea163d

Observation 61f99978-a860-4a60-828d-9d1ff0eeb965 · outbound

This paper cites Learning to walk in minutes using massively parallel deep reinforcement learning.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Learning to walk in minutes using massively parallel deep reinforcement learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:bde1fd21f66f1110a695b2dcd6807de74040ff6b14a8d56db8a3b22bea4fadff

Observation 7f1f459e-8eb2-4ee1-a0ac-7e6defe77733 · outbound

This paper cites Trial without error: Towards safe reinforcement learning via human intervention.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Trial without error: Towards safe reinforcement learning via human intervention

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:4a46b61b7723bf9539453ee92141dddb96fc07faa350147fbc11130d941bc984

Observation 258d8482-dc10-4668-9ef7-1a446d1fb73c · outbound

This paper cites Proximal Policy Optimization Algorithms.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Proximal Policy Optimization Algorithms

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:11addd5a91ea570122e2da4049b099fa87b21c514cffd528b36f4ec461ef5163

Observation 71ae7c4e-07b1-4d91-9455-667a6c5dbe44 · outbound

This paper cites Smith, J.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Smith, J

Reference 16

Resolution
verified exact
doi, observed 2026-07-13T05:49:23.761511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:a62caeeb83434f9f38b1f32d4dec92f5a678fc30f27ff718d60103630e51be0b

Observation 47bd04e3-9676-486c-956f-e380a6bf15f0 · outbound

This paper cites Learning to be Safe: Deep RL with a Safety Critic.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Learning to be Safe: Deep RL with a Safety Critic

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:29a0b9ea9491f871480729925a3354776977a6191170ffebc4030289cf92cab0

Observation 6fbf0afc-f137-487a-8e8e-33a50aeca87e · outbound

This paper cites Sim-to-real: Learning agile locomotion for quadruped robots.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Sim-to-real: Learning agile locomotion for quadruped robots

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:f37eb17fd33ced6b4465f81d6d9f2e275415870dadf5b9c59cf325defd0a87e5

Observation a91ff146-8d42-498a-bb9e-1893a6f293b2 · outbound

This paper cites Reward Constrained Policy Optimization.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Reward Constrained Policy Optimization

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:0aa474d6b0f7388b09fcf2e0b3cb083372bc46e40672a537ff8a63877fc52bdc

Observation 56cdc0bd-9fc4-48a2-a58c-b8b6cc9060b6 · outbound

This paper cites Emanuel Todorov, Tom Erez, and Yuval Tassa.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Emanuel Todorov, Tom Erez, and Yuval Tassa

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:299bd1f1c0de865546894accd2f3be4fadaaf38bb1a164e0f3ba4feac68703c2

Observation 06645f40-c119-4ad3-a0dd-57568a859669 · outbound

This paper cites Mark Towers, Jordan K.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Mark Towers, Jordan K

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:89059f1574a4b5cd181de8bca600459426f2e35cfd4150bcba717a19424cc3b9

Observation ffc01a51-20ed-4a4c-bfd2-176bcd8158ed · outbound

This paper cites URLhttps://doi.org/10.24963/ijcai.2024/913.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions URLhttps://doi.org/10.24963/ijcai.2024/913

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:3f76f11f9ad5fd261e3363a806561836adafc51e78e1ecdc8c73fba2c22f3379

Observation f6355c3b-f2b4-4c47-90bf-b0fef8aed300 · outbound

This paper cites Kevin Zakka, Yuval Tassa, and MuJoCo Menagerie Contributors.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Kevin Zakka, Yuval Tassa, and MuJoCo Menagerie Contributors

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:2e072d3a1f89b9e03b7037ee147f752ab1f9e65ffa9bc9844a5f86c143aaddd4

Observation 4da7ee70-3e09-4f31-b7a0-41fbf86662a1 · outbound

This paper cites URLhttps://doi.org/10.24963/ijcai.2023/763.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions URLhttps://doi.org/10.24963/ijcai.2023/763

Reference 24

Resolution
verified exact
doi, observed 2026-07-13T05:49:23.749250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:13d3e62c412b801898b9ee6891e57e281a6e97644a9c7330193c01661bea46ec

Observation 2fe6394f-3268-4bcf-9cc2-11947860fb17 · outbound

This paper cites Clipping bounds variance downstream of the singularity rather than removing it.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Clipping bounds variance downstream of the singularity rather than removing it

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:dded2bff43422335d4000a4adff97bb9ee640efdcb718fa52725dd3705790d65

Observation 30d9b33d-1e35-4927-a0c1-db54f7ecacd9 · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:1648f33b09ed6ce84ff27c919344db2b52fdd25c47dff57f70294c8612f65a84

Observation 7879240e-bc99-45df-b474-a43106d4cbba · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:20c1467de7fffab879fe031e226655d37380322dc9ea1ef5e3299aa84d428974

Observation 281778e4-fe75-4cbe-b09f-56f2c4bb271f · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:a973282fb0b9ba9799f3bd1f9ba8e21b4de2450b0de536a75bb8e6d02c6eff7e

Observation 9722ad14-7b7e-44dd-89cf-5a72673eead8 · outbound

This paper cites CPO and PPO-Lagrangian additionally carry the constraint hyperparameters their objective requires (Section C.3).

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions CPO and PPO-Lagrangian additionally carry the constraint hyperparameters their objective requires (Section C.3)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:563f354a9059f1a4cc5dc82aff071ed580e1a4bac19867a84c4b026b35f0a6fe

Observation 36000acf-9365-4be8-a4b0-19692dbfcb5d · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:016ae4602541d339d372f117fdd0e1165413053c3f2bb3a2d955a47f94eadbde

Observation 5209fcc6-7cba-41d5-8ac8-b87faf89dc3e · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:4b2827502445e3d23a3249ac2c88d2fd953eb5f3b9d056a251347d26b856b047

Observation ea2c3ada-c1ea-42b8-a5d1-ffe5a9adc628 · outbound

This paper cites global”). An alternative “per-segment.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions global”). An alternative “per-segment

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:9530bb5619d721b85716f372dc90b6b3600e3dad8bb99c269ffc6ac7a7c857de

Observation b2ce74f8-3002-4757-8960-f03e60a71d89 · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:f2e77ef5811181c2cbf0f14c6e51d943da1eb69c2f9ab679f22e36a6e7fba4ad

Observation a347c806-4456-4019-bf06-729146d76549 · outbound

This paper cites an unresolved cited work.

SafeExplorer: An Unbiased Policy Gradient for Reinforcement Learning with Recovery Interventions Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-13T05:46:03.704125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T05:46:03.704125Z digest=sha256:66605a2ca21932ffa524cfd45e93277963d0327697c72565b314013ddda41f65

Pith citing papers

No inbound Pith citation observations are available.