Pith. sign in

Paper Citation Record · LEDGER

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation

As of 16 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2412.11138.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11138 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:20:56.545141Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact4
  • verified fuzzy28
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0800d1b-1221-4dde-9f03-2fb971020bf8 · outbound

This paper cites write newline.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.249835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.249835Z digest=sha256:233c5838161925730ea9326f994ecee6ebe47a18e05bdf736f606e6b43935516

Observation a360c338-1592-4384-8cbe-68d871306fc0 · outbound

This paper cites Constrained policy optimization, 2017.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained policy optimization, 2017

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.460464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.255303Z digest=sha256:140770dff27ea4c13d109aa07aa0082e66d0010e9f3c918a6dd7353a4212ffce

Observation a71c32c2-d606-4952-b20a-8bda488a1571 · outbound

This paper cites Constrained Markov decision processes, volume 7.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained Markov decision processes, volume 7

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.446656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.260145Z digest=sha256:0e3ab102494a4081e7d753ca3c0252db2cb4d04f1d062a534274a4d7e835ad5a

Observation d9d7241f-bd17-43b4-910c-2346b9547840 · outbound

This paper cites Constrained Policy Optimization via Bayesian World Models.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained Policy Optimization via Bayesian World Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.263978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.263978Z digest=sha256:44706ba7e6875b8a1fef6b2286b80102f54a37a2a8a9fd14db478e5bfe32674c

Observation b5b70bcd-f0d2-40d2-b46a-593fd00d3b87 · outbound

This paper cites Robots that interact with humans: a review of safety technologies and standards.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Robots that interact with humans: a review of safety technologies and standards

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.433993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.269227Z digest=sha256:e86a1af95a1f4c91d1f42dc9fcc9c2bb7e14e1c2aa25c55470053f274b3d52f2

Observation cc377145-760d-4253-b4b2-f0818b5ed36c · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Risk-constrained reinforcement learning with percentile risk criteria

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.273249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.273249Z digest=sha256:71acfb25d89a1faba281dd83af3c699c3ea8f419b68373856b2fac908d895e92

Observation feb4e2bc-1b61-424a-980b-d7003ef5fec4 · outbound

This paper cites Model-Augmented Actor-Critic: Backpropagating through Paths.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Model-Augmented Actor-Critic: Backpropagating through Paths

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.278127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.278127Z digest=sha256:72a584161466dd575bd06adc38dbf614d0be6479ba19da30b12d4b1c33965cfc

Observation 940c8509-2d63-4ecf-9739-8d5275988a23 · outbound

This paper cites Augmented proximal policy optimization for safe reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Augmented proximal policy optimization for safe reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.408861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.283465Z digest=sha256:6c8f0e59bc7beaeb14359340ad56d6197e43e522d726506b10954586b6ff7d68

Observation fa0132e8-2c18-4e22-a256-bf12ef92e5d8 · outbound

This paper cites Safe RLHF : Safe reinforcement learning from human feedback.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Safe RLHF : Safe reinforcement learning from human feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.290448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.290448Z digest=sha256:e35b760cafa2075a9a0f4de992a76f1a2ed17cab7efb44148f3b09123df43432

Observation 69df94db-b3f8-4867-9d6b-23b7f539f038 · outbound

This paper cites A differentiable physics engine for deep learning in robotics.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A differentiable physics engine for deep learning in robotics

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.388846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.296012Z digest=sha256:f368b70c008457e4613150e369be29cade395d9c5b005d3db3d58958f621a47c

Observation 7e2fba4c-a9f7-4be6-bcaa-350a35f3c12b · outbound

This paper cites A., Farouk, H., and Mofreh, E.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A., Farouk, H., and Mofreh, E

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.376244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.302087Z digest=sha256:02d19fba5000d70dd80bb790cd96420e7de850d001ddd52c340081ec69f59cbc

Observation 7ec16299-e684-4742-9a86-ad0a989acecd · outbound

This paper cites D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.364745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.306751Z digest=sha256:5b35e4a1f79c0ba3f23e0ae559108129cb17ef362d7c771f6f954b5670ff63c9

Observation 998678a0-0f6a-4dc2-8e83-c9718926bce9 · outbound

This paper cites A Review and Outlook on Predictive Cruise Control of Vehicles and Typical Applications Under Cloud Control System.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A Review and Outlook on Predictive Cruise Control of Vehicles and Typical Applications Under Cloud Control System

Reference 13

Resolution
verified exact
doi, observed 2026-08-11T15:20:56.597981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.311323Z digest=sha256:1badc38c2632c96e36698e3cbe075ea36338ba28df30664d571af6e716759235

Observation 6da0f8c1-27d8-4aa7-8fb6-e54e96503af9 · outbound

This paper cites and Fern \'a ndez, F.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Fern \'a ndez, F

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.318141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.318141Z digest=sha256:8ffc7fe13fc25dc2ccece9fd482f77f37412a3021e6df6c691774ac3fe5ec6ba

Observation 2c72b972-cdd9-4fbb-9475-3d4fc74d445a · outbound

This paper cites Bullet-safety-gym: A framework for constrained reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Bullet-safety-gym: A framework for constrained reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.345538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.322058Z digest=sha256:03e9646e37cb2032b2543739cd9975e1c81708dc49f63225c5601b82980929a9

Observation 2bc3246a-e9f7-4e8a-9fbe-5c0dc54e1b5d · outbound

This paper cites and Bhatnagar, S.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Bhatnagar, S

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.333323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.327290Z digest=sha256:0278d2603149c4d7ddda776acab8621b58ff0e989001879621acc33b7e6454a1

Observation af741505-959b-40aa-86ed-88620819c127 · outbound

This paper cites Personalized robotic control via constrained multi-objective reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Personalized robotic control via constrained multi-objective reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.314490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.331700Z digest=sha256:8ca57c25869dd12b169a4465de46e0378600c38942a685976d1103159b2d212f

Observation dd8f973c-9032-4d22-a699-aebf289f3952 · outbound

This paper cites Dojo: A Differentiable Physics Engine for Robotics.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Dojo: A Differentiable Physics Engine for Robotics

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.336180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.336180Z digest=sha256:73482fc80f5ed84636f028a295931a412de0a357801091d8a23a1a2edee76507

Observation 542a055a-db30-4c96-9c4f-51a54f4500b6 · outbound

This paper cites Deep differentiable reinforcement learning and optimal trading.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Deep differentiable reinforcement learning and optimal trading

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.287611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.341108Z digest=sha256:0dcb6a1d6344099511d4bb2164898e75ce20977b5f59b7dd447189c459147288

Observation d9df688e-ce39-4a8d-8168-6830c17bb556 · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation AI Alignment: A Comprehensive Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.344908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.344908Z digest=sha256:86c33b99f0648ff3c8979528b6bdd2082f73cf051589a6d4e27464ef1c0538a1

Observation bbb0b150-bbd1-441f-988a-92c973798b93 · outbound

This paper cites Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.349062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.349062Z digest=sha256:c736a16991e835e93deeec071ed8608850055f13aa587126d2179012e5990f29

Observation b1796f7d-c120-452b-b1a1-62d0045df476 · outbound

This paper cites OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.353163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.353163Z digest=sha256:b2f033938645e9548be9213126af08d1a5ecb4e4711cb941e386fffeb9f1f280

Observation fe97cd66-6674-40cc-8114-c2d85b8c6e79 · outbound

This paper cites Aligner: Efficient Alignment by Learning to Correct.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Aligner: Efficient Alignment by Learning to Correct

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.357555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.357555Z digest=sha256:a843dead4da280fef2593980d1fec0bf5fe664caaab2b44e4c030c6109db300b

Observation ffc5257a-0570-4ec5-a019-8bb2b1a35488 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.273181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.361653Z digest=sha256:1b20f76b1707d7a49e48a265d344a2857b095f492e0808131ae42e3c25e23c12

Observation 0c3ca86f-84c8-49a4-a559-a0d2e92c2df0 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.261362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.365601Z digest=sha256:4a7e0931688a340c7bc8f6fab1e1ca1dd276ab48aa8e09ad4fbe99b0469fdc7b

Observation 1cbd7cbe-0eea-4de5-b983-a42d7b8e4fea · outbound

This paper cites and Langford, J.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Langford, J

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.249699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.369026Z digest=sha256:f7510e8e79aa736f50223f21f1b6c7fe32810643adbfd34e9dc4803b43946543

Observation ef741c55-a3d1-444e-949e-cf7bdb033418 · outbound

This paper cites C., Jain, R., and Nuzzo, P.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation C., Jain, R., and Nuzzo, P

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.238206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.372814Z digest=sha256:e5d1df2a7ce7dc432551f48f3ed8cc592ce4e17602e5b1b753c1126fa9d35fac

Observation 305c086e-6476-4c01-a751-90a7ed9ea8bc · outbound

This paper cites Reparameterization gradient for non-differentiable models.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Reparameterization gradient for non-differentiable models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.225583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.376100Z digest=sha256:838abdcb308a535b37589be07040d9a66eb5ef314640f4557c1063a8e82d2727

Observation 376d8cbb-5fde-4d28-bcf4-32f3a7da07f5 · outbound

This paper cites Constrained variational policy optimization for safe reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained variational policy optimization for safe reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.379471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.379471Z digest=sha256:daf56fddcb523cea44a5dbfc67ac70eba7bc499a06675044aa2ce468a770b668

Observation c807d510-2679-498a-a0dd-ba2240b1b7dd · outbound

This paper cites An off-policy trust region policy optimization method with monotonic improvement guarantee for deep reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation An off-policy trust region policy optimization method with monotonic improvement guarantee for deep reinforcement learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.383211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.383211Z digest=sha256:7a0a29a419653be001b410d2a4b4a16f05f06fb8669d77268d8b21a0d5b3d5e4

Observation cbd0547a-70e9-47d5-906c-7bfe07d721c8 · outbound

This paper cites Gradients are Not All You Need.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Gradients are Not All You Need

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.386945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.386945Z digest=sha256:b2ef7200847e59629d50f6bee68381eef7d7dded22839dad113864b1e895b902

Observation 2bc24076-c5a0-407f-b5b7-b541570fd9c8 · outbound

This paper cites Monte carlo gradient estimation in machine learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Monte carlo gradient estimation in machine learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.204095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.390858Z digest=sha256:26ac6a2d53f14511f60639190a557b4c45179d4cc7e122a6d56a0d576a3f1303

Observation 49bd8828-605e-4244-a788-311f55409221 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.191282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.394340Z digest=sha256:d7f3731446bd956b914747df61bc55680226625f7c8174028d7a878326c7128f

Observation 7acdd6df-4a9a-4ef2-bff5-8f5a322e9b8b · outbound

This paper cites A focused backpropagation algorithm for temporal pattern recognition.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A focused backpropagation algorithm for temporal pattern recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.178946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.398561Z digest=sha256:7399ac5ddbc3bdb1b2c1f570813af8aa413b2a68a23da90b01347112fd8aa959

Observation a63348ae-db32-4cce-ab11-1df098d90384 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.165377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.403481Z digest=sha256:77c44c4a65f6bf2a781f6bb21e3fd6b1b2810e1f76f0999d7986444710854337

Observation cbbc98fb-67aa-411f-bb0a-9dd7e69882c8 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.151741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.407548Z digest=sha256:01b966e41f7f58f40139115cb997f2e46942507b368fdefd55f24e11446d2867

Observation bf742935-2efa-4048-97d0-0bba24bf954d · outbound

This paper cites Trajectory planning with miscellaneous safety critical zones**this work was supported by ffi - strategic vehicle research and innovation.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Trajectory planning with miscellaneous safety critical zones**this work was supported by ffi - strategic vehicle research and innovation

Reference 37

Resolution
verified exact
doi, observed 2026-08-11T15:20:56.584281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.412556Z digest=sha256:e9248326089aab458bceb2d479a47264bdb5c11364e2ae4e6450b03722edf483

Observation c1761bfb-2b04-466e-9cb2-61bcc64c4769 · outbound

This paper cites M., Smaby, N., and Cutkosky, M.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation M., Smaby, N., and Cutkosky, M

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.134741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.418519Z digest=sha256:b470e5c6d48e1cca1a5fb275d0a12fef8de6d32ec14e802f4a3956c5c962988b

Observation 712e9719-cd9b-4042-80e5-8bb490350a44 · outbound

This paper cites Training language models to follow instructions with human feedback.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Training language models to follow instructions with human feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.423648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.423648Z digest=sha256:e6a78c07e58c5fe642c83eeaa0e07af7ea753d4ce77d4a4e496e29823ca07355

Observation 139bfd47-8804-48ff-9cb5-13faf20c5b9b · outbound

This paper cites Model-based reinforcement learning with scalable composite policy gradient estimators.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Model-based reinforcement learning with scalable composite policy gradient estimators

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.112583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.432253Z digest=sha256:7c2ac6f5aa69118bc10ad10485acc81fced923398406b5e167ced75fbc45c242

Observation 3114b496-2514-448b-852b-9f77beb9670a · outbound

This paper cites Model-based reinforcement learning with scalable composite policy gradient estimators.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Model-based reinforcement learning with scalable composite policy gradient estimators

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.095465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.437224Z digest=sha256:ccde8fcd552cc4532c0bf0e5fd196ed38d8e1a3728121c37e0b503941a2e9ac2

Observation 633aab94-066a-433e-8a49-dcfaa483f501 · outbound

This paper cites and Barr, A.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Barr, A

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.082293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.441781Z digest=sha256:8ab4a23697e30419b0c6168bdaa01d88f37be19a067a0856162c1b4668558229

Observation 682a5e1f-371a-4840-b23b-9977cfb85249 · outbound

This paper cites E., Perescu-Popescu, L., and Mastorakis, N.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation E., Perescu-Popescu, L., and Mastorakis, N

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.067430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.446967Z digest=sha256:f350098713d4f82cb123a7507f9bb012bc92ebb9016464c11b1645dc1c26d86b

Observation efaa3de5-6dd0-45e0-9e11-2630796a9d4a · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.451697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.451697Z digest=sha256:92d89ed11718cb1141266582391270a12027895c0966cfdde0b4712f55209bbb

Observation d5a0be68-f149-4d2b-a9ca-6fb830cd6854 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.456035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.456035Z digest=sha256:0030065b4ae9d909f8058916285de6af3e5ab76320e9f7e18869bbade11795f4

Observation 5af2260d-6e97-4eb8-8543-0fd2555934bb · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.460994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.460994Z digest=sha256:9f367ba02edbcf456ef4cb4da459f119a7df5b5facba4b961658d143c9dda9ae

Observation 4cb0e4ea-7af1-4b76-bb84-fdafd028042e · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.465505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.465505Z digest=sha256:9288e922e47110add76816c49a7bb67a5a034a6c5ff05d1c88da1803bc704b70

Observation 3a617aa4-eb65-40b0-9b18-09692828d39e · outbound

This paper cites Trust region policy optimization.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Trust region policy optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.470091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.470091Z digest=sha256:7e50d5021304d2048d1fa97aec139bed7b4f21fd8ea16adcfe2b9d46115bbf27

Observation 30eb5d28-521f-42c7-9a2d-90b69df97748 · outbound

This paper cites TBQ($\sigma$): Improving Efficiency of Trace Utilization for Off-Policy Reinforcement Learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation TBQ($\sigma$): Improving Efficiency of Trace Utilization for Off-Policy Reinforcement Learning

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:20:56.676622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.473886Z digest=sha256:4dab39639a55e8196605d2ffee41d675b246a5c005dae56e9c6c42d34d387d0a

Observation 2164f7e4-e74d-4d9f-9cc3-e67046c62bb1 · outbound

This paper cites J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.478429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.478429Z digest=sha256:99106c2e9f90e05eafa32a177de77b263b98d5c9ddd8013206a67151f73a5305

Observation 5d2a9931-94a1-4178-8e73-2206990cc789 · outbound

This paper cites Mastering the game of go without human knowledge.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Mastering the game of go without human knowledge

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.482832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.482832Z digest=sha256:7ff908930d0514e2bda437d9c4cbf8a4aaa05758734058d86f33fc7996c2c268

Observation e4475aa1-6272-45d6-8567-4328e096c74f · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.008450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.486674Z digest=sha256:955e57ec71a7128fa2546fa1c6fb7dc6d384d599a42897a78a0d36f5e8f49435

Observation 3b87282a-bebf-4bd3-99f1-0c0868e88f05 · outbound

This paper cites Responsive safety in reinforcement learning by pid lagrangian methods.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Responsive safety in reinforcement learning by pid lagrangian methods

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.994990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.490259Z digest=sha256:e49f0489b35ab26dc6523df5372833d5be101bd56fa3d53a6f8f43d8f184d5ea

Observation 70b62e8d-cedb-4af6-9ad0-e4ab9cc69157 · outbound

This paper cites J., Simchowitz, M., Zhang, K., and Tedrake, R.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation J., Simchowitz, M., Zhang, K., and Tedrake, R

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.980047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.494887Z digest=sha256:d3dc5ec389efaca4ef5bed4738896b213150df834c217634ab18e7f74baec9c8

Observation b60e1a78-6ea9-4152-85f0-a6a952922a4b · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.499664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.499664Z digest=sha256:083b25f93efc87e6605c36e3e7bfbb797a8e7fb79a92b15b2efc435fa5f7bc37

Observation f5a7f6ff-f92e-4263-875e-631c4c3bad3c · outbound

This paper cites W., Wang, T., Shang, Y., and Wu, Z.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation W., Wang, T., Shang, Y., and Wu, Z

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.956471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.505091Z digest=sha256:547f850ffb76d196474c5e426d1444d49b90592fd39b81bdca93f1d3d717b796

Observation a320c451-bbe7-4f3a-b23b-6681bab8ddc3 · outbound

This paper cites Development of a humanoid robot control system based on ar-bci and slam navigation.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Development of a humanoid robot control system based on ar-bci and slam navigation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.943923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.509463Z digest=sha256:63dadc34d3b569564bd9174cf637a317f45fb87b7f2e97ee5139409322f36b40

Observation cd42231a-daa8-4f4a-9662-c64b4836664a · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:56.931716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.515075Z digest=sha256:d569d2cb76219c191348a7cc903272f74960d785a21018fb6f28ede6e26982eb

Observation 39328285-6824-4b56-97bc-6b1ea74a029b · outbound

This paper cites FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.521513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.521513Z digest=sha256:24954711a64060017cd64959632875eeca5daf2ac3e27dfbcd5684902fd9ae10

Observation 9f83f98c-d743-4be3-8402-224876a928a8 · outbound

This paper cites Trustworthy Reinforcement Learning Against Intrinsic Vulnerabilities: Robustness, Safety, and Generalizability.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Trustworthy Reinforcement Learning Against Intrinsic Vulnerabilities: Robustness, Safety, and Generalizability

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.526507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.526507Z digest=sha256:f86c4da604d0c8f48bd75d0fcd5c6c82b986e0f1f1d9399ebb2afd8141da0093

Observation 1df34a0f-ee63-4136-a3d1-302b68d056d5 · outbound

This paper cites A Unified Approach for Multi-step Temporal-Difference Learning with Eligibility Traces in Reinforcement Learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A Unified Approach for Multi-step Temporal-Difference Learning with Eligibility Traces in Reinforcement Learning

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:20:56.632413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.531090Z digest=sha256:6571170e39e4eaf86d4eaa4dcab8745557db8768b14609fb20778d512d010600

Observation 73e89b4b-0b3d-413f-a788-dfa37e890dae · outbound

This paper cites Constrained update projection approach to safe policy optimization.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained update projection approach to safe policy optimization

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.917830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.536115Z digest=sha256:a83de96336438dd9ea2cfb96131de9584403d710f0283c5a3a436b226b70ee7b

Observation f7d4c6dd-6a6a-4971-9ab3-bf5fcb76046f · outbound

This paper cites Projection-Based Constrained Policy Optimization.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Projection-Based Constrained Policy Optimization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.540973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.540973Z digest=sha256:dd08e5f691280ab90412eeeb9913683cb3c237cd9984ed945588d27a40756475

Observation 4b6b7e47-0add-45be-89ac-407b5fee5bee · outbound

This paper cites First order constrained optimization in policy space.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation First order constrained optimization in policy space

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.901732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.545141Z digest=sha256:08287cb180392320c9323fdfdcb91d9030d4591e36abc7f481c465858e049a63

Pith citing papers

No inbound Pith citation observations are available.