Pith. sign in

Paper Citation Record · LEDGER

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation

As of 16 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 0 inbound Pith citation observations for arXiv:2412.11138.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11138 v1

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:20:56.545141Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

64 of 64 outbound references displayed

  • verified exact4
  • verified fuzzy28
  • unresolved32
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f0800d1b-1221-4dde-9f03-2fb971020bf8 · outbound

This paper cites write newline.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.249835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.249835Z digest=sha256:adf686dbc8fff88a3779d81f01b8770accf37773a0f0adce38290b698477a09c

Observation a360c338-1592-4384-8cbe-68d871306fc0 · outbound

This paper cites Constrained policy optimization, 2017.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained policy optimization, 2017

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.460464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.255303Z digest=sha256:b599afce1ac40b32feaf16dada441c3bd8b03f9b808d7f005a1630c201cfe749

Observation a71c32c2-d606-4952-b20a-8bda488a1571 · outbound

This paper cites Constrained Markov decision processes, volume 7.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained Markov decision processes, volume 7

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.446656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.260145Z digest=sha256:b474a9f7cd812aebe786b665e4c7dc98d6eac1319a7a538de24207b3307ec7f3

Observation d9d7241f-bd17-43b4-910c-2346b9547840 · outbound

This paper cites Constrained Policy Optimization via Bayesian World Models.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained Policy Optimization via Bayesian World Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.263978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.263978Z digest=sha256:4f2a4aecd3cd5c1107dce124505a17dd847a8616db2490f931304a0811cb269b

Observation b5b70bcd-f0d2-40d2-b46a-593fd00d3b87 · outbound

This paper cites Robots that interact with humans: a review of safety technologies and standards.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Robots that interact with humans: a review of safety technologies and standards

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.433993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.269227Z digest=sha256:70385f2516b3ac941eb1f5fa225ce5b879abe4f5000a6a4a579f14928fae3680

Observation cc377145-760d-4253-b4b2-f0818b5ed36c · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Risk-constrained reinforcement learning with percentile risk criteria

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.273249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.273249Z digest=sha256:a2474519f613d8e01c5d85785f667740457eac54b179a8b80c7a73fec4ee8012

Observation feb4e2bc-1b61-424a-980b-d7003ef5fec4 · outbound

This paper cites Model-Augmented Actor-Critic: Backpropagating through Paths.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Model-Augmented Actor-Critic: Backpropagating through Paths

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.278127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.278127Z digest=sha256:c52e05e7d1b28d3b4527f96f8b1cee2ef5474600dcc890b6ceada86176ad1581

Observation 940c8509-2d63-4ecf-9739-8d5275988a23 · outbound

This paper cites Augmented proximal policy optimization for safe reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Augmented proximal policy optimization for safe reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.408861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.283465Z digest=sha256:48b3a567d1d5850ea324fb83b957f9f39dcd21c2cf4fd2d8a0eac6b84e714704

Observation fa0132e8-2c18-4e22-a256-bf12ef92e5d8 · outbound

This paper cites Safe RLHF : Safe reinforcement learning from human feedback.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Safe RLHF : Safe reinforcement learning from human feedback

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.290448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.290448Z digest=sha256:043a8f9a37646a529dfd35a1c1bcb38eb0f2302e6f23ef9f031d3ee1c25d11c2

Observation 69df94db-b3f8-4867-9d6b-23b7f539f038 · outbound

This paper cites A differentiable physics engine for deep learning in robotics.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A differentiable physics engine for deep learning in robotics

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.388846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.296012Z digest=sha256:04a8546fd0addfb1a1f8e6b8db30bf68b8233e713395ee9a0850d420834dc600

Observation 7e2fba4c-a9f7-4be6-bcaa-350a35f3c12b · outbound

This paper cites A., Farouk, H., and Mofreh, E.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A., Farouk, H., and Mofreh, E

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.376244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.302087Z digest=sha256:b62e2604ba928b005d48e5f08042a64c2e1dbcbe01122b7f0f2cb1ba2fc17548

Observation 7ec16299-e684-4742-9a86-ad0a989acecd · outbound

This paper cites D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation D., Frey, E., Raichuk, A., Girgin, S., Mordatch, I., and Bachem, O

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.364745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.306751Z digest=sha256:c847cae36b99db72f635e2ac91f94027853151c0f3465507f67560f745d9afd5

Observation 998678a0-0f6a-4dc2-8e83-c9718926bce9 · outbound

This paper cites A Review and Outlook on Predictive Cruise Control of Vehicles and Typical Applications Under Cloud Control System.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A Review and Outlook on Predictive Cruise Control of Vehicles and Typical Applications Under Cloud Control System

Reference 13

Resolution
verified exact
doi, observed 2026-08-11T15:20:56.597981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.311323Z digest=sha256:65b963b5930bbba808f096b47f89759a957ed832bbd18e548b4bbc398b74bd86

Observation 6da0f8c1-27d8-4aa7-8fb6-e54e96503af9 · outbound

This paper cites and Fern \'a ndez, F.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Fern \'a ndez, F

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.318141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.318141Z digest=sha256:fd662d2fff3b1d5ac75190cb582738e16d357b86131b22dfc992f9138c746f6d

Observation 2c72b972-cdd9-4fbb-9475-3d4fc74d445a · outbound

This paper cites Bullet-safety-gym: A framework for constrained reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Bullet-safety-gym: A framework for constrained reinforcement learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.345538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.322058Z digest=sha256:bda12fbb2068b4108061d6246457c92d972f3127e02ef8be2c8330fb291d405c

Observation 2bc3246a-e9f7-4e8a-9fbe-5c0dc54e1b5d · outbound

This paper cites and Bhatnagar, S.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Bhatnagar, S

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.333323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.327290Z digest=sha256:d1da26e305cc8c0113e379ebd30549cee2d4ef309df8b24cd38555da0e7b9417

Observation af741505-959b-40aa-86ed-88620819c127 · outbound

This paper cites Personalized robotic control via constrained multi-objective reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Personalized robotic control via constrained multi-objective reinforcement learning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.314490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.331700Z digest=sha256:2bb5689af6c77a8dfc59170b1453c6dc1e3f13ca1da43be99d28e9e281d8ca3c

Observation dd8f973c-9032-4d22-a699-aebf289f3952 · outbound

This paper cites Dojo: A Differentiable Physics Engine for Robotics.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Dojo: A Differentiable Physics Engine for Robotics

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.336180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.336180Z digest=sha256:6dcb857975f347c5c78c41c9b9e972d946d9daf06d0c4805bd5746a346a02e28

Observation 542a055a-db30-4c96-9c4f-51a54f4500b6 · outbound

This paper cites Deep differentiable reinforcement learning and optimal trading.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Deep differentiable reinforcement learning and optimal trading

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.287611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.341108Z digest=sha256:08b38e173eee168d69259e46de3f3e467f476402581dca409511463de71c688c

Observation d9df688e-ce39-4a8d-8168-6830c17bb556 · outbound

This paper cites AI Alignment: A Comprehensive Survey.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation AI Alignment: A Comprehensive Survey

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.344908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.344908Z digest=sha256:0e3c06488da9163f6e8f694ac45cfdebd3dcb64e06aa3ba44f0a023059ea9fdc

Observation bbb0b150-bbd1-441f-988a-92c973798b93 · outbound

This paper cites Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Safety-Gymnasium: A Unified Safe Reinforcement Learning Benchmark

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.349062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.349062Z digest=sha256:ec780d8d83efb90c3fd84df48dce9ba23e9ab49cc4a38baa117c8bf546965fff

Observation b1796f7d-c120-452b-b1a1-62d0045df476 · outbound

This paper cites OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation OmniSafe: An Infrastructure for Accelerating Safe Reinforcement Learning Research

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.353163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.353163Z digest=sha256:529f2a08750f97242f49ec7c195bdb92e054a215fb0e9cd650de1031e1f29c10

Observation fe97cd66-6674-40cc-8114-c2d85b8c6e79 · outbound

This paper cites Aligner: Efficient Alignment by Learning to Correct.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Aligner: Efficient Alignment by Learning to Correct

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.357555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.357555Z digest=sha256:75f538f9b0d8d485a466ba7f5add5fbf49be5b28e3305857e199d448847dd3fd

Observation ffc5257a-0570-4ec5-a019-8bb2b1a35488 · outbound

This paper cites Beavertails: Towards improved safety alignment of llm via a human-preference dataset.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Beavertails: Towards improved safety alignment of llm via a human-preference dataset

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.273181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.361653Z digest=sha256:1b541c2a7ccaf014de8f6660571961c04c5af8825af8260ac8fc9ceb46d6f328

Observation 0c3ca86f-84c8-49a4-a559-a0d2e92c2df0 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.261362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.365601Z digest=sha256:14af6fbd4f057b7195f5584c9cc84060de62656143201af9e8369e4d8f055c78

Observation 1cbd7cbe-0eea-4de5-b983-a42d7b8e4fea · outbound

This paper cites and Langford, J.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Langford, J

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.249699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.369026Z digest=sha256:4fd77c5474d9f3fda53e36137a41ce451d2ec2f3144105a8fd7d7605154af036

Observation ef741c55-a3d1-444e-949e-cf7bdb033418 · outbound

This paper cites C., Jain, R., and Nuzzo, P.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation C., Jain, R., and Nuzzo, P

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.238206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.372814Z digest=sha256:afd9c5ed32b389066565d0432b2e7442b4ce60c5a481d229e502756127a3a28d

Observation 305c086e-6476-4c01-a751-90a7ed9ea8bc · outbound

This paper cites Reparameterization gradient for non-differentiable models.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Reparameterization gradient for non-differentiable models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.225583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.376100Z digest=sha256:a8bf90a5eb0f6a151c9b93a0f7c22e7c2bde2c2033d3433d5535e8013c20ee1f

Observation 376d8cbb-5fde-4d28-bcf4-32f3a7da07f5 · outbound

This paper cites Constrained variational policy optimization for safe reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained variational policy optimization for safe reinforcement learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.379471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.379471Z digest=sha256:8095682aeed72b288944fb5385ce3c7ce429857bec3ffebda24bd395f2e8d10f

Observation c807d510-2679-498a-a0dd-ba2240b1b7dd · outbound

This paper cites An off-policy trust region policy optimization method with monotonic improvement guarantee for deep reinforcement learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation An off-policy trust region policy optimization method with monotonic improvement guarantee for deep reinforcement learning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.383211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.383211Z digest=sha256:cd2720d1662da9bf287fc8310760db9f3841b8d46ce9fe6bc9371855036289db

Observation cbd0547a-70e9-47d5-906c-7bfe07d721c8 · outbound

This paper cites Gradients are Not All You Need.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Gradients are Not All You Need

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.386945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.386945Z digest=sha256:56af3c9c01bd1ef19780991e63ff32f9817b6cb9428a2797b8deba201c7583c4

Observation 2bc24076-c5a0-407f-b5b7-b541570fd9c8 · outbound

This paper cites Monte carlo gradient estimation in machine learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Monte carlo gradient estimation in machine learning

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.204095Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.390858Z digest=sha256:4a7b2666573fdc54a81f8c3652219101dc26008e236dd51560ecd4a8619d8519

Observation 49bd8828-605e-4244-a788-311f55409221 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.191282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.394340Z digest=sha256:e700445b23199b53fc347c316565f5c0429bfd4c3e19f4f40d4e06186d51f44c

Observation 7acdd6df-4a9a-4ef2-bff5-8f5a322e9b8b · outbound

This paper cites A focused backpropagation algorithm for temporal pattern recognition.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A focused backpropagation algorithm for temporal pattern recognition

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.178946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.398561Z digest=sha256:59f6185656136739d53af567ee5ae8c9a4356b34e0f48d38d7aca01758b0fabd

Observation a63348ae-db32-4cce-ab11-1df098d90384 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.165377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.403481Z digest=sha256:77f34160df005c250fe2302cfbc795fc1730a927cba7969c37151e743689faeb

Observation cbbc98fb-67aa-411f-bb0a-9dd7e69882c8 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.151741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.407548Z digest=sha256:a6208c686c6d6beb24557eca9a599bb81d2b553cf3423822467ba9342f9a5a2d

Observation bf742935-2efa-4048-97d0-0bba24bf954d · outbound

This paper cites Trajectory planning with miscellaneous safety critical zones**this work was supported by ffi - strategic vehicle research and innovation.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Trajectory planning with miscellaneous safety critical zones**this work was supported by ffi - strategic vehicle research and innovation

Reference 37

Resolution
verified exact
doi, observed 2026-08-11T15:20:56.584281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.412556Z digest=sha256:f2bcf46e2210de3d574d5dc7b29dbc67eeca8e9758fde97a3cc1e9b7ae0ca9bc

Observation c1761bfb-2b04-466e-9cb2-61bcc64c4769 · outbound

This paper cites M., Smaby, N., and Cutkosky, M.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation M., Smaby, N., and Cutkosky, M

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.134741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.418519Z digest=sha256:d582c6e7e7b7d655e0a7ac836da1326c022d1ba32216b9f7636131b47dbdd450

Observation 712e9719-cd9b-4042-80e5-8bb490350a44 · outbound

This paper cites Training language models to follow instructions with human feedback.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Training language models to follow instructions with human feedback

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.423648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.423648Z digest=sha256:81c5f4910a18087bfece4b416fb6f2dff0a6f9f4433ec56e75f90e27cee31dfd

Observation 139bfd47-8804-48ff-9cb5-13faf20c5b9b · outbound

This paper cites Model-based reinforcement learning with scalable composite policy gradient estimators.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Model-based reinforcement learning with scalable composite policy gradient estimators

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.112583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.432253Z digest=sha256:3525d89b6852a5f82939156382f86c849358a0bcb233419ef550e965236ac3f9

Observation 3114b496-2514-448b-852b-9f77beb9670a · outbound

This paper cites Model-based reinforcement learning with scalable composite policy gradient estimators.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Model-based reinforcement learning with scalable composite policy gradient estimators

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.095465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.437224Z digest=sha256:f969172312463a1a0e390b30e3e6af422f6d40b7e745c78242a76a5dbd430061

Observation 633aab94-066a-433e-8a49-dcfaa483f501 · outbound

This paper cites and Barr, A.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation and Barr, A

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.082293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.441781Z digest=sha256:4883465d8454bb47c997415822c52f682fad13ce119b0f6dbdfe64e436999f00

Observation 682a5e1f-371a-4840-b23b-9977cfb85249 · outbound

This paper cites E., Perescu-Popescu, L., and Mastorakis, N.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation E., Perescu-Popescu, L., and Mastorakis, N

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:57.067430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.446967Z digest=sha256:351635fd33204aec991e72035f41451d9ba6edff461bb1a0a8c0f1b8ce0f5ecd

Observation efaa3de5-6dd0-45e0-9e11-2630796a9d4a · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.451697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.451697Z digest=sha256:e686a1c865be7efdd92b038c8faceb2475a5c1f59fe35dabe1d8213005b6c100

Observation d5a0be68-f149-4d2b-a9ca-6fb830cd6854 · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.456035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.456035Z digest=sha256:409171f5aacf6e57843f81f19751ca27b5da455afb0effe3991afa499d3cc105

Observation 5af2260d-6e97-4eb8-8543-0fd2555934bb · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.460994Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.460994Z digest=sha256:1d7587172d306e2d3acfe8d2d52b332e56271e9d954a725cf5322141ed10e4a3

Observation 4cb0e4ea-7af1-4b76-bb84-fdafd028042e · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.465505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.465505Z digest=sha256:ab1f258c3236cc5954261c0f8f8d0e691d868c1e4559e028d22e727c938d9c53

Observation 3a617aa4-eb65-40b0-9b18-09692828d39e · outbound

This paper cites Trust region policy optimization.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Trust region policy optimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.470091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.470091Z digest=sha256:bc767e864ff1c23871db767175a2e66ccb8861f75eb1be63dbaf6b30fc7fcdfb

Observation 30eb5d28-521f-42c7-9a2d-90b69df97748 · outbound

This paper cites TBQ($\sigma$): Improving Efficiency of Trace Utilization for Off-Policy Reinforcement Learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation TBQ($\sigma$): Improving Efficiency of Trace Utilization for Off-Policy Reinforcement Learning

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:20:56.676622Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.473886Z digest=sha256:fafcfca0c783b5802f4695896a4c4fd3d5669131b6f802e622c684dffcc01cdc

Observation 2164f7e4-e74d-4d9f-9cc3-e67046c62bb1 · outbound

This paper cites J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation J., Guez, A., Sifre, L., Van Den Driessche, G., Schrittwieser, J., Antonoglou, I., Panneershelvam, V., Lanctot, M., et al

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.478429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.478429Z digest=sha256:3ac452a94491f3c2eaff7066e34ff8a109458085038a68810fcb0caee2057176

Observation 5d2a9931-94a1-4178-8e73-2206990cc789 · outbound

This paper cites Mastering the game of go without human knowledge.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Mastering the game of go without human knowledge

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.482832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.482832Z digest=sha256:2d9ce8603e4d0086b46d5e814ec45e155e51c1af25dd4940c4479b5f79cc1144

Observation e4475aa1-6272-45d6-8567-4328e096c74f · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:57.008450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.486674Z digest=sha256:e1f10c5ec9dbec6c63b5e9a88dea4ade3edf6a14d2da4202319af50be1e7eb7e

Observation 3b87282a-bebf-4bd3-99f1-0c0868e88f05 · outbound

This paper cites Responsive safety in reinforcement learning by pid lagrangian methods.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Responsive safety in reinforcement learning by pid lagrangian methods

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.994990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.490259Z digest=sha256:f2905a1fc58fabafdcb669c09f85077fe5df73e4e32d514f14b80ed7372a4224

Observation 70b62e8d-cedb-4af6-9ad0-e4ab9cc69157 · outbound

This paper cites J., Simchowitz, M., Zhang, K., and Tedrake, R.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation J., Simchowitz, M., Zhang, K., and Tedrake, R

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.980047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.494887Z digest=sha256:f6b1e08dee4a2cc30ee978d186a45e99fdb16169af64aebe5f175a0c687c13df

Observation b60e1a78-6ea9-4152-85f0-a6a952922a4b · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.499664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.499664Z digest=sha256:7d51029cf43a38715d3458fed3440b6175aeac6f39e5e2d5ea5c9381466fa5e5

Observation f5a7f6ff-f92e-4263-875e-631c4c3bad3c · outbound

This paper cites W., Wang, T., Shang, Y., and Wu, Z.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation W., Wang, T., Shang, Y., and Wu, Z

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.956471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.505091Z digest=sha256:2fc02fc9bac068a1540f919c8fb15188272fac8896227fa0ded662b0c2cead99

Observation a320c451-bbe7-4f3a-b23b-6681bab8ddc3 · outbound

This paper cites Development of a humanoid robot control system based on ar-bci and slam navigation.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Development of a humanoid robot control system based on ar-bci and slam navigation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.943923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.509463Z digest=sha256:4f38c42b30e660b23e0c8b406f689c6a6cfb829d99e1394ca06ab510878d6803

Observation cd42231a-daa8-4f4a-9662-c64b4836664a · outbound

This paper cites an unresolved cited work.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:20:56.931716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.515075Z digest=sha256:713861c6c30b6caa5647a5b345837bb10a67ea4807b9b57ef3a6974d285cd9be

Observation 39328285-6824-4b56-97bc-6b1ea74a029b · outbound

This paper cites FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation FluidLab: A Differentiable Environment for Benchmarking Complex Fluid Manipulation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.521513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.521513Z digest=sha256:0ae5ae05c5ac073733c6a2f8e290b0e34dfadf63ad9f592daddc4443235e3533

Observation 9f83f98c-d743-4be3-8402-224876a928a8 · outbound

This paper cites Trustworthy Reinforcement Learning Against Intrinsic Vulnerabilities: Robustness, Safety, and Generalizability.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Trustworthy Reinforcement Learning Against Intrinsic Vulnerabilities: Robustness, Safety, and Generalizability

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.526507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.526507Z digest=sha256:26c8fa4af2ae84bf8dc1ba2cdec5a7d05d061c7ca58a3a7f8db55a11eb7f9602

Observation 1df34a0f-ee63-4136-a3d1-302b68d056d5 · outbound

This paper cites A Unified Approach for Multi-step Temporal-Difference Learning with Eligibility Traces in Reinforcement Learning.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation A Unified Approach for Multi-step Temporal-Difference Learning with Eligibility Traces in Reinforcement Learning

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-11T15:20:56.632413Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.531090Z digest=sha256:e36911259d0c83dfd9208b2e1830841b84a12fe63f19fb100cb55f68d7f2cb40

Observation 73e89b4b-0b3d-413f-a788-dfa37e890dae · outbound

This paper cites Constrained update projection approach to safe policy optimization.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Constrained update projection approach to safe policy optimization

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.917830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.536115Z digest=sha256:d5076837da6cf97ed4a0a2e22b8ed0f2db8bd59f9d0b74463b02215052f34e54

Observation f7d4c6dd-6a6a-4971-9ab3-bf5fcb76046f · outbound

This paper cites Projection-Based Constrained Policy Optimization.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation Projection-Based Constrained Policy Optimization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T15:20:56.540973Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:20:56.540973Z digest=sha256:7fcbced529d583a398db6569b82cdc45b909348f3081b37ae619bec5e40728ca

Observation 4b6b7e47-0add-45be-89ac-407b5fee5bee · outbound

This paper cites First order constrained optimization in policy space.

Safe Reinforcement Learning using Finite-Horizon Gradient-based Estimation First order constrained optimization in policy space

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:20:56.901732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-11T15:20:56.545141Z digest=sha256:b5b20f01f30ad1f04fd2692d54104a61ade2bcc1f1714d0e74d330f4ee3fdce2

Pith citing papers

No inbound Pith citation observations are available.