Pith. sign in

Paper Citation Record · LEDGER

A Provable Approach for End-to-End Safe Reinforcement Learning

As of 15 August 2026, this Paper Citation Record lists 61 of 61 outbound references and 0 inbound Pith citation observations for arXiv:2505.21852.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21852 v1

Coverage vector

measured 61 of 61 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:30:26.255011Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

61 of 61 outbound references displayed

  • verified exact1
  • verified fuzzy26
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5caa45fb-6d24-48eb-8d23-c96a11334fb7 · outbound

This paper cites Achiam, D.

A Provable Approach for End-to-End Safe Reinforcement Learning Achiam, D

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:37.889234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:19.727315Z digest=sha256:54155978dc27e5fb6edfc354da40b1a9afab409b9d0bc28ac2a82ac4d3869e2e

Observation 1ea45edf-8f6a-4412-922d-31cdef7e8a52 · outbound

This paper cites Alshiekh, R.

A Provable Approach for End-to-End Safe Reinforcement Learning Alshiekh, R

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:37.708658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:19.857266Z digest=sha256:cba6a2f93ddf558c0540e1f45c7116465e858446cccd4bcd3d3472310697394d

Observation 4ef9eeff-cff6-48f8-a334-a5a74e082f37 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:37.456951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:19.994834Z digest=sha256:97c7d21e010d1813d3bde56b6e3e38ed98319902ea99c30e9b05ca9b08ea6a73

Observation 7a62189b-0e02-41e8-b30e-bfed14d055d6 · outbound

This paper cites Concrete Problems in AI Safety.

A Provable Approach for End-to-End Safe Reinforcement Learning Concrete Problems in AI Safety

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:20.092151Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:20.092151Z digest=sha256:00bfa14b32a9d74e9664b8cccc8e201934afc50f39d184201f84b8b9a28aef54

Observation fc6ff8e2-20c2-4c3b-86b9-90bfad22a305 · outbound

This paper cites Berkenkamp, M.

A Provable Approach for End-to-End Safe Reinforcement Learning Berkenkamp, M

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:37.287804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:20.255035Z digest=sha256:50b220e449b628263419af28409de8485f3416cbac6c625b5721684fb6d2bdb9

Observation 74ccda11-804f-4a36-9ef6-bd3f6e011539 · outbound

This paper cites Bhatnagar and K.

A Provable Approach for End-to-End Safe Reinforcement Learning Bhatnagar and K

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.998638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:20.443200Z digest=sha256:afef0b08fbe5e784b9ee5e754145dfda7b59fc879b4dcdc04ad31cf05deb563c

Observation 5a2a7343-b76f-4fad-ba6e-a7991a8f7ccb · outbound

This paper cites Black, M.

A Provable Approach for End-to-End Safe Reinforcement Learning Black, M

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.763546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:20.564755Z digest=sha256:76c6b32b9e0f942f101623df2995a2603f67f732cdda971d603a403431169373

Observation 32c8fa42-bdd5-451d-86ac-026819081ba9 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:36.610803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:20.657269Z digest=sha256:e3b5628bcb2de0d4aedd638ef22bb1f7d40ddd9d572520a40db837078f074afb

Observation 8d0dea74-9e4d-4cc5-a7a2-0adfaa9148bc · outbound

This paper cites Brandfonbrener, A.

A Provable Approach for End-to-End Safe Reinforcement Learning Brandfonbrener, A

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.444884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:20.882152Z digest=sha256:ee54508fa75f51fd86917d3eef07d70b55c6ca227a78e06f35fe98ab398ab52d

Observation 4f6df005-2369-4bd7-a90c-d7466538cc85 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:36.211969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:20.964019Z digest=sha256:e1d26272dda2af46739c31d82e709d1bbd7d6f112ea3ee5b48e55bf762ed1a14

Observation c9dddd1d-31a7-4780-9147-a4281bd3b44e · outbound

This paper cites Cheng, G.

A Provable Approach for End-to-End Safe Reinforcement Learning Cheng, G

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:36.005293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:21.094841Z digest=sha256:e11cb61250510048177432ad3c7e6a154316b210d85889baa95545a3786818f4

Observation 6e2775c4-5441-4b23-a1c7-88659cd599c9 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:35.833182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:21.171133Z digest=sha256:4c76228b215a34bccea01ac982e2a3dc68ea469661f89227b535bdc132024d1d

Observation 674dfce5-6ac5-4f54-bc8a-68d633f7205e · outbound

This paper cites Da Costa, M.

A Provable Approach for End-to-End Safe Reinforcement Learning Da Costa, M

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:21.294753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:21.294753Z digest=sha256:34d6c7c94ee25dc26fa6e7415b5e68e1d59a0506c89f4a3074ac111440a8e324

Observation b13a745c-023b-40b8-8a72-08602e0cbedc · outbound

This paper cites Emmons, B.

A Provable Approach for End-to-End Safe Reinforcement Learning Emmons, B

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.700611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:21.394376Z digest=sha256:888b8d1b40083039a37616536690ad9fe8ae263c186003f2322ba3b470e87ce3

Observation 1e5e9932-f7f1-489a-bb75-3eba8f490110 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:35.525264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:21.491943Z digest=sha256:b79205cf6a66827b1c843657538dc1fcfe7c0ba65ffd8ba8694e353e9717b67d

Observation 092bc828-ae16-4493-8bed-b153b1723c04 · outbound

This paper cites Fujimoto, D.

A Provable Approach for End-to-End Safe Reinforcement Learning Fujimoto, D

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.288966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:21.614956Z digest=sha256:9a40461e2d4f0ed6d55e99acb9f2faa4ca1b3360b17a476bd698ec14ae9c9e11

Observation f5e404af-41b3-4494-966c-3c5cc3b0f525 · outbound

This paper cites Fulton and A.

A Provable Approach for End-to-End Safe Reinforcement Learning Fulton and A

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:35.105286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:21.692574Z digest=sha256:325a1504040423629fde890184e476df88dc009e7d05df6283a16cb415e8c0d6

Observation a7274ed4-4e90-4f4e-83fa-9173b709f53a · outbound

This paper cites Garcıa and F.

A Provable Approach for End-to-End Safe Reinforcement Learning Garcıa and F

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.901866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:21.790603Z digest=sha256:d44ecee398496bb2f69831e23d1c2575956f6a1cfc66e3df40e59f430e4ca0f6

Observation 97fcee7b-9fd3-4b4c-8a43-f7e7b001e6aa · outbound

This paper cites Gronauer.

A Provable Approach for End-to-End Safe Reinforcement Learning Gronauer

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.699316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:21.937413Z digest=sha256:644139735e977403a39a3418cf25292f4e7d1748de52892b1fa04fcb93b20959

Observation c32c7df1-e8cf-4c9d-b28f-55a1f4bf4c8d · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:34.500351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:22.049343Z digest=sha256:1da0a514c12385ca21dc498ccad44706e5943a7f2626c2032565f46f703c7317

Observation adbed44e-9a9b-4514-a306-d885ac1df103 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

A Provable Approach for End-to-End Safe Reinforcement Learning DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.159218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.159218Z digest=sha256:8781bfe0778c3c0b97de08a1226f33a257eb5818ec1ca7be22ef246caf3a82f8

Observation 690ac445-dd81-405f-8f18-d5424fdf5081 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:34.344303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:22.245173Z digest=sha256:8c7ce0bf1bb6e326e7f87bcabff542261d9496068ac98dcb1611856fc43f01d6

Observation dc55872e-60ae-495a-90c6-99f555156613 · outbound

This paper cites Hambly, R.

A Provable Approach for End-to-End Safe Reinforcement Learning Hambly, R

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:34.049121Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:22.326144Z digest=sha256:f6f9766cc0e9f8579ac76514b1a241720af53613ce3479ce3f589993bb3c670f

Observation 00fbdec3-2754-4655-acb8-94c756ce42f1 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:33.895147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:22.391642Z digest=sha256:56247054b50a6e114a6c4c25426dab81cd16edd3f97d41b745ba9732e62b2243

Observation 6897be8e-84ed-4d96-868a-a3899270c6c7 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:33.732780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:22.504885Z digest=sha256:6627c76d5f475a0eda4da2f9d26da3a7d7ca5d029a1183d5404930d8dc726207

Observation 8e5fce55-3c99-4ee3-a2a8-9332af96a67c · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:33.570967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:22.637476Z digest=sha256:f92c734c178c802ce96d13720351e160d7bfedd7282e916b303e292b96657a50

Observation cb9f6868-d166-4fd2-9fb6-db9ed9a6026c · outbound

This paper cites Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking.

A Provable Approach for End-to-End Safe Reinforcement Learning Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.701820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.701820Z digest=sha256:7b9464104975f55c324ffafb39c4ad45e154c70afd02c3a96c647211c4f24616

Observation 02dac6d4-7dbd-4165-993c-83a9995dff40 · outbound

This paper cites Kumar, J.

A Provable Approach for End-to-End Safe Reinforcement Learning Kumar, J

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:33.369305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:22.786585Z digest=sha256:ad70c25d36e0b86b390eb9c0efe2a90b5f7446e25dc8e86c88a971ab928b4082

Observation 0124c1ed-df2a-44cc-8b23-59c16ca05f04 · outbound

This paper cites Reward-Conditioned Policies.

A Provable Approach for End-to-End Safe Reinforcement Learning Reward-Conditioned Policies

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:22.843583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:22.843583Z digest=sha256:fef1455ed81e4615794b716283c1042b3c140b8a74dc3c5058305988115b4ec5

Observation 139671e5-860b-468f-97ce-5e25eb919d11 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:33.098066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:22.975095Z digest=sha256:108c048c08af7e2a031be07f272d66ad16ced7052fa305a8ccdf5912bca94f37

Observation febdd373-fd80-472d-9a82-065076405ea4 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.875494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:23.089264Z digest=sha256:ecc7bd8d7c51e4c17df89c5e2226f46682b7e0e4b47c97648627b1cef4e231fe

Observation 6ea14459-59f1-4823-b1e6-d0a9d8720cdb · outbound

This paper cites Levine, C.

A Provable Approach for End-to-End Safe Reinforcement Learning Levine, C

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:32.638209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:23.184252Z digest=sha256:783eab57f66544ae65d060087692555e2fec2659ecc6cfb7e06a96feb9c156d9

Observation dd2a1b61-3e23-49c2-97e9-7bb23ee939f8 · outbound

This paper cites Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems.

A Provable Approach for End-to-End Safe Reinforcement Learning Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.279554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.279554Z digest=sha256:3789b3ee9113118d35445852ce0797186fd5c21c47b87c4248fa5b901b31504a

Observation dc0847ab-7a35-4193-904f-bde708311a34 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.392354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:23.333956Z digest=sha256:9a8d8b44781f0610cb41dd3c606241d0011832acb301de7e79716a1fe6efedef

Observation 93140e8d-fce0-41ad-a807-5a00ac9cbac0 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.244747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:23.434878Z digest=sha256:560982e544758c5ad8efc15a2ee58c33991f70f1a287a997f914d29f9f569953

Observation f0751b9c-ed33-46ac-bf04-49112ec2ab7d · outbound

This paper cites Datasets and Benchmarks for Offline Safe Reinforcement Learning.

A Provable Approach for End-to-End Safe Reinforcement Learning Datasets and Benchmarks for Offline Safe Reinforcement Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.525259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.525259Z digest=sha256:b42c5e2796d1da8345d20ddf4ac76d6b73c2171bfac9c1dd34e4134a6fb80e29

Observation da4a3e35-5a01-478f-b769-9c6494246e5e · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:32.017371Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:23.654841Z digest=sha256:95e1ef5dbe0eceed12daefb79e20a198653d041ff43d5f213d1aedbecfb8d2a7

Observation dc6ab268-0dca-444c-80f9-9dea7274a561 · outbound

This paper cites Ouyang, J.

A Provable Approach for End-to-End Safe Reinforcement Learning Ouyang, J

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.804742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:23.724974Z digest=sha256:bb2b3a6a25f72469c476d5f233e48c14a308ae143c666e3e661bfe1c745acedf

Observation 10e93eee-0ed5-4ce2-8e51-32cbdd160ab2 · outbound

This paper cites Safe Policies for Reinforcement Learning via Primal-Dual Methods.

A Provable Approach for End-to-End Safe Reinforcement Learning Safe Policies for Reinforcement Learning via Primal-Dual Methods

Reference 39

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:30:27.002527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:23.813683Z digest=sha256:7cc54f5f75a218fe66ee04dcb1f34bb762ae9627148677eb7646824e35875d8c

Observation 4d25c2c2-a451-4962-96da-cd2e252c5317 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:23.872152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:23.872152Z digest=sha256:231278f6baa8606d1621f558ab69bd829c53a896cec7f680dbb568e97f524ede

Observation 3871230a-2563-4852-b4bb-ec764f5f4fae · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:31.514817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:23.963485Z digest=sha256:eae9cd1009105eab90171550320ab4dc0a2fa7e456457d103b3bb703bf2cca6a

Observation 4ff7817e-ca91-4258-844d-b00f8440da8f · outbound

This paper cites Satija, P.

A Provable Approach for End-to-End Safe Reinforcement Learning Satija, P

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.347590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:24.009443Z digest=sha256:c8ed919140a8824317238a9b8affe4238500142d29ae0330199d735a2a81b603

Observation 90cddf7c-eeb2-4a7d-bc5b-4a7371aa3876 · outbound

This paper cites Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions.

A Provable Approach for End-to-End Safe Reinforcement Learning Reinforcement Learning Upside Down: Don't Predict Rewards -- Just Map Them to Actions

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.086262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.086262Z digest=sha256:049f15f75ccfa974703619437a89e126ad37525d657504589b9398c15bfabff0

Observation add3ead9-d96b-4ea1-8ec3-d1b2e0eb72dc · outbound

This paper cites Sootla, A.

A Provable Approach for End-to-End Safe Reinforcement Learning Sootla, A

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:31.171392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:24.163137Z digest=sha256:844aae561939c4d5fdc3080e7eaf87d9a035e451fe05dce0bc65d5121ac56e1d

Observation d66c4c7e-dc5a-4ada-9017-3768a69efe62 · outbound

This paper cites Srinivas, A.

A Provable Approach for End-to-End Safe Reinforcement Learning Srinivas, A

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.974743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:24.214658Z digest=sha256:56d7e22f02a58ce7f5bb0eff628bf7bed729ac9eafded68530275d63dd2a7577

Observation fc5c38cc-235b-47fc-a44a-eabe153027c2 · outbound

This paper cites Training Agents using Upside-Down Reinforcement Learning.

A Provable Approach for End-to-End Safe Reinforcement Learning Training Agents using Upside-Down Reinforcement Learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.270417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.270417Z digest=sha256:bedd098b62c28e0d5eacbe9b4e7e95339ae9c9827fb801a4736489d668a04a4d

Observation 926187b4-167e-4697-8738-b0f9704de70d · outbound

This paper cites Stooke, J.

A Provable Approach for End-to-End Safe Reinforcement Learning Stooke, J

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.791525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:24.345683Z digest=sha256:667d395bd1441f01cdf6d8410ca2d1af5931d917f5b35b6124a956c9cace4e96

Observation d7812556-4fad-48c9-89e9-0720b8ceea70 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:30.584763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:24.514869Z digest=sha256:5a9470a031d782a2a8c9d0e732616d20b024e3e077d695d3061fcaccdb3b9399

Observation 3469f3d0-0a5b-49e6-8e86-57cabe6f424d · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:30.442925Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:24.636311Z digest=sha256:fc244d7ef9ff0e5fc4532b7bcc5ca85bad4821d5b9063941c4787705420eb971

Observation 47120573-a122-4411-92b6-ec129f9b34d5 · outbound

This paper cites Turchetta, F.

A Provable Approach for End-to-End Safe Reinforcement Learning Turchetta, F

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.277183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:24.735023Z digest=sha256:df236fc6d25c9e8d06cb3187fe65b0dda1aca914b8c5c4a8a32d3a74868f0f56

Observation b1411c21-861b-4134-81b4-67a501ca461e · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T13:30:24.804655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:30:24.804655Z digest=sha256:a949a928842823ee3b753215fbabec38d5e23697c43e20cc472033c5ab3475da

Observation d32a2ba4-93e9-441a-a0b4-beb49d816e7d · outbound

This paper cites Wachi and Y.

A Provable Approach for End-to-End Safe Reinforcement Learning Wachi and Y

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:30.104278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:24.914873Z digest=sha256:d1a5f7462e6a15ee1b50cfe8308b15017f2255f024732818dd56bb27eea822e4

Observation a0062dd8-7c4c-476a-aea2-d501f869a7db · outbound

This paper cites Wachi, W.

A Provable Approach for End-to-End Safe Reinforcement Learning Wachi, W

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.874805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:25.018298Z digest=sha256:c751d96aa3aac81137c0954724c5e8e51299d7d3f711dc9d42c2b01bb195fa2d

Observation 327c2ee6-cc89-40f5-8722-4dfe318abc7f · outbound

This paper cites Wachi, X.

A Provable Approach for End-to-End Safe Reinforcement Learning Wachi, X

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.713700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:25.234827Z digest=sha256:a8e124340d8b5a9dff1cbd2a0a3decf432420877ce26fac0e002cad12ad64576

Observation 3cfc6953-bd2c-471a-80eb-5e8de9f2eb1b · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:29.524828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:25.375135Z digest=sha256:6bb5065791aa0cb50fa2493cb5eeb24035111dc6d5a89108b903422c5e38fc37

Observation 4400f6c7-0285-4bb1-9c33-50dd9d5be1a7 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 56

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:29.299540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:25.644735Z digest=sha256:13bd903014abcc929b43f8a9c367fb1766242414604bc92ca94d679069141b30

Observation 23b55a16-b18d-49ce-a66f-b55830d03c47 · outbound

This paper cites Yang and M.

A Provable Approach for End-to-End Safe Reinforcement Learning Yang and M

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:29.074292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:25.815177Z digest=sha256:be8bf20bb7d8a503823ee51efc768b5f1447b711e47941aedea6c25d7e6ee799

Observation 701d1300-f321-4120-aba9-54be7671f6b7 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:28.892426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:25.934915Z digest=sha256:f880943fe1fabc6a3f04153c68495be8b4137383e830e2148bb0b178d4dc687c

Observation 92c4b7fb-3b1e-4c95-9a80-e8f6a8c5dfe6 · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 59

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:28.742128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:26.009155Z digest=sha256:cc41423dd50fbbe11d4723b6fd7a019a59b829632e558912dc257a4844ce64e3

Observation d7a6d77d-4f82-41b8-9618-8e2501cd0dcc · outbound

This paper cites an unresolved cited work.

A Provable Approach for End-to-End Safe Reinforcement Learning Unresolved cited work

Reference 60

Resolution
unresolved
raw_fallback, observed 2026-08-07T13:30:28.579001Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:26.125071Z digest=sha256:ca0c20c86ee51ffe9c5cd855923d08410258433af09b58a28d930655aa9e8007

Observation 6362a78c-7d61-49b9-95da-5e8cb5b3a191 · outbound

This paper cites Let β : X → ∆(A) be a behavior policy and D := {Ξ(i)}n i=1 ∼ (Pβ)n be a collection of n i.i.d.

A Provable Approach for End-to-End Safe Reinforcement Learning Let β : X → ∆(A) be a behavior policy and D := {Ξ(i)}n i=1 ∼ (Pβ)n be a collection of n i.i.d

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:30:28.467897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:30:26.255011Z digest=sha256:730b14634ae962bf8477902e1e883f87e13f7559af95f40a90f83ad7043f605e

Pith citing papers

No inbound Pith citation observations are available.