Pith. sign in

Paper Citation Record · LEDGER

Safe Online Learning via Smooth Safety-Structured Policy Composition

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2606.31320.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.31320 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:12:07.125494Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact10
  • verified fuzzy6
  • unresolved3
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5c4f1ad0-9cb6-446b-bef9-689b86ce0c60 · outbound

This paper cites Control Barrier Functions: Theory and Applications.

Safe Online Learning via Smooth Safety-Structured Policy Composition Control Barrier Functions: Theory and Applications

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:45:40.499189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:e4d68cbcd6c171c3fbad7ffecf4a2343aa1a85c9791ea47518633627dad2ad17

Observation 7f051ca2-e846-4547-8969-46da3a60e246 · outbound

This paper cites doi: 10.1145/3744351.

Safe Online Learning via Smooth Safety-Structured Policy Composition doi: 10.1145/3744351

Reference 2

Resolution
verified exact
doi, observed 2026-07-01T06:15:26.390161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:c61392b9be36bd5a3ee1aa04c9ffbd98b175a5a596e2f913f246309a75f678c4

Observation 38be1716-538c-4c67-aee2-ab0b15a0d9bc · outbound

This paper cites Lee, Matthew Tan, Yuke Zhu, and Jeannette Bohg.

Safe Online Learning via Smooth Safety-Structured Policy Composition Lee, Matthew Tan, Yuke Zhu, and Jeannette Bohg

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T06:15:26.394256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:61555dfbf2502f27d540933e381c2320953c6880058773d57d538741bf9515a2

Observation ceb67701-16a4-4252-8dda-dade35141081 · outbound

This paper cites Learning to Walk in the Real World with Minimal Human Effort.

Safe Online Learning via Smooth Safety-Structured Policy Composition Learning to Walk in the Real World with Minimal Human Effort

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:40.491380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:c5f51fcf0028c906c68f816c3eb8b64bb234f5c06eab4f93935060baa8b65d4a

Observation 8615bfd7-5425-4501-bc16-d0f9940f54e9 · outbound

This paper cites Annual Review of Control, Robotics, and Autonomous Systems7(2024).https://doi.org/10.1146/ANNUREV-CONTROL-071723-102940.

Safe Online Learning via Smooth Safety-Structured Policy Composition Annual Review of Control, Robotics, and Autonomous Systems7(2024).https://doi.org/10.1146/ANNUREV-CONTROL-071723-102940

Reference 5

Resolution
verified exact
doi, observed 2026-07-01T06:15:26.400128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:e2356e523123098925cf4f72e7e42432273c22a8e3a69f0dbf6b68e086785e3b

Observation 18af1742-2420-49d3-b5e6-59d05959123d · outbound

This paper cites Ibarz, J.

Safe Online Learning via Smooth Safety-Structured Policy Composition Ibarz, J

Reference 6

Resolution
metadata mismatch
doi, observed 2026-07-01T06:15:26.386008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:9d7158a75387a5b0da2e10d4ea9ae3e29a571ee7dfce55948110429e02a2a8ea

Observation cb52c102-f21b-463e-8509-e94104ead8ff · outbound

This paper cites cc/paper/2021/hash/85ea6fd7a2ca3960d0cf5201933ac998-Abstract.html.

Safe Online Learning via Smooth Safety-Structured Policy Composition cc/paper/2021/hash/85ea6fd7a2ca3960d0cf5201933ac998-Abstract.html

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T19:12:50.865649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:1069d4ee7638716522d92add27016807a3cafcf11da845c06e3c91ac2e2ed366

Observation 210dbb4f-da32-40ff-ba46-b8ffdd8c9943 · outbound

This paper cites Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking.

Safe Online Learning via Smooth Safety-Structured Policy Composition Provably Safe Reinforcement Learning: Conceptual Analysis, Survey, and Benchmarking

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:40.501783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:e436e64d078e8d337299edaaabc43dd1947281736fee8dfde632bbc257a77e25

Observation 20ef1545-8009-4f72-abe6-15934da4e59f · outbound

This paper cites Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?.

Safe Online Learning via Smooth Safety-Structured Policy Composition Safe Reinforcement Learning using Action Projection: Safeguard the Policy or the Environment?

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:45:40.504080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:5e8fc10839c329b89c148680458d5c200934714c0d396c8b2d772f7d6febf0c0

Observation ee1fb851-bcdc-4668-9f5c-065fe208ec7b · outbound

This paper cites arXiv preprint arXiv:2509.21014 (2025).

Safe Online Learning via Smooth Safety-Structured Policy Composition arXiv preprint arXiv:2509.21014 (2025)

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:40.488515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:97dc2f47d5e8307d3381022a05a07dabdeea6fbf5a0388fa689cc25d544f459a

Observation da4d45d8-3c5f-4df3-8dd0-cdea7da1886a · outbound

This paper cites In: NASA Formal Methods Symposium.

Safe Online Learning via Smooth Safety-Structured Policy Composition In: NASA Formal Methods Symposium

Reference 11

Resolution
metadata mismatch
doi, observed 2026-07-01T06:15:26.418657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:b50cd523817b879d93f3a7561b6762fa7cd3eca33e5a742dcf3ac06f55544203

Observation 637a075c-bb27-4f0c-89b6-dc5bcad91686 · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

Safe Online Learning via Smooth Safety-Structured Policy Composition Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 12

Resolution
malformed identifier
local_arxiv, observed 2026-07-01T09:45:40.506809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:b1b4ccd069c087022434adcdde8e25991b0c19cb1b8e3b3c87956c0ac0a56225

Observation 22ee596a-adbc-427e-9425-0243d819bfc2 · outbound

This paper cites Proximal Policy Optimization Algorithms.

Safe Online Learning via Smooth Safety-Structured Policy Composition Proximal Policy Optimization Algorithms

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T09:45:40.493661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:38aedb695a41078a0eef378e83b761475afc2b96f27d6db6e3a4d02fce657a8e

Observation 5a60eb90-bd82-4a9d-ad67-aaf7218c74dd · outbound

This paper cites an unresolved cited work.

Safe Online Learning via Smooth Safety-Structured Policy Composition Unresolved cited work

Reference 14

Resolution
verified exact
doi, observed 2026-07-01T06:15:26.396724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:ce6e3d953f4e2d59c0ada8426b051d036f7ceb7e430dcff36efeae3102d9a039

Observation d106a98a-4357-401f-8c5a-21a19dc16afe · outbound

This paper cites Reinforcement Learning with Adaptive Regularization for Safe Control of Critical Systems.

Safe Online Learning via Smooth Safety-Structured Policy Composition Reinforcement Learning with Adaptive Regularization for Safe Control of Critical Systems

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T09:45:40.498343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:9728641204368d7e5f8acbedad2b887e8488e67334a87620007440dce4518552

Observation 66b11d40-83c4-4d4a-8b22-81ff3d9dd91e · outbound

This paper cites Linear model predictive safety certification for learning-based control.

Safe Online Learning via Smooth Safety-Structured Policy Composition Linear model predictive safety certification for learning-based control

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T09:45:40.495697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:8cd215fd73a7242ceaeb50afcebec51690ff77535b6498ff91f07a8e1f13608b

Observation 582fa6b4-6588-4cf6-910e-19f539cd82cf · outbound

This paper cites ISBN 978-1-4503-1996-6.

Safe Online Learning via Smooth Safety-Structured Policy Composition ISBN 978-1-4503-1996-6

Reference 17

Resolution
verified exact
doi, observed 2026-07-01T06:15:26.382697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:d5259b78f67fb608b629849f3bb5efd191964abd06080229b899a132be80fe52

Observation c3aa37ef-0bb5-41e8-8925-9293bbb4cc38 · outbound

This paper cites Linhai Xie, Sen Wang, Stefano Rosa, Andrew Markham, and Niki Trigoni.

Safe Online Learning via Smooth Safety-Structured Policy Composition Linhai Xie, Sen Wang, Stefano Rosa, Andrew Markham, and Niki Trigoni

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:40.503798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:170b3ef0bf8c7979cdea0068d756648a67f4b946eb5fcbf86589da886ffec3c9

Observation 498762bc-ae7e-41fc-b637-80a375a18c4c · outbound

This paper cites Stable and Safe Reinforcement Learning via a Barrier-Lyapunov Actor-Critic Approach.

Safe Online Learning via Smooth Safety-Structured Policy Composition Stable and Safe Reinforcement Learning via a Barrier-Lyapunov Actor-Critic Approach

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-07-01T09:45:40.488370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:ae9498ccd5fd2ede4b5b9ce1b0c099bdacf931c2dfded8aaac75c67bbbc027d3

Observation 1c28926c-b38e-48d9-8234-37adc57e3e17 · outbound

This paper cites That is, there exist b(s)∈Randg(s)∈Rm such that ˆ∆(s,a) =b(s) +g(s) ⊤a+εlin(s,a),|εlin(s,a)|≤ϵlin(s)(22) for all actionsaon this segment.

Safe Online Learning via Smooth Safety-Structured Policy Composition That is, there exist b(s)∈Randg(s)∈Rm such that ˆ∆(s,a) =b(s) +g(s) ⊤a+εlin(s,a),|εlin(s,a)|≤ϵlin(s)(22) for all actionsaon this segment

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T19:12:50.873816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:69c69e9321330b036308860dc5e43685174cdd2d8d606a466c3ecd136c682231

Observation cae868ec-ab63-4c34-b105-f53be0792bca · outbound

This paper cites an unresolved cited work.

Safe Online Learning via Smooth Safety-Structured Policy Composition Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-07-06T19:12:50.889570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:78c428bc19bb20e0597dd99ade95f1917862687ca94dfcd500f16a835ab0de90

Observation d73eea61-3bfc-4ced-ad97-6aa64ca6c260 · outbound

This paper cites an unresolved cited work.

Safe Online Learning via Smooth Safety-Structured Policy Composition Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-07-06T19:12:50.882814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:2a9bf161f8b2afff493bf897f0124079e8f337f0a3a491abb6c93fe27bff1396

Observation b6865a8c-3410-4022-896d-835c977c9034 · outbound

This paper cites qCMdcjpSuxuYQzuZyBqqjO9S8DY=.

Safe Online Learning via Smooth Safety-Structured Policy Composition qCMdcjpSuxuYQzuZyBqqjO9S8DY=

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T19:12:50.886386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:01a442bec277f228002f52c799a7f690d69972975f77df449dfabaae65416d30

Observation e45dc118-bb26-487b-b5ea-27586c9bb944 · outbound

This paper cites an unresolved cited work.

Safe Online Learning via Smooth Safety-Structured Policy Composition Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-07-06T19:12:50.881211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:a4c055994c4c52c8fcb111b36f322174535b840dfef07b6e5650ab72febaff7c

Observation 56d335bd-1817-4bf2-a3f2-c1197373fce8 · outbound

This paper cites The system model can be found at (Tian et al., 2024).

Safe Online Learning via Smooth Safety-Structured Policy Composition The system model can be found at (Tian et al., 2024)

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T19:12:50.886281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:fbdb7b8b492ae42fbd7447e3d00ba80c8a105a01e56d960115c4305ad8440da6

Observation ae0de3cb-6074-4832-b6ea-d2fe5963d2c6 · outbound

This paper cites In our case study, we set the initial position of the quadrotor ass0xyz ={1.5,1.5,1.5}and the target position of the quadrotor asˆsxyz ={2.5,2.5,2.5}.

Safe Online Learning via Smooth Safety-Structured Policy Composition In our case study, we set the initial position of the quadrotor ass0xyz ={1.5,1.5,1.5}and the target position of the quadrotor asˆsxyz ={2.5,2.5,2.5}

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T19:12:50.878238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:a356ae79e6f1156a2be618ea67e64e13a3b4dc0a3c74db9182a6396871831428

Observation b7ae71bf-47ee-4e78-9539-9526fc9cf917 · outbound

This paper cites We observe that training of theSimplexis graduallydivergingwitha largecriticloss, asshown inFig.11.

Safe Online Learning via Smooth Safety-Structured Policy Composition We observe that training of theSimplexis graduallydivergingwitha largecriticloss, asshown inFig.11

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-06T19:12:50.875621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:6670f22f8ab911f28c629f17045a11fdc87d45d96e4b24ba0782a8d514e8d17d

Observation 2838b799-7e82-4d6b-97e1-696d672721dc · outbound

This paper cites an unresolved cited work.

Safe Online Learning via Smooth Safety-Structured Policy Composition Unresolved cited work

Reference 28

Resolution
malformed identifier
raw_fallback, observed 2026-07-06T19:12:50.883144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:0935e6ad75fa1a2b8b2fb10f0c18d7418ca7a6d7880cf8f04480dd230423183e

Observation 2e692e5f-97aa-449a-ba4a-11c885d7b2b0 · outbound

This paper cites For simple tasks, such as cartpole and glucose, the agent could learn using the data generated by the safe policy.

Safe Online Learning via Smooth Safety-Structured Policy Composition For simple tasks, such as cartpole and glucose, the agent could learn using the data generated by the safe policy

Reference 29

Resolution
malformed identifier
raw_fallback, observed 2026-07-06T19:12:50.867841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-01T06:12:07.125494Z digest=sha256:d159e3f0a1cbd7a9e347ac832e59c79922111cea645d7ccafc0ebb3a4e8d10e7

Pith citing papers

No inbound Pith citation observations are available.