Pith. sign in

Paper Citation Record · LEDGER

Q-learning-based Model-free Safety Filter

As of 15 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 1 inbound Pith citation observation for arXiv:2411.19809.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.19809 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T05:55:14.710636Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:35:50.981161Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T13:35:52.227123Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact6
  • verified fuzzy7
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9b7c3d74-eaa0-4e75-8e10-735f684148b2 · outbound

This paper cites Rapid Locomotion via Reinforcement Learning.

Q-learning-based Model-free Safety Filter Rapid Locomotion via Reinforcement Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.602434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.602434Z digest=sha256:666ddeeb3ef0c2c9a4c93517c85e642a893de4d019ea4821f9f20ec3eac6b478

Observation f7ee1373-3bf3-48f5-ad14-d924720bef00 · outbound

This paper cites Deep Reinforcement Learning for Robotic Manipulation with Asynchronous Off-Policy Updates.

Q-learning-based Model-free Safety Filter Deep Reinforcement Learning for Robotic Manipulation with Asynchronous Off-Policy Updates

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.606161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.606161Z digest=sha256:b0ad0aba27e464b995299d459897d78cb61d9de73b602f5fe236d0c652b6b606

Observation 46372c24-a799-4920-9c07-ec66d651a4f1 · outbound

This paper cites Barrier- certified adaptive reinforcement learning with applications to brushbot navigation,.

Q-learning-based Model-free Safety Filter Barrier- certified adaptive reinforcement learning with applications to brushbot navigation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.610100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.610100Z digest=sha256:4990ef03a5d5a63cb40de4e0b494432bf66a9a19b694fa2b0d2b8fdea28a5a35

Observation 5471a02c-9c3a-4a40-bc87-74051ea2fdca · outbound

This paper cites Safe reinforcement learning using robust control barrier functions,.

Q-learning-based Model-free Safety Filter Safe reinforcement learning using robust control barrier functions,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:55:15.082600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.613396Z digest=sha256:b2e2df2cb4e63a526e630c07b3a1f7e97eed19948e76a89c9f271923d9028922

Observation c4b7b39d-0e86-4723-8098-50dd7e88249b · outbound

This paper cites End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks.

Q-learning-based Model-free Safety Filter End-to-End Safe Reinforcement Learning through Barrier Functions for Safety-Critical Continuous Control Tasks

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.620452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.620452Z digest=sha256:4bb101f4d86bf7067456a38585526d37e88515d1f3fd34001e860c38e03653b9

Observation 858acfdf-2c74-47c8-94f9-2d8d33c1a165 · outbound

This paper cites Probabilistic model predictive safety certification for learning-based control.

Q-learning-based Model-free Safety Filter Probabilistic model predictive safety certification for learning-based control

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:55:14.901804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.624401Z digest=sha256:1d7b847849c60f78016f35f8b82c30536c97e800b539f0e75e1d5aaa4666ef9d

Observation 9327079e-9592-4a50-8d16-f357ece1aa49 · outbound

This paper cites A General Safety Framework for Learning-Based Control in Uncertain Robotic Systems.

Q-learning-based Model-free Safety Filter A General Safety Framework for Learning-Based Control in Uncertain Robotic Systems

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.627923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.627923Z digest=sha256:e3a6971969e6d57201339bb2e5c1a7ed3f48feb18e921d992f81d9a28dc86881

Observation 6e3901c7-0b86-4289-b347-5af2be9e7cab · outbound

This paper cites Agile But Safe: Learning Collision-Free High-Speed Legged Locomotion.

Q-learning-based Model-free Safety Filter Agile But Safe: Learning Collision-Free High-Speed Legged Locomotion

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.631395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.631395Z digest=sha256:8091f1a4141622bebb3503289fe72823f5383969f00dd738fba12e93bedc2bd4

Observation ac379a51-0a69-4ca4-bb16-2d0bc11f4b42 · outbound

This paper cites Constrained policy optimization,.

Q-learning-based Model-free Safety Filter Constrained policy optimization,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:55:15.073760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.634750Z digest=sha256:759e875f9442c2c2d7a8f9330a1289a4af423e1917cdaaaf0b44f4d8be0e8d78

Observation 718ed895-76f5-4502-98ff-c3d91556111f · outbound

This paper cites Value functions are control barrier functions: Verification of safe policies using control theory.

Q-learning-based Model-free Safety Filter Value functions are control barrier functions: Verification of safe policies using control theory

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:55:15.064669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.638267Z digest=sha256:9db8882b9b7ff0bd33bb03dd1ca15c4fbb028284b91129c67e5eb94c55aa8159

Observation e06bc79a-b879-4f01-b3c3-6ac56719ca4a · outbound

This paper cites Bridging hamilton-jacobi safety analysis and reinforcement learning,.

Q-learning-based Model-free Safety Filter Bridging hamilton-jacobi safety analysis and reinforcement learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:55:15.053967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.645900Z digest=sha256:a1fd99a4a202f481773f0d6e3d9aefdbf0af967c96ef217cff6468911354deb7

Observation 84a06bb5-228b-432c-8639-997a0bf8979a · outbound

This paper cites Altman, Constrained Markov Decision Processes.

Q-learning-based Model-free Safety Filter Altman, Constrained Markov Decision Processes

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:55:15.043635Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.649013Z digest=sha256:44249510fe35692ab9f2b11bf98f4eccd47b1ec1047cad45efe2980bbf1658f6

Observation 18533d8f-d73e-4b17-bec4-aa740caec948 · outbound

This paper cites Adaptive dynamic programming for nonaffine nonlinear optimal control problem with state constraints,.

Q-learning-based Model-free Safety Filter Adaptive dynamic programming for nonaffine nonlinear optimal control problem with state constraints,

Reference 15

Resolution
malformed identifier
no resolver link, observed 2026-08-12T05:55:14.651966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.651966Z digest=sha256:85c753768efb772615ee9ad161e6c2e59930f0dd4e0a1a80af3a87fc048257e1

Observation ffada0da-e970-4f27-a5f1-4c090bbfc393 · outbound

This paper cites Reward Constrained Policy Optimization.

Q-learning-based Model-free Safety Filter Reward Constrained Policy Optimization

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.654949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.654949Z digest=sha256:dd77765aa52c054eb895291f18eab285ee9caaa49d6517cd680b2df789ba394d

Observation 42fabd6b-c9c7-491f-9087-fd471dbf8f9d · outbound

This paper cites Benchmarking safe exploration in deep reinforcement learning,.

Q-learning-based Model-free Safety Filter Benchmarking safe exploration in deep reinforcement learning,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.657949Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.657949Z digest=sha256:aa924b7c4b91f89cb85fe5504590c4519aa917b01bbc37931d72b26a21bdd96d

Observation d0cb7765-bfe3-436d-9425-50d9c8bfde69 · outbound

This paper cites Integrated Decision and Control: Towards Interpretable and Computationally Efficient Driving Intelligence.

Q-learning-based Model-free Safety Filter Integrated Decision and Control: Towards Interpretable and Computationally Efficient Driving Intelligence

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:55:14.842867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.666742Z digest=sha256:8afab2682dad11b08f60825a35adc91daf06ac34dc5f94404267125e8c55f93f

Observation 5fc2dc98-115d-4ab1-8287-2facf86f4007 · outbound

This paper cites Learning to be Safe: Deep RL with a Safety Critic.

Q-learning-based Model-free Safety Filter Learning to be Safe: Deep RL with a Safety Critic

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.669840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.669840Z digest=sha256:e8ae963c619653621684f1aeb728443ec50b028dfa36ddfa587160418194c44a

Observation c01971df-14d8-4442-b45e-ff6d2f050f3a · outbound

This paper cites Re- covery rl: Safe reinforcement learning with learned recovery zones,.

Q-learning-based Model-free Safety Filter Re- covery rl: Safe reinforcement learning with learned recovery zones,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:55:15.025971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.672985Z digest=sha256:049da1d6b10e49d0039daf41c44577bf7aadd83c3bdaf8aa8ae7353ee0d88324

Observation e3dccc95-03b9-47d7-b388-d61c78a5f02c · outbound

This paper cites Safety Optimized Reinforcement Learning via Multi-Objective Policy Optimization.

Q-learning-based Model-free Safety Filter Safety Optimized Reinforcement Learning via Multi-Objective Policy Optimization

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:55:14.821779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.675768Z digest=sha256:9cdb76239d8e41666ad1cd071f070ee45c563c917eaf0cf04302ba6570535606

Observation cecd247b-e3da-4f8c-842e-df1afbc8b5bc · outbound

This paper cites Safe Reinforcement Learning by Imagining the Near Future.

Q-learning-based Model-free Safety Filter Safe Reinforcement Learning by Imagining the Near Future

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:55:14.809338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.678976Z digest=sha256:d03f0751b57efce87f97fb570aa5a5d97cc63c1e5399fd70ecfc2522aadaffac

Observation 70fa88f3-eb34-4674-ab67-d16b00b05eac · outbound

This paper cites Learning safety critics via a non-contractive binary bellman operator.

Q-learning-based Model-free Safety Filter Learning safety critics via a non-contractive binary bellman operator

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:55:14.871426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.682096Z digest=sha256:cc5d7956cefeca547fca4c0b1413f92b61a80a8d0319607fa48824e8c6aa47f8

Observation 2c8fc713-064f-4d8f-b960-2555d1c80c68 · outbound

This paper cites Constrained Policy Optimization via Bayesian World Models.

Q-learning-based Model-free Safety Filter Constrained Policy Optimization via Bayesian World Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.684951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.684951Z digest=sha256:a03fcb5ec1b0ecb6a5f2793db59ee4813d5bb08c989252fc36a6616390540eb7

Observation d26d5cb9-9839-4e09-af70-96ef11405211 · outbound

This paper cites Safe Continuous Control with Constrained Model-Based Policy Optimization.

Q-learning-based Model-free Safety Filter Safe Continuous Control with Constrained Model-Based Policy Optimization

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-08-12T05:55:14.788931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.688366Z digest=sha256:c044ed383c86aee6cc77d82a1f17d54dfe1bd71ad5cdacd972a03a4149418858

Observation 0b7932dc-7da0-4293-a712-d0e32fa210a3 · outbound

This paper cites an unresolved cited work.

Q-learning-based Model-free Safety Filter Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.691637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.691637Z digest=sha256:33f62315b5a26627fb6fca205c52de43f78042919124cfaba401cf61fd9b7d3b

Observation 09c2efcb-a79d-47fa-b98e-8ed4799321c8 · outbound

This paper cites Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor.

Q-learning-based Model-free Safety Filter Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.694667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.694667Z digest=sha256:e13fd20c1e94db9d7f58902331dc051e717ffe499ec85b6fa3ad608eb0e0ab3b

Observation be8e5220-3762-43f4-8493-954afa1db199 · outbound

This paper cites Playing Atari with Deep Reinforcement Learning.

Q-learning-based Model-free Safety Filter Playing Atari with Deep Reinforcement Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.697770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.697770Z digest=sha256:fbf20622a02886415d50c17c4f53b6947d4a33a101029fb58d795e0b83f37245

Observation 2a6b6002-6338-43b3-ab35-bd2b7df9dfed · outbound

This paper cites Proximal Policy Optimization Algorithms.

Q-learning-based Model-free Safety Filter Proximal Policy Optimization Algorithms

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.700908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.700908Z digest=sha256:0a4ef77ce9264ac5fdcc6fd242706d9e618c9ec30681f2dad88b43fa10e9a0c1

Observation f9e2ffd1-65e2-486f-b06a-2db503f9bb38 · outbound

This paper cites Continuous control with deep reinforcement learning.

Q-learning-based Model-free Safety Filter Continuous control with deep reinforcement learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.704002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.704002Z digest=sha256:05d9ffeca1645191b5ea2dc8f1fcb9073ad048aac3d2e6cc38e526434ef5d6ba

Observation 6fc480f0-a8f1-4e36-a9fc-4cb1f5df8264 · outbound

This paper cites Addressing Function Approximation Error in Actor-Critic Methods.

Q-learning-based Model-free Safety Filter Addressing Function Approximation Error in Actor-Critic Methods

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.707219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.707219Z digest=sha256:07b0e55ca3e68c52aa28c1b826e36a3b1d280a4ae1af4e3bd23f388b7e6c29b7

Observation 266b26f6-922f-4181-a350-432b9eff8ec6 · outbound

This paper cites Optimal control for a shape memory alloy actuated soft digit using iterative learning control,.

Q-learning-based Model-free Safety Filter Optimal control for a shape memory alloy actuated soft digit using iterative learning control,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T05:55:15.008728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T05:55:14.710636Z digest=sha256:9b57e1778862689bba424e8569a26469d99bce899fc4059bf49542bda85988df

Observation d1474a7f-e3fa-4f85-b16c-05acae6b17cd · outbound

This paper cites Model-Based Reinforcement Learning for Atari.

Q-learning-based Model-free Safety Filter Model-Based Reinforcement Learning for Atari

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.599070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.599070Z digest=sha256:83a343d78797c7f62f0326f9955c46060a2873433fc34d0e487cc12ba4bd982f

Observation 57d2901b-701d-40f7-88a2-c35fb3078c04 · outbound

This paper cites Projection-Based Constrained Policy Optimization.

Q-learning-based Model-free Safety Filter Projection-Based Constrained Policy Optimization

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.664113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.664113Z digest=sha256:fc274943a54e5ec795934cb54b767cbb6f429f6b5fa716980bc25a5ed8dcec1f

Observation fefe6960-2dc7-4e1d-af49-7c0ed08d36d4 · outbound

This paper cites Safe Reinforcement Learning Using Robust Control Barrier Functions.

Q-learning-based Model-free Safety Filter Safe Reinforcement Learning Using Robust Control Barrier Functions

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T05:55:14.616781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T05:55:14.616781Z digest=sha256:e5d07d8bbb8d2f40f1cc529c6a79688b492db71b7b29ef2baddda65be22043af

Pith citing papers

Observation 85f2e082-e725-4692-8836-3ec34871f147 · inbound

Verifiable Safety Q-Filters via Hamilton-Jacobi Reachability and Multiplicative Q-Networks cites this paper.

Verifiable Safety Q-Filters via Hamilton-Jacobi Reachability and Multiplicative Q-Networks Q-learning-based Model-free Safety Filter

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:35:52.284753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-07T13:35:50.981161Z digest=sha256:04b613b768077391175c8dbacbfbb501a06bc14c8f3089031f74e6b740c47db6