Pith. sign in

Paper Citation Record · LEDGER

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 3 inbound Pith citation observations for arXiv:2507.00485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00485 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:20:29.473633Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:11:29.709863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:07:28.391121Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 811bcd8d-b9a8-448b-8aae-9c2e93087fd6 · outbound

This paper cites Constrained policy optimiza- tion.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Constrained policy optimiza- tion

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.673697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.382296Z digest=sha256:611446b5198e68b1a5783d700d0519b2ad4f33d752833fc8f3a441b6738191ea

Observation 7d838393-106d-4239-9e0d-dc73312485de · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.412845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.412845Z digest=sha256:cb238245bf2a2ee503d2870dbebc65c8ddb9790224222e01020589557c1d7b30

Observation 0b917f32-af8c-465d-89e6-100f6087d487 · outbound

This paper cites Enhancing the robustness of qmix against state-adversarial attacks.Neurocomputing, 572:127191,.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Enhancing the robustness of qmix against state-adversarial attacks.Neurocomputing, 572:127191,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.674761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.422276Z digest=sha256:0b3ad4ba1e81f097db03d9eff9ea8ae6a6d0e98dfba593d4f6603f7a9a365052

Observation 2cb0052d-2385-4b95-8162-4339f3f95b34 · outbound

This paper cites Robust training in multiagent deep reinforcement learning against optimal adversary.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Robust training in multiagent deep reinforcement learning against optimal adversary

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.665359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.425982Z digest=sha256:8d72759dc4ce815cffb8ba42ed8624d6d99a3da7b16c35f8f5ac7515377fba47

Observation cb575ed0-9d50-41b5-bbc9-64257d2373f2 · outbound

This paper cites Backdoor attacks on safe reinforcement learning- enabled cyber–physical systems.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Backdoor attacks on safe reinforcement learning- enabled cyber–physical systems

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.632972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.435260Z digest=sha256:34d05f67ee66f8cf5417e36cf8ae233684fdb61b063479d0da5a9eae1eb76f19

Observation a9bf24f5-be90-40d1-9626-508c8ce35f1a · outbound

This paper cites Trojdrl: Evaluation of back- door attacks on deep reinforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Trojdrl: Evaluation of back- door attacks on deep reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.621937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.437848Z digest=sha256:f562cf033d8bbde07b3ba377565b83fbf286f8181c50fcfe9d0871e8eb44892b

Observation 95836e9a-628d-41a0-afd8-5ef045ce71b3 · outbound

This paper cites Con- strained variational policy optimization for safe reinforce- ment learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Con- strained variational policy optimization for safe reinforce- ment learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.611753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.444969Z digest=sha256:1adf8bdeca96f6679b79ba19885ca5ec4428b7ce43b06f6694aad0d686b31268

Observation 7b450561-be08-4f34-a4e9-55a1dec67e0a · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Towards deep learning models resistant to adversarial attacks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.601839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.448131Z digest=sha256:fe41a00a2c00faa3e3beaa243591ed4f8b4d3dc574f1300cebfcfdb2e71e6579

Observation dd5e120c-00f2-4d39-b49c-0cc62563a593 · outbound

This paper cites Marl sim2real transfer: Merging physical reality with digital virtuality in meta- verse.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Marl sim2real transfer: Merging physical reality with digital virtuality in meta- verse

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.591243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.451164Z digest=sha256:3da6813e2391c3db8d73ab51285fca4fec8ce5f3ca2e77cdf735effd462ff82a

Observation 7b95e5ce-3118-4800-b57f-0e94d1ee0e4c · outbound

This paper cites Responsive safety in reinforcement learn- ing by PID lagrangian methods.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Responsive safety in reinforcement learn- ing by PID lagrangian methods

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.580573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.454841Z digest=sha256:a7cc27072c848371ff86629c7a3a72eaaa4718fb0f77187c2a2444d05e924bf3

Observation 23b3a5a9-d5fa-49ff-967e-d5a4953c7e70 · outbound

This paper cites Mankowitz, and Shie Mannor.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Mankowitz, and Shie Mannor

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.570764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.458015Z digest=sha256:fe41c2705027b36965793c0158d0a77d60d5ebfa0b228244c445f0fb3ac82cd6

Observation 318e91ba-365c-41eb-aefa-2ba262f9fbc5 · outbound

This paper cites Backdoorl: Backdoor attack against competitive reinforcement learn- ing.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Backdoorl: Backdoor attack against competitive reinforcement learn- ing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.561421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.461226Z digest=sha256:30d24743c260e57904e9913821c91cebeef2904fe812abc5f212d7e1c082e02f

Observation 6c2d81ce-7fbd-4c6a-97b1-bac205a7e6c5 · outbound

This paper cites Partially observable mean field multi- agent reinforcement learning based on graph attention net- work for uav swarms.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Partially observable mean field multi- agent reinforcement learning based on graph attention net- work for uav swarms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.552002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.466839Z digest=sha256:f91503807a1de121e2c8bb80298c9325818cc005e7d8c4445e4e427127d7d205

Observation ca4b967f-7f06-4d96-b152-23f71921dd26 · outbound

This paper cites First order constrained optimization in policy space.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning First order constrained optimization in policy space

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.542180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.469695Z digest=sha256:c0a46e049d27d99308bfd65dd4b562013e9da409366aa33fed8a213f882f7d2c

Observation 97f76f8e-0a90-4f66-9a50-aacb331ddb06 · outbound

This paper cites A robust mean-field actor-critic rein- forcement learning against adversarial perturbations on agent states.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning A robust mean-field actor-critic rein- forcement learning against adversarial perturbations on agent states

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.531670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.473633Z digest=sha256:3083eb577104a1efd0e5c0f65c95dc746c2fb6666cf364091ff1b41018b8947f

Observation 3ebc93a3-82df-491e-a0d5-d1d7ab5a2b9c · outbound

This paper cites Safety gymna- sium: A unified safe reinforcement learning benchmark.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Safety gymna- sium: A unified safe reinforcement learning benchmark

Reference 1994

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.642657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.431771Z digest=sha256:e31d0633160d197e6da53b56512aeb17f811a0d8bb290c43317ada0f3a3e51f9

Observation 6f5a24a6-b521-4a36-b29c-d2e8c3d88137 · outbound

This paper cites Constrained policy optimiza- tion via bayesian world models.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Constrained policy optimiza- tion via bayesian world models

Reference 1998

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.651583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.388852Z digest=sha256:ea6fd50cb738cce1f64007faec3bbce5bcf700d43b994587b5c79a6ce610689a

Observation c2955608-c556-44a5-96dd-e1e86def435f · outbound

This paper cites Context-aware safe reinforcement learning for non- stationary environments.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Context-aware safe reinforcement learning for non- stationary environments

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.618862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.399212Z digest=sha256:4ed02fe489dfb8128bba992180170dd941bb9ae69c54aab11a1fd683d3165558

Observation 0b3580d5-cdf5-45dd-8836-748dda812fe9 · outbound

This paper cites an unresolved cited work.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Unresolved cited work

Reference 2012

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:20:30.629309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.396002Z digest=sha256:243bc6ae35ccbfe061b09ba3b9200a2f320783d7e4313d5d2ada602664284b4d

Observation b5722f24-2869-49cc-9651-df085326ca0d · outbound

This paper cites Policycleanse: Backdoor detection and mitiga- tion for competitive reinforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Policycleanse: Backdoor detection and mitiga- tion for competitive reinforcement learning

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.686052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.419570Z digest=sha256:8d64c424a1f7b7277304f08316603a9b3aa7bb4ca76157fc5791db88dda9157d

Observation 67b1924f-a84d-4766-b9f7-3a45409974a0 · outbound

This paper cites Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.662359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.385947Z digest=sha256:97d652132b4284f4e93274287b5661e94bb87ec60b36a3684f345ed01c12d94f

Observation 66f54b19-7271-4baa-84ea-a223e1c319cc · outbound

This paper cites Badrl: Sparse targeted backdoor attack against reinforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Badrl: Sparse targeted backdoor attack against reinforcement learning

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.585085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.409436Z digest=sha256:8ef08d122f57fa6cb25459b64b66bf6446f00ad9e6da49a240049d073c975752

Observation 2eaad8fe-ae38-46ee-bc78-4adf6c952ebe · outbound

This paper cites Goodfellow, Jonathon Shlens, and Christian Szegedy.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Goodfellow, Jonathon Shlens, and Christian Szegedy

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.695760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.416677Z digest=sha256:817d1fd4c1c6a3aff1f5ae3a311d815f741dc35e1d4a2dab00433b91bcd5597d

Observation 7f1d205e-0fd2-4481-bc3e-a208e4e89b08 · outbound

This paper cites Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.441050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.441050Z digest=sha256:4199d03e9a917f01fef47612ba42a9c2f61520cf69e8ca42cb5227ad6743b1f0

Observation c45262ad-fe4e-41fd-b2cf-23cf9243c6ec · outbound

This paper cites Projection-Based Constrained Policy Optimization.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Projection-Based Constrained Policy Optimization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.464007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.464007Z digest=sha256:e78a6110af04cf366c702b071b19ed41b5224dbbe224b023c16325d482df4d6c

Observation 07ac8b65-1c42-4f9a-994f-04fea91ff6ba · outbound

This paper cites An online actor–critic algorithm with function approximation for constrained markov decision processes.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning An online actor–critic algorithm with function approximation for constrained markov decision processes

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.639990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.392517Z digest=sha256:529658990146b8cc1ea6c08825abb63291157574abccc04f79562d8e0649a0f2

Observation 3a957276-6d76-4cf5-89eb-9270ba8b54bd · outbound

This paper cites Robust multi- agent reinforcement learning method based on adversar- ial domain randomization for real-world dual-uav co- operation.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Robust multi- agent reinforcement learning method based on adversar- ial domain randomization for real-world dual-uav co- operation

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.607737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.403636Z digest=sha256:59b64764eea3517403873cf7eca50baa9bb576c66cd4896034275a5ef7d67c5a

Observation f538e655-37da-4bcc-b5d4-43d5ddcaee51 · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Risk-constrained reinforcement learning with percentile risk criteria

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.597019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.406412Z digest=sha256:8ba790f14532ee51ca1dc5381bf918e5184d7dc8809bbe5e3da366ab92ed7f79

Observation a0ef48ee-e03a-45d8-917b-b50a1fb35b02 · outbound

This paper cites Consideration of risk in re- inforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Consideration of risk in re- inforcement learning

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.653990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T21:20:29.429077Z digest=sha256:d39750d557dba32b8fd7c90390cbef093dda6a6ceb015027693d0b4321e124b8

Pith citing papers

Observation e9f5d610-b76f-44a1-8cda-dbe946532e22 · inbound

Dataset Poisoning Attacks on Behavioral Cloning Policies cites this paper.

Dataset Poisoning Attacks on Behavioral Cloning Policies PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:11:29.709863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:11:29.709863Z digest=sha256:85e4b3e91d6cedf977df71699df35061e7c0b6140958b76eea60135c1581c464

Observation ed111060-4680-4267-9b22-ee791dc26641 · inbound

Trojan Attacks on Neural Network Controllers for Robotic Systems cites this paper.

Trojan Attacks on Neural Network Controllers for Robotic Systems PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:24:01.980022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:24:01.980022Z digest=sha256:a8cf666ad16f4837b40a584b01e568592987be41d233a154f80e85cab0f4a622

Observation a0217317-2144-4599-956a-f6d5e1852ac3 · inbound

Safe-RULE: Safe Reinforcement UnLEarning cites this paper.

Safe-RULE: Safe Reinforcement UnLEarning PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:07:28.392658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-06-27T17:28:09.686163Z digest=sha256:e25206a2f94acecdd0c583f53d66fd60fe365e1ce6330544e663de11b1c86bb7