Pith. sign in

Paper Citation Record · LEDGER

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

As of 16 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 3 inbound Pith citation observations for arXiv:2507.00485.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00485 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:20:29.473633Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T20:11:29.709863Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T00:07:28.391121Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy25
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 811bcd8d-b9a8-448b-8aae-9c2e93087fd6 · outbound

This paper cites Constrained policy optimiza- tion.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Constrained policy optimiza- tion

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.673697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.382296Z digest=sha256:928027f82f040d41fce8d9c9614ee65c44dc71a628a0c92efdf1a0e6ac891234

Observation 7d838393-106d-4239-9e0d-dc73312485de · outbound

This paper cites Benchmarking Batch Deep Reinforcement Learning Algorithms.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Benchmarking Batch Deep Reinforcement Learning Algorithms

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.412845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.412845Z digest=sha256:cb238245bf2a2ee503d2870dbebc65c8ddb9790224222e01020589557c1d7b30

Observation 0b917f32-af8c-465d-89e6-100f6087d487 · outbound

This paper cites Enhancing the robustness of qmix against state-adversarial attacks.Neurocomputing, 572:127191,.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Enhancing the robustness of qmix against state-adversarial attacks.Neurocomputing, 572:127191,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.674761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.422276Z digest=sha256:7b3887da979b4b2e0e0384da23afd65697ada96519bb8c4b802616295d09dd5c

Observation 2cb0052d-2385-4b95-8162-4339f3f95b34 · outbound

This paper cites Robust training in multiagent deep reinforcement learning against optimal adversary.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Robust training in multiagent deep reinforcement learning against optimal adversary

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.665359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.425982Z digest=sha256:2c1f356666183fabe863d261f0dac1a1ae7d1d081d3849ab6298ade219fdf62d

Observation cb575ed0-9d50-41b5-bbc9-64257d2373f2 · outbound

This paper cites Backdoor attacks on safe reinforcement learning- enabled cyber–physical systems.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Backdoor attacks on safe reinforcement learning- enabled cyber–physical systems

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.632972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.435260Z digest=sha256:ac9e9157e29101f93cbb1988f8802b2379b04986994c1510eade997e5fc52d17

Observation a9bf24f5-be90-40d1-9626-508c8ce35f1a · outbound

This paper cites Trojdrl: Evaluation of back- door attacks on deep reinforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Trojdrl: Evaluation of back- door attacks on deep reinforcement learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.621937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.437848Z digest=sha256:c276bd0333be05a6ab07ee91cbfff753f85fac57df35395c58f0802715b399c0

Observation 95836e9a-628d-41a0-afd8-5ef045ce71b3 · outbound

This paper cites Con- strained variational policy optimization for safe reinforce- ment learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Con- strained variational policy optimization for safe reinforce- ment learning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.611753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.444969Z digest=sha256:779adef5886b299cd8feb6ee3222262a205772dfab2fb93828e0193baacb8e3f

Observation 7b450561-be08-4f34-a4e9-55a1dec67e0a · outbound

This paper cites Towards deep learning models resistant to adversarial attacks.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Towards deep learning models resistant to adversarial attacks

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.601839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.448131Z digest=sha256:7f6d40a59bf0bc28a422283975211ad5618388e6be87ef30d124fd7f570acb5a

Observation dd5e120c-00f2-4d39-b49c-0cc62563a593 · outbound

This paper cites Marl sim2real transfer: Merging physical reality with digital virtuality in meta- verse.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Marl sim2real transfer: Merging physical reality with digital virtuality in meta- verse

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.591243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.451164Z digest=sha256:de4f319ffa90d6a8e99a87d71855777d93c4db309692f530cb526b37ad5f50c4

Observation 7b95e5ce-3118-4800-b57f-0e94d1ee0e4c · outbound

This paper cites Responsive safety in reinforcement learn- ing by PID lagrangian methods.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Responsive safety in reinforcement learn- ing by PID lagrangian methods

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.580573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.454841Z digest=sha256:a5322d2319e6554048cbee5afd541ee4e1dee67e946e0aa5a6cac6f0be5b3ddf

Observation 23b3a5a9-d5fa-49ff-967e-d5a4953c7e70 · outbound

This paper cites Mankowitz, and Shie Mannor.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Mankowitz, and Shie Mannor

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.570764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.458015Z digest=sha256:3b814986ac0682fef95cb18390e241bec3876adb7a2cfe2e32e4d0d5e2e7857b

Observation 318e91ba-365c-41eb-aefa-2ba262f9fbc5 · outbound

This paper cites Backdoorl: Backdoor attack against competitive reinforcement learn- ing.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Backdoorl: Backdoor attack against competitive reinforcement learn- ing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.561421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.461226Z digest=sha256:2038c7ef10c8db81293bb25fe9e10a0381c6f5ae009c55b5f0cc1014188bb114

Observation 6c2d81ce-7fbd-4c6a-97b1-bac205a7e6c5 · outbound

This paper cites Partially observable mean field multi- agent reinforcement learning based on graph attention net- work for uav swarms.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Partially observable mean field multi- agent reinforcement learning based on graph attention net- work for uav swarms

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.552002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.466839Z digest=sha256:a39b5d815de3ee0817339249e3feb4a0b48cd02f2636d7a35d11363be42ffe35

Observation ca4b967f-7f06-4d96-b152-23f71921dd26 · outbound

This paper cites First order constrained optimization in policy space.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning First order constrained optimization in policy space

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.542180Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.469695Z digest=sha256:2fc2b84ebee088867d7f8b092d127228f0bd9f456adbb65e491bb6d94d6c0e4e

Observation 97f76f8e-0a90-4f66-9a50-aacb331ddb06 · outbound

This paper cites A robust mean-field actor-critic rein- forcement learning against adversarial perturbations on agent states.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning A robust mean-field actor-critic rein- forcement learning against adversarial perturbations on agent states

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.531670Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.473633Z digest=sha256:570dacabeb23a804804d9d18b4c4a95a393338c39a4e737a3fd4b7926be33204

Observation 3ebc93a3-82df-491e-a0d5-d1d7ab5a2b9c · outbound

This paper cites Safety gymna- sium: A unified safe reinforcement learning benchmark.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Safety gymna- sium: A unified safe reinforcement learning benchmark

Reference 1994

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.642657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.431771Z digest=sha256:d9b8ee0d811e080becbaf7c6586381c564553adbaecb0918820234ab89a4754e

Observation 6f5a24a6-b521-4a36-b29c-d2e8c3d88137 · outbound

This paper cites Constrained policy optimiza- tion via bayesian world models.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Constrained policy optimiza- tion via bayesian world models

Reference 1998

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.651583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.388852Z digest=sha256:456ba613cadbe4cc9bd939405d23e5aadbb2ea235a446780157648fff62dcde6

Observation c2955608-c556-44a5-96dd-e1e86def435f · outbound

This paper cites Context-aware safe reinforcement learning for non- stationary environments.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Context-aware safe reinforcement learning for non- stationary environments

Reference 2005

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.618862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.399212Z digest=sha256:ba395a13a20dd7a748c462ef1b6ebb1528b52b32ed77dc89e9d2663a33cc3974

Observation 0b3580d5-cdf5-45dd-8836-748dda812fe9 · outbound

This paper cites an unresolved cited work.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Unresolved cited work

Reference 2012

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:20:30.629309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.396002Z digest=sha256:88c3864a92ba57c1b1cbea211c1c9987ea5156effa284fe3f96cf5c4bfdf14b9

Observation b5722f24-2869-49cc-9651-df085326ca0d · outbound

This paper cites Policycleanse: Backdoor detection and mitiga- tion for competitive reinforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Policycleanse: Backdoor detection and mitiga- tion for competitive reinforcement learning

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.686052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.419570Z digest=sha256:b661ccb719c1a0fb704ead608d0f0d5ae71f7f5ab5ce3e2c8bfe22f8e2d60e63

Observation 67b1924f-a84d-4766-b9f7-3a45409974a0 · outbound

This paper cites Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Constrained markov decision processes with total cost criteria: Lagrangian approach and dual linear program

Reference 2017

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.662359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.385947Z digest=sha256:3a9ec57252f433d325352770b92d5ea682e7ebf460d2cbac0d4612c86aaa1e28

Observation 66f54b19-7271-4baa-84ea-a223e1c319cc · outbound

This paper cites Badrl: Sparse targeted backdoor attack against reinforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Badrl: Sparse targeted backdoor attack against reinforcement learning

Reference 2018

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.585085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.409436Z digest=sha256:028feab563385d8ba87b293141cf026b896659f8c2686f1ac2ba0c8369665199

Observation 2eaad8fe-ae38-46ee-bc78-4adf6c952ebe · outbound

This paper cites Goodfellow, Jonathon Shlens, and Christian Szegedy.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Goodfellow, Jonathon Shlens, and Christian Szegedy

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.695760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.416677Z digest=sha256:61b4ebcca77e2f763f1843632c644d2b1b952bdbbdf94fce5448298bc8e93088

Observation 7f1d205e-0fd2-4481-bc3e-a208e4e89b08 · outbound

This paper cites Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Accelerated Primal-Dual Policy Optimization for Safe Reinforcement Learning

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.441050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.441050Z digest=sha256:4199d03e9a917f01fef47612ba42a9c2f61520cf69e8ca42cb5227ad6743b1f0

Observation c45262ad-fe4e-41fd-b2cf-23cf9243c6ec · outbound

This paper cites Projection-Based Constrained Policy Optimization.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Projection-Based Constrained Policy Optimization

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:20:29.464007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:20:29.464007Z digest=sha256:e78a6110af04cf366c702b071b19ed41b5224dbbe224b023c16325d482df4d6c

Observation 07ac8b65-1c42-4f9a-994f-04fea91ff6ba · outbound

This paper cites An online actor–critic algorithm with function approximation for constrained markov decision processes.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning An online actor–critic algorithm with function approximation for constrained markov decision processes

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.639990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.392517Z digest=sha256:3b669e9923b91dca759a63971ff22d52a2e51c140667bbf5ab6215674ba3013a

Observation 3a957276-6d76-4cf5-89eb-9270ba8b54bd · outbound

This paper cites Robust multi- agent reinforcement learning method based on adversar- ial domain randomization for real-world dual-uav co- operation.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Robust multi- agent reinforcement learning method based on adversar- ial domain randomization for real-world dual-uav co- operation

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.607737Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.403636Z digest=sha256:57bd01ac4086a45d28bdf4afbe9bbad7b697f0a6334ae0997a9d47c355a6f048

Observation f538e655-37da-4bcc-b5d4-43d5ddcaee51 · outbound

This paper cites Risk-constrained reinforcement learning with percentile risk criteria.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Risk-constrained reinforcement learning with percentile risk criteria

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:30.597019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.406412Z digest=sha256:8de246b3eeca079950d4e1835e1c5c1a496e954321a7c59d0c53280dc5c3f1d8

Observation a0ef48ee-e03a-45d8-917b-b50a1fb35b02 · outbound

This paper cites Consideration of risk in re- inforcement learning.

PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning Consideration of risk in re- inforcement learning

Reference 2025

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:20:29.653990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T21:20:29.429077Z digest=sha256:a573e5abad9e227200549f06c5c149e2426bb4744e3f45544a746b8147aae1f9

Pith citing papers

Observation e9f5d610-b76f-44a1-8cda-dbe946532e22 · inbound

Dataset Poisoning Attacks on Behavioral Cloning Policies cites this paper.

Dataset Poisoning Attacks on Behavioral Cloning Policies PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T20:11:29.709863Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T20:11:29.709863Z digest=sha256:85e4b3e91d6cedf977df71699df35061e7c0b6140958b76eea60135c1581c464

Observation ed111060-4680-4267-9b22-ee791dc26641 · inbound

Trojan Attacks on Neural Network Controllers for Robotic Systems cites this paper.

Trojan Attacks on Neural Network Controllers for Robotic Systems PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T04:24:01.980022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T04:24:01.980022Z digest=sha256:a8cf666ad16f4837b40a584b01e568592987be41d233a154f80e85cab0f4a622

Observation a0217317-2144-4599-956a-f6d5e1852ac3 · inbound

Safe-RULE: Safe Reinforcement UnLEarning cites this paper.

Safe-RULE: Safe Reinforcement UnLEarning PNAct: Crafting Backdoor Attacks in Safe Reinforcement Learning

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-07-03T00:07:28.392658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-27T17:28:09.686163Z digest=sha256:6b8dfb96c73a3e173e2392f00a213523a57469407c7a31e033bb1f52584c633c