Pith. sign in

Paper Citation Record · LEDGER

Agent Safety Alignment via Reinforcement Learning

As of 8 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 4 inbound Pith citation observations for arXiv:2507.08270.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.08270 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:29:36.837135Z

measured 34 of 34 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:14:57.067936Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T02:16:26.560954Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved16
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 180d6887-487a-4761-b359-a24be9ab44c2 · outbound

This paper cites Narasimhan, and Yuan Cao.

Agent Safety Alignment via Reinforcement Learning Narasimhan, and Yuan Cao

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:39.278568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:35.375189Z digest=sha256:8065972b68df4ffb85fea0082e497b9b508cd096394555ba84f922f0c565bd85

Observation 9b8cfb93-2cd6-414b-bca4-c1d04723b526 · outbound

This paper cites AutoGPT: An open-source autonomous agent framework.

Agent Safety Alignment via Reinforcement Learning AutoGPT: An open-source autonomous agent framework

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:39.040807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:35.487365Z digest=sha256:4a21b639539e3a43626b470e8a7af5812e1e262be822ad0377bc27afdd33fc18

Observation 8c747e8a-2736-4a40-863c-48cef5300cb1 · outbound

This paper cites BabyAGI: Experimental self-building autonomous agent.

Agent Safety Alignment via Reinforcement Learning BabyAGI: Experimental self-building autonomous agent

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:38.892749Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:35.542656Z digest=sha256:4144ba4e0f66fab6a8db89255af370046403dc5a6714f3bd73a3fc875a863365

Observation 30daa421-c6f6-4f62-871b-9fbce1301e31 · outbound

This paper cites AgentGPT: Configure and deploy autonomous ai agents.

Agent Safety Alignment via Reinforcement Learning AgentGPT: Configure and deploy autonomous ai agents

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:38.670365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:35.591544Z digest=sha256:ccd4178d8e2bbd3d9ad58a32b6fbe2d2675f0985bd4611d92a243bcc39142715

Observation 6d323b6e-662d-4c35-8a7b-b8dd7e2d3ac3 · outbound

This paper cites AI agents under threat: A survey of key security challenges and future pathways.

Agent Safety Alignment via Reinforcement Learning AI agents under threat: A survey of key security challenges and future pathways

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:38.493698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:35.624350Z digest=sha256:32b181f2daf732c84c8d2e8613408ff5082e026a336726dcc35d6ed37defdb3b

Observation e903261e-c05d-428b-a610-f02e4534c4f9 · outbound

This paper cites Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents.

Agent Safety Alignment via Reinforcement Learning Navigating the Risks: A Survey of Security, Privacy, and Ethics Threats in LLM-Based Agents

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:35.685754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:35.685754Z digest=sha256:f4a9aad26323893f0f9a405a7872bf2999a79f9614d713bc8f121d19a51d3ac4

Observation 277557af-1ef4-4596-8267-f1f6d14cfb36 · outbound

This paper cites Retool: Reinforcement learning for strategic tool use in llms, 2025.

Agent Safety Alignment via Reinforcement Learning Retool: Reinforcement learning for strategic tool use in llms, 2025

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:38.295437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:35.730506Z digest=sha256:17c214601ec561e2ef0cc3acfad9ab1444bde93b28e3f9ff1af09cb31015464e

Observation 9064441a-04bc-4858-bf9f-bd04b31bf99b · outbound

This paper cites SEM: Reinforcement Learning for Search-Efficient Large Language Models.

Agent Safety Alignment via Reinforcement Learning SEM: Reinforcement Learning for Search-Efficient Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:35.794563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:35.794563Z digest=sha256:e8326edaf29ca73a862ec3b621714f81826e10a240319ffaad78955f0f39223e

Observation 11389ede-4427-4e8b-9513-85c8adb70359 · outbound

This paper cites Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training.

Agent Safety Alignment via Reinforcement Learning Agent-R: Training Language Model Agents to Reflect via Iterative Self-Training

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:35.842058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:35.842058Z digest=sha256:2c108eceb0dc2d841da2e3165148426ac7c0045d8fa45a530f4c7b7e5c697590

Observation 9db288de-ae8a-40ce-922e-1f3dbc026ce4 · outbound

This paper cites Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen.

Agent Safety Alignment via Reinforcement Learning Pan, Wen Zhang, Huajun Chen, Fan Yang, Zenan Zhou, and Weipeng Chen

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:38.183883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:35.896986Z digest=sha256:aea16a567ae1359c5870b82a4e8c9b1796ebc30cfecf413d22722e303ec28a21

Observation 81cbd57d-026c-4c08-b680-a2d98e1d6acf · outbound

This paper cites Agent security bench (ASB): formalizing and benchmarking attacks and defenses in llm-based agents.

Agent Safety Alignment via Reinforcement Learning Agent security bench (ASB): formalizing and benchmarking attacks and defenses in llm-based agents

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:38.041081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:35.957481Z digest=sha256:1b70dffcd0df8cee8e9d62a0aff7cd0d192e58175e500d5bec5cdb921f4dc2d7

Observation 29aeb465-7be8-4e24-a417-a491faef2aff · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

Agent Safety Alignment via Reinforcement Learning Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.015210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.015210Z digest=sha256:aff77c46b1bf768f2c4c127491a539795cd0e1f7f1d16d206c8f1188bfe6e450

Observation f0ee323a-2ed7-4e0d-bd10-97e50f9bb5ae · outbound

This paper cites Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction.

Agent Safety Alignment via Reinforcement Learning Think Twice Before You Act: Enhancing Agent Behavioral Safety with Thought Correction

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.076016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.076016Z digest=sha256:d89c6aa7b0f3ffb7f5809675c73b32e94c000de817780da88a62e4a9c5b59528

Observation 810eec96-44f3-4628-8fde-c766f6a55667 · outbound

This paper cites AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models.

Agent Safety Alignment via Reinforcement Learning AgentAlign: Navigating Safety Alignment in the Shift from Informative to Agentic Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.127202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.127202Z digest=sha256:771e111f61d5b016da07f52e412d7803b3ee0047f4127f168b4ca005f00a06b2

Observation 80905211-b7a7-441a-a286-19e9f3ddb51a · outbound

This paper cites Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025.

Agent Safety Alignment via Reinforcement Learning Deepseek-r1: Incentivizing reasoning capability in llms via reinforcement learning, 2025

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.164795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.164795Z digest=sha256:2026412200d0a9c8fa3da9cde633a6f0d6ef7cdb30d881ded3d8b086b96d8f25

Observation 17a08604-0c36-4cad-8775-c863bc923e5b · outbound

This paper cites Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents.

Agent Safety Alignment via Reinforcement Learning Injecagent: Benchmarking indirect prompt injections in tool-integrated large language model agents

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:37.906726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:36.202956Z digest=sha256:9b0f98e84b534190f01c56448b3b1f32533e9d43ca7b7c45f19c9348ce44fad4

Observation e0617b99-43ab-43ee-ad94-10026b7b850b · outbound

This paper cites Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E.

Agent Safety Alignment via Reinforcement Learning Patil, Huanzhi Mao, Charlie Cheng-Jie Ji, Fanjia Yan, Vishnu Suresh, Ion Stoica, and Joseph E

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:37.752821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:36.248215Z digest=sha256:36c433728a5c7c6fe3b07db2408acaca919b8c01f3a5f2a6f77fd322bbacd311

Observation b2d897fa-1c92-4adf-9b0a-edc24e9241d6 · outbound

This paper cites Large Language Model Agent: A Survey on Methodology, Applications and Challenges.

Agent Safety Alignment via Reinforcement Learning Large Language Model Agent: A Survey on Methodology, Applications and Challenges

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.302719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.302719Z digest=sha256:5eb70a567fc93a30c368d54442cfe08463eafa193dcdfa59770aa3f8f8172c26

Observation 6e81ee76-f028-4cc3-bc71-108bc49d7958 · outbound

This paper cites An In-depth Survey of Large Language Model-based Artificial Intelligence Agents.

Agent Safety Alignment via Reinforcement Learning An In-depth Survey of Large Language Model-based Artificial Intelligence Agents

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.361217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.361217Z digest=sha256:b383d4ff51d60c895c2b81217ff948d0c82df162fe57d316925e0d4973753fea

Observation f10aa2d4-cfbb-496c-82c6-09d82e7f1ce8 · outbound

This paper cites Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L.

Agent Safety Alignment via Reinforcement Learning Sumers, Shunyu Yao, Karthik Narasimhan, and Thomas L

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:37.597508Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:36.410757Z digest=sha256:c245d434f6c940f87f9bb672d3f14ea11481b9199f47eb5f3c0483114c72b57d

Observation a273817c-b321-431e-bd36-6a183d9853e9 · outbound

This paper cites AutoAgent: A Fully-Automated and Zero-Code Frame- work for LLM Agents, 2025.

Agent Safety Alignment via Reinforcement Learning AutoAgent: A Fully-Automated and Zero-Code Frame- work for LLM Agents, 2025

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:37.477865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:36.458385Z digest=sha256:107fdd4894f7b82dc340e28158dbb34f99ae37a09ae78eab47da78047510c786

Observation dad2e0b3-a7d9-4aea-8958-819993d89801 · outbound

This paper cites ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning.

Agent Safety Alignment via Reinforcement Learning ReSearch: Learning to Reason with Search for LLMs via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.509257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.509257Z digest=sha256:f6f197a7c3428d8fef4250ae5322e5c43b154088377e64cfc2fb68546d70648d

Observation 17698f00-7f3f-49a3-a4d3-fc47cef63484 · outbound

This paper cites Deepresearcher: Scaling deep research via reinforcement learning in real-world environments, 2025.

Agent Safety Alignment via Reinforcement Learning Deepresearcher: Scaling deep research via reinforcement learning in real-world environments, 2025

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.552229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.552229Z digest=sha256:15f99ffb0416ee2b2a23cfa2bd0e1a6b7b04c1e029f557526273e84255488ca9

Observation c292e407-2abc-4880-9c67-6926d025440e · outbound

This paper cites Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol (MCP) Ecosystem.

Agent Safety Alignment via Reinforcement Learning Beyond the Protocol: Unveiling Attack Vectors in the Model Context Protocol (MCP) Ecosystem

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.599007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.599007Z digest=sha256:c937765351714d3398295379aa82e86485a7dcad10408e1de68309662eb7cfc2

Observation af481968-d1c7-4ba9-aca5-a42cdc8392d5 · outbound

This paper cites Progent: Securing AI Agents with Privilege Control.

Agent Safety Alignment via Reinforcement Learning Progent: Securing AI Agents with Privilege Control

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.649473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.649473Z digest=sha256:0bf6899bbf647bf86d8bf4fd37f3e54f4379ce892e0ff6858801885648e8d5da

Observation 3d8ff5a5-ea96-4e0b-91cb-22d95099f4d2 · outbound

This paper cites Fox in the Henhouse: Supply-Chain Backdoor Attacks Against Reinforcement Learning.

Agent Safety Alignment via Reinforcement Learning Fox in the Henhouse: Supply-Chain Backdoor Attacks Against Reinforcement Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.704014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.704014Z digest=sha256:fd44add5143d929f00dae7b898b3606a63e138cede2dc1c2be6a70ed2e4a6096

Observation f0adbf1b-2904-4aa7-954c-6cf7b23e3bc7 · outbound

This paper cites A practical memory injection attack against LLM agents.

Agent Safety Alignment via Reinforcement Learning A practical memory injection attack against LLM agents

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.746281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.746281Z digest=sha256:e03d267ce9bfcc90689c5cacc4aa0a4ab34be5e581752a5a6e5570e1648eeefa

Observation 17663713-ff13-49cf-816d-39dcd70da839 · outbound

This paper cites Agentpoison: Red- teaming LLM agents via poisoning memory or knowledge bases.

Agent Safety Alignment via Reinforcement Learning Agentpoison: Red- teaming LLM agents via poisoning memory or knowledge bases

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:29:37.380765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T18:29:36.793821Z digest=sha256:3a0ead7657629131ae6857b2e737f8f33287d6debe06cc6c38fdb872393201ab

Observation 474d8d38-3066-4322-b55d-e140835d4b2f · outbound

This paper cites Safeagentbench: A benchmark for safe task planning of embodied LLM agents.

Agent Safety Alignment via Reinforcement Learning Safeagentbench: A benchmark for safe task planning of embodied LLM agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T18:29:36.837135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:36.837135Z digest=sha256:66b41302c22f476ba530624b365b7eb429117e07103b3cd5c7467ffcb2188eec

Observation 86d6065a-488b-4180-8dbc-b3fb19586d46 · outbound

This paper cites an unresolved cited work.

Agent Safety Alignment via Reinforcement Learning Unresolved cited work

Reference 2023

Resolution
parse uncertain
no resolver link, observed 2026-08-06T18:29:35.441056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:29:35.441056Z digest=sha256:950bca40b51c63c5b72eadb951ccef29f04b84632f7118ec412af1c172673799

Pith citing papers

Observation 4e0b924d-c7a3-4880-af77-73a752de6b5f · inbound

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner cites this paper.

S3LoRA: Safe Spectral Sharpness-Guided Pruning in Adaptation of Agent Planner Agent Safety Alignment via Reinforcement Learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T18:12:35.170680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T18:12:35.170680Z digest=sha256:99d4ff667d5401eb734c3b4b3a31742a3f759787f370f2b623e6c343767b41bc

Observation df93920d-a25d-43be-90bb-e4a56a80e3e5 · inbound

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety cites this paper.

RUBAS: Rubric-Based Reinforcement Learning for Agent Safety Agent Safety Alignment via Reinforcement Learning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-02T02:16:26.562656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-28T11:07:47.814115Z digest=sha256:a046650dd07cc044c5ced17feb9c56c79595d38b7f2d0e774fff91362f838422

Observation 302533e2-3be7-4e66-9313-72f4e359b5d8 · inbound

SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing cites this paper.

SAFETY SENTRY: Context-Aware Human Intervention via EXECUTE-ASK-REFUSE Routing Agent Safety Alignment via Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-02T04:49:08.700306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T04:49:08.700306Z digest=sha256:2d21dacc7de8b37fa4377b1310335fab45f091f312e729dc4b50a55cdff7c489

Observation 1de972a0-65c9-483d-bc7a-6cef8dca0b61 · inbound

$S^3$: Improving Agent Safety through Multi-Stage Defense cites this paper.

$S^3$: Improving Agent Safety through Multi-Stage Defense Agent Safety Alignment via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:14:57.067936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:14:57.067936Z digest=sha256:46632c2b7d52b27e42d9dc19112271b9d9440621a0cc1e1a5f0edad61bd772fa