Pith. sign in

Paper Citation Record · LEDGER

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents

As of 7 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2605.16282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.16282 v1

Coverage vector

measured 71 of 71 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-21T01:42:55.693115Z

measured 72 of 72 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T00:45:58.094591Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

71 of 71 outbound references displayed

  • verified exact69
  • verified fuzzy0
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d7448029-fd2b-46b8-8f01-bbf7ad43d487 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.964497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:4cecac60ebf47baf1218f02136e038c9a0ca75a3c263881838822ddff785314e

Observation a63e0232-3732-49d0-9208-c68909232097 · outbound

This paper cites Arora, S.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Arora, S

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.968973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:a1ad9e79c867b4b57bc110234f6b4525edb58c3ef640f7c364e1e62d1cd18cc1

Observation 11199b3f-40b2-4b92-a653-44732c62ff5b · outbound

This paper cites Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.997040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:2a25fb804b1e98dd65dbba95035bbe9ad483d1e53cb0347711b45de93edb8acb

Observation 6457398a-7e11-4d66-841d-cf2780bdb754 · outbound

This paper cites RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.973932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:ee61ff3effe0cae862f4934e246d5ec2c25c79f845d098f04b5a4cccda287c29

Observation f7055d05-e992-49c0-854f-8125434520f9 · outbound

This paper cites Bordes, C.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Bordes, C

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.991937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:03ed94f73729aaca1840c571c5abf2afc1a3553cfdb6c608b2580b9d036983be

Observation 4b9fdf3f-7778-40d7-9dd2-9f12810ad2a9 · outbound

This paper cites AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:57.001946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:777806a4302792b9f9cf7f4fa961ec1e7456aeba36c7042fca5a773636b01749

Observation 97325393-fe10-40ab-8464-4e60c294076c · outbound

This paper cites Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.931932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:4b38e3100094b661d6e08f2d5621b2ae48f324048020d4bd4d65d4372df911fe

Observation 2cbda156-ec48-46de-92e9-fed450dadf5a · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents OR-Bench: An Over-Refusal Benchmark for Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.924885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:8293265d226e2956fa5072a2ab22600fd4596e22fca05bb21313852d11652803

Observation 317b4b9f-2ad6-44b5-87e3-6709fa52abb7 · outbound

This paper cites AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.937545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:bb2818448e35c4ef7a35afc985cbf14d9c1ecf1717b73eb280a41c975e22a67a

Observation e25cc844-6f5a-4aa3-a9b8-e4b14881adfc · outbound

This paper cites Yann Dubois, Balázs Galambosi, Percy Liang, and Tat- sunori B Hashimoto.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Yann Dubois, Balázs Galambosi, Percy Liang, and Tat- sunori B Hashimoto

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.942984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:fc7455d849c5e809a0ddd93425b65437b183db153b93677ab0631bd58dc95963

Observation f3d770a8-7c0d-42e7-bb3a-150ffd0b512d · outbound

This paper cites Agentleak: A full-stack benchmark for privacy leakage in multi-agent llm systems.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Agentleak: A full-stack benchmark for privacy leakage in multi-agent llm systems

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.913592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:16fe7358709e2d5052d9e18754dad9a57b804a4f133003ae208d9506c88c0346

Observation 7ffffaaa-f023-4b25-9ec1-55900731891a · outbound

This paper cites WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.902195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:319ee6b631d5edd8297792426de01781fe674c7a4d814ac913f61977500cf448

Observation 04461fc8-bafd-4be1-9322-7842dff66de3 · outbound

This paper cites Backdooragent: A unified framework for backdoor attacks on llm-based agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Backdooragent: A unified framework for backdoor attacks on llm-based agents

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.880888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:987b42b2d4572a10ab51f3f041e447a89bcf8ea432a2de5697fd5fc5313f1c25

Observation 85b0d4e7-a8a2-42ee-86d0-03b4d12aa8f5 · outbound

This paper cites RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.870990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:b739707c66b9be20efa79dda7b67d96482e9f573edf38fae0edf4ce5f7e09962

Observation 853c3ada-9d65-4089-9c8a-cfeaece0b7a2 · outbound

This paper cites Alignment faking in large language models.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Alignment faking in large language models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.885882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:5553c2a278e38756813c16b668dddc05698bcff98c5ba6ca03974395a5d1bfed

Observation b32ee26c-2bd6-4fec-9609-431acca58607 · outbound

This paper cites Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.875654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:baab52d44a5661b85c6cc409e87b6d01e4530d84f747431053516ca629262a2b

Observation 4e9279dc-b024-4673-880b-171ed4c94d20 · outbound

This paper cites Hadeliya, M.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Hadeliya, M

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.907448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:7ed261153a2ff4b7d32d6f34e6b07d009b49969966f1a84921ad190101ce1a01

Observation 509f0dc6-de5c-46f4-b529-ff370c2b0f9d · outbound

This paper cites Multi-Agent Risks from Advanced AI.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Multi-Agent Risks from Advanced AI

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.982090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:94be14261d9ef435a3821b1c4815eec235338c5ec87de0d8296fef507db2af4b

Observation 56f9752d-e8a3-4b38-84f7-a6f74d7b8cd7 · outbound

This paper cites Hopman, J.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Hopman, J

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.833946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:1a176ff230e714b9fe19719897cc4344fae92ca5590d39137639fa231591f1ba

Observation a7fbe2f2-f1cf-40bb-ac9d-29c6534da786 · outbound

This paper cites TrustAgent: Towards Safe and Trustworthy LLM-based Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents TrustAgent: Towards Safe and Trustworthy LLM-based Agents

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.839185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:77f8b3ab4ee38b4eb4ca8cf4d7944a7bfb34af2f4fba96bac0840cee075d88a5

Observation 86227a48-bf0a-49a8-828e-3d27de4fa793 · outbound

This paper cites Risks from Learned Optimization in Advanced Machine Learning Systems.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Risks from Learned Optimization in Advanced Machine Learning Systems

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.828334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:33c60af9c5a69ebc80c1124d9e55a301deb5e920f5286b6d93901524a3cc549e

Observation 594844c5-4a6c-4666-baa5-72e0b5f1fc29 · outbound

This paper cites Jiang, Y.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Jiang, Y

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.844416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:be4948c8320d3fc5366f5ea581bfbe76d4d614d1e4ef2beec2f0aed0f28d98cd

Observation f6a8b501-c52d-49cb-a0c3-d818f5b7d17f · outbound

This paper cites Juneja, J.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Juneja, J

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.818083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:0d2b0a83bc783e84abb4bb444814a269261f2e5d4b30a98f67190a44954c0ab0

Observation 323dda98-9909-4554-bd9a-ebe310de255c · outbound

This paper cites Kavathekar, H.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Kavathekar, H

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.801027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:0ec4ffbd4227002cbe5d86be1118900f53b6a998772ec7f0b163a88ec2a8bf73

Observation c6157c38-f118-4488-965b-126a0e324d1a · outbound

This paper cites SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.805773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:7168d914acebc27666d0cbeea95a579a130822f6e647b7ee1ebe8493dff9591b

Observation 68b27e3c-0afc-42c0-997f-a0fff24e7e37 · outbound

This paper cites Bradley Knox, and Kimin Lee.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Bradley Knox, and Kimin Lee

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.812705Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:dc1d05eb3cff120dfcffbd57582ba0a6667214d8a1766686c11c8c0c94bce67d

Observation fccabdc4-e26a-4a40-8ad7-a2d8392829cf · outbound

This paper cites ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-06-05T02:16:16.756993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:f27b226bd7e3ed827ebe432e1cb8cfb833cf073d8baabb6cdfc9986251fd9764

Observation 7b1a9fd3-3093-4323-8df2-f34968cf18c4 · outbound

This paper cites A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.891144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:a899f97c8f7e2038d094bf1ea2eb0b212e7f6f83269354e6b50d84800e414d82

Observation cff8d4f4-6bba-42b0-a4bf-9a0969f32cca · outbound

This paper cites SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.639271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:146fa4d233b72594598c741b690c81403cf4771f8c74957b26b10a7bd88f4007

Observation d963c482-fd03-4cc5-8179-dcf80d320a36 · outbound

This paper cites Agentsafe: Benchmarking the safety of embodied agents on hazardous instructions.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Agentsafe: Benchmarking the safety of embodied agents on hazardous instructions

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.644461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:00748cf8f4da5dec3c2675d0b0e5cda0ee14ede386ee16d529ca47f6f2cc312e

Observation 393ab57d-bca6-42ba-8cc7-9a0f32feb5de · outbound

This paper cites Is- bench: Evaluating interactive safety of vlm-driven embodied agents in daily household tasks.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Is- bench: Evaluating interactive safety of vlm-driven embodied agents in daily household tasks

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:57.007268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:e97d62fc024ecdc3624e9f8931e4bdeff8fc200d7d4ab85264473fc0d68b8df2

Observation 0ae9007d-334b-4abd-b96b-75870300796d · outbound

This paper cites Agentauditor: Human-level safety and security evaluation for llm agents.arXiv preprint arXiv:2506.00641.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Agentauditor: Human-level safety and security evaluation for llm agents.arXiv preprint arXiv:2506.00641

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:57.012454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:983db534d52e246872a969f55f4ef95ca7ba581e60a4b1613488c978e9a494e9

Observation c1eafb32-4e8a-4616-bc67-c910bcde6c26 · outbound

This paper cites Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.823188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:43a256136f57a922dd0eba28c0655698f46900019a7f68ba7a62fd723d877bec

Observation e7318f8f-69f2-489e-a78a-6ef3d401c1fe · outbound

This paper cites Natural Emergent Misalignment from Reward Hacking in Production RL.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Natural Emergent Misalignment from Reward Hacking in Production RL

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.655093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:69fec2e03307ab41a3f1a3456fae5289c1fd7fb6f4716a89d0d3b3c00ed8e307

Observation 102ae941-79ac-43c2-8907-e5b90e58b0f3 · outbound

This paper cites McGregor, V.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents McGregor, V

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.860696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:1d15017400ea6f8a9584895fb4dda9351158c988da37d237aaeb62d6ad739f04

Observation bc03938c-acf7-4e92-8a54-1d47d8a45a91 · outbound

This paper cites Frontier Models are Capable of In-context Scheming.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Frontier Models are Capable of In-context Scheming

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.790333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:dacfb25b51229fc28f69c2da0588277691343194354ce3c2def363ac7a009cb2

Observation f0295ea1-fa65-467c-be28-0d8d34d5f688 · outbound

This paper cites an unresolved cited work.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-21T01:43:57.213984Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:c96da4de75cf50f13a79c1450244c8a79bf09fca15cd301c34065cb76b2ae69f

Observation b79c7602-7f7c-4d27-a357-45e108c435fa · outbound

This paper cites Evaluation and Benchmarking of LLM Agents: A Survey.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Evaluation and Benchmarking of LLM Agents: A Survey

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.680558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:b8efa42150cc28f61630dacfed0cfa4cba0c39a12ed66e7b5459bb07c25bd674

Observation fc1e67d8-bf15-4319-a5e5-da783eb72ca5 · outbound

This paper cites AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-06-23T04:13:38.836403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:2628148f318ee71ee7bcbcac4e63cf3de438211500a57762f417b9af7b511104

Observation e94a4214-ee0a-462d-b4cb-42256b2cd592 · outbound

This paper cites Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-28T02:04:14.756178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:2a28c963c951a3e9b84bf6bd14ce9aec95962d5615364808b30f08fe8939c77a

Observation 869c3a57-64a1-4ff9-ae17-3ec45f6b8da3 · outbound

This paper cites N \"o ther, A.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents N \"o ther, A

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:57.035744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:4ff591e8d5e50705123f973788a61c0c60ba4edd697f40a5523588a93ac26b72

Observation e8a3e7e2-a852-48ae-8c54-3e7a98dff719 · outbound

This paper cites Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.709741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:1b8d88113d050ed60629c1336870c5e112170c5a2b96f12e00aa3b0b6b0ba337

Observation 5a9e2519-e0b9-4629-aa8f-4bf3ca85ff2b · outbound

This paper cites When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.736724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:b25d9355b83ba729e396f9c562d1969e8ade0b55366a9b36fb6b7939e401e5b4

Observation 905caca1-94da-40e4-9adc-a3a18b023ce9 · outbound

This paper cites Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.649657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:769a1da19b505b2fad1005afa511443cdb8fa8f38cf770476447f5b512fc2e09

Observation 8fc44ee4-99b8-476f-860e-7e8645fc0122 · outbound

This paper cites BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.704660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:4991f498cb3dd6e35879b327ab21a5cd5b53c233501fb56d5a5752650bcd876f

Observation aa9ecd7e-3111-4372-8f79-b5d03991a1df · outbound

This paper cites Identifying the Risks of LM Agents with an LM-Emulated Sandbox.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Identifying the Risks of LM Agents with an LM-Emulated Sandbox

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.692475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:8ab87b760b7ce7537247ea32516cc7dfa76d3e1cc1f42d8646ee3b88a429feee

Observation c36f65e9-2496-45e3-bd1c-e9d03e27654a · outbound

This paper cites Schlatter, B.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Schlatter, B

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.700164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:60e9c855b03e62951b6e6d3b8a59726ba8fcb70485d9e8556b5ce221bb41322c

Observation 919ac7ed-ea23-413a-b3c7-9eb712fe5cfb · outbound

This paper cites PropensityBench: Evaluating propensity under pressure.arXiv preprint arXiv:2511.20703.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents PropensityBench: Evaluating propensity under pressure.arXiv preprint arXiv:2511.20703

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.855920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:57f4a105984613c9eec41b520b8372b1b38a01314cc1a1c6342e057246e3de43

Observation 6c0c96f1-3b89-4fd2-837e-27a291dffcfe · outbound

This paper cites an unresolved cited work.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Unresolved cited work

Reference 49

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.762811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:40ac2c5fa8c421962d8fab1b9f100447c6f8718116623bf76788093c9aeb86bf

Observation bef34a4a-98b5-41f4-847a-44179fde9bcb · outbound

This paper cites Vijayvargiya, A.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Vijayvargiya, A

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.953464Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:276a58950fd3a577eb97272fcd23db1094fad4bd2cb1ae5096b4d810e497c1a7

Observation 061f3076-96ea-4941-bda7-b4a3b59a52aa · outbound

This paper cites AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:57.021306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:f4b8b90675ef0c1e6f428efe3cdb0d3a38e4874327095f95c56112ed0a109b72

Observation 9a3344e2-5e67-4e03-bc5a-0669f6456152 · outbound

This paper cites A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.850450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:bd8bb712f238e28e728b0f71caf073c16687e0fadbd560d1b4b5623317675cd0

Observation 9465a880-24f8-4a52-90af-93158c92b1b1 · outbound

This paper cites A Survey on Large Language Model based Autonomous Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents A Survey on Large Language Model based Autonomous Agents

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.717642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:75b1c8a0d1b4f8d1b572c33f7cda2297ee405402283bc4533fc92b16567cd8e4

Observation c24b1e4b-2cb1-4fa4-a55c-6cb45de64082 · outbound

This paper cites The Rise and Potential of Large Language Model Based Agents: A Survey.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents The Rise and Potential of Large Language Model Based Agents: A Survey

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.725147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:76a1b7124ecf24702eab395fb41985dca8b50fcbcb6c04444c8902eff040bb02

Observation e1ef2c08-0e70-4678-b327-3718a55b296e · outbound

This paper cites SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.686578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:0b1720fa917842d7d781ef87140c96395eccdbe6aea09ca3bbf781bb83153169

Observation 3dc32d50-682f-4799-860e-6f84d6ceb9db · outbound

This paper cites GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:57:50.455179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:8524f1f4f40b874a1a876a78459f1a26c94e24087aacfc2bdd0c401f2263c7ce

Observation 6792b4dd-5bfc-41f9-95a2-70c74ecb788f · outbound

This paper cites an unresolved cited work.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Unresolved cited work

Reference 57

Resolution
unresolved
raw_fallback, observed 2026-05-21T01:43:57.210432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:5642da22ff5bc6578b146a892ab0398d5581bc865696241b08253f7a477946ab

Observation 907bf584-306e-46bd-b3cc-2a95e6a32d88 · outbound

This paper cites Survey on Evaluation of LLM-based Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Survey on Evaluation of LLM-based Agents

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.713747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:f88b078f1724eb9fe84b4941a3ecefd809f0379608af0f6d05b7830bf1d0aa5b

Observation 87a7a443-e182-4e75-b277-cc4a36b9f1e9 · outbound

This paper cites SafeAgentBench: A benchmark for safe task planning of embodied LLM agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents SafeAgentBench: A benchmark for safe task planning of embodied LLM agents

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.742827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:bb03bc9a0d41161633b1f3936629cf3a89000fb1b009180489cf758e99ee9feb

Observation 8b25fa8c-80dd-4d2a-96df-b8eab9b2499c · outbound

This paper cites How Should AI Safety Benchmarks Benchmark Safety?.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents How Should AI Safety Benchmarks Benchmark Safety?

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-08-06T02:07:04.187677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:6ad3a8b4f75b9cd53a7aacd81995be5fde27122f84de066b7a09a2a7916670a8

Observation 6afe3eec-e3fa-4324-b7c2-178e295152dc · outbound

This paper cites A Survey on Trustworthy LLM Agents: Threats and Countermeasures.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents A Survey on Trustworthy LLM Agents: Threats and Countermeasures

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.947706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:b71db39627f16a06d59221e7a2065365750ffbec602e2426a5788523bc02aca1

Observation 967c5c0a-2f0d-4fd8-b5e8-bcd54bff5ea2 · outbound

This paper cites R-Judge: Benchmarking Safety Risk Awareness for LLM Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents R-Judge: Benchmarking Safety Risk Awareness for LLM Agents

Reference 62

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.667888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:368101466242bae5bf2af14ceed066e62a943d47787f68e0d7923fcf2c8fb563

Observation 273019c1-ba34-48ec-a485-1e7b818939f6 · outbound

This paper cites Nothing humbles you like telling your OpenClaw ``confirm before acting'' and watching it speedrun deleting your inbox.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Nothing humbles you like telling your OpenClaw ``confirm before acting'' and watching it speedrun deleting your inbox

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.919539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:d3531305031b27311b91cf938923212961cc2de9d778a342cf116f6d6c6c11ac

Observation c4f743fa-3509-46c5-afde-7f86ab5ef670 · outbound

This paper cites InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.987491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:ed2c14425aaeed71207eace483158ad6188fecf99be1d6c019ccf0c41e867296

Observation d6bdd1a1-6a24-4dad-b7be-123f13f5c09c · outbound

This paper cites Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:57.025992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:303e1719b2c3ad57bdc26fb1142d18c2f82bb79e84016e775f13b2e6fbe47501

Observation 4f527707-ce98-47f7-9f26-9d30dff09837 · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.795455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:162acf075ea49314242b14d20660c28d3a829b304bcadd65ab73136c66dc93bf

Observation 57f6ad12-bcd2-428d-a55f-e1bb3dde946e · outbound

This paper cites PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:57.030529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:b53226cc4286ac52bddafab9fc807c249f74ba62bc9d07ca902403c955dcd972

Observation e5c1383d-de45-46e6-9327-ea19121a83e1 · outbound

This paper cites Safepro: Evaluating the safety of professional-level ai agents.arXiv preprint arXiv:2601.06663, 2026.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Safepro: Evaluating the safety of professional-level ai agents.arXiv preprint arXiv:2601.06663, 2026

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:57.016964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:486e0fd67a9784fa4fab19f2473b8b770ac1247aea31e1e31cc70ea4d1cb1ae2

Observation 27d9d08c-85b0-49c3-8ef9-23467c42abe9 · outbound

This paper cites Mcp-safetybench: A benchmark for safety evaluation of large language models with real-world mcp servers.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Mcp-safetybench: A benchmark for safety evaluation of large language models with real-world mcp servers

Reference 69

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.958592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:55e8392e122af9ee8546e993173ba25c51bceed7f8185d9248888496f21fc812

Observation dfac3783-4610-42ff-9440-cb100574a342 · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-21T01:43:56.662569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:65e40f7a61cb5f5c74b8c6200eec6287929bb44f656ebc021b8d1ab0c7bae127

Observation 287ada0c-ae60-400a-a549-643165fec13f · outbound

This paper cites PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models

Reference 71

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.674346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:39f9a55d9d89fa64aa32391054ef2262290abc7ca3f51530d56f178944697981

Pith citing papers

Observation 257adf34-3152-4825-9e69-0a95f0fbc0d7 · inbound

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks cites this paper.

Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T00:45:58.094591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:45:58.094591Z digest=sha256:fdd2a8378830464393614da721d4409c5e4d93b3b701b8daf37e710b9d4ebed1