Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T01:42:55.693115Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 71 of 71 outbound references and 1 inbound Pith citation observation for arXiv:2605.16282.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-21T01:42:55.693115Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-03T00:45:58.094591Z
A source-named dated measurement, never combined with another source.
Source: cited_works
71 of 71 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d7448029-fd2b-46b8-8f01-bbf7ad43d487 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a63e0232-3732-49d0-9208-c68909232097 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Arora, S
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 11199b3f-40b2-4b92-a653-44732c62ff5b · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Training a Helpful and Harmless Assistant with Reinforcement Learning from Human Feedback
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6457398a-7e11-4d66-841d-cf2780bdb754 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents RepliBench: Evaluating the Autonomous Replication Capabilities of Language Model Agents
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f7055d05-e992-49c0-854f-8125434520f9 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Bordes, C
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4b9fdf3f-7778-40d7-9dd2-9f12810ad2a9 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents AgentPoison: Red-teaming LLM Agents via Poisoning Memory or Knowledge Bases
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 97325393-fe10-40ab-8464-4e60c294076c · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Agentic AI Security: Threats, Defenses, Evaluation, and Open Challenges
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 2cbda156-ec48-46de-92e9-fed450dadf5a · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents OR-Bench: An Over-Refusal Benchmark for Large Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 317b4b9f-2ad6-44b5-87e3-6709fa52abb7 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents AgentDojo: A Dynamic Environment to Evaluate Prompt Injection Attacks and Defenses for LLM Agents
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e25cc844-6f5a-4aa3-a9b8-e4b14881adfc · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Yann Dubois, Balázs Galambosi, Percy Liang, and Tat- sunori B Hashimoto
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f3d770a8-7c0d-42e7-bb3a-150ffd0b512d · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Agentleak: A full-stack benchmark for privacy leakage in multi-agent llm systems
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7ffffaaa-f023-4b25-9ec1-55900731891a · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents WASP: Benchmarking Web Agent Security Against Prompt Injection Attacks
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 04461fc8-bafd-4be1-9322-7842dff66de3 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Backdooragent: A unified framework for backdoor attacks on llm-based agents
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 85b0d4e7-a8a2-42ee-86d0-03b4d12aa8f5 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents RAS-Eval: A Comprehensive Benchmark for Security Evaluation of LLM Agents in Real-World Environments
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 853c3ada-9d65-4089-9c8a-cfeaece0b7a2 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Alignment faking in large language models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b32ee26c-2bd6-4fec-9609-431acca58607 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Not what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4e9279dc-b024-4673-880b-171ed4c94d20 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Hadeliya, M
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 509f0dc6-de5c-46f4-b529-ff370c2b0f9d · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Multi-Agent Risks from Advanced AI
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 56f9752d-e8a3-4b38-84f7-a6f74d7b8cd7 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Hopman, J
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation a7fbe2f2-f1cf-40bb-ac9d-29c6534da786 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents TrustAgent: Towards Safe and Trustworthy LLM-based Agents
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 86227a48-bf0a-49a8-828e-3d27de4fa793 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Risks from Learned Optimization in Advanced Machine Learning Systems
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 594844c5-4a6c-4666-baa5-72e0b5f1fc29 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Jiang, Y
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f6a8b501-c52d-49cb-a0c3-d818f5b7d17f · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Juneja, J
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 323dda98-9909-4554-bd9a-ebe310de255c · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Kavathekar, H
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c6157c38-f118-4488-965b-126a0e324d1a · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents SHADE-Arena: Evaluating Sabotage and Monitoring in LLM Agents
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 68b27e3c-0afc-42c0-997f-a0fff24e7e37 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Bradley Knox, and Kimin Lee
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fccabdc4-e26a-4a40-8ad7-a2d8392829cf · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 7b1a9fd3-3093-4323-8df2-f34968cf18c4 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents A Benchmark for Evaluating Outcome-Driven Constraint Violations in Autonomous AI Agents
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation cff8d4f4-6bba-42b0-a4bf-9a0969f32cca · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents SafeRAG: Benchmarking Security in Retrieval-Augmented Generation of Large Language Model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d963c482-fd03-4cc5-8179-dcf80d320a36 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Agentsafe: Benchmarking the safety of embodied agents on hazardous instructions
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 393ab57d-bca6-42ba-8cc7-9a0f32feb5de · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Is- bench: Evaluating interactive safety of vlm-driven embodied agents in daily household tasks
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 0ae9007d-334b-4abd-b96b-75870300796d · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Agentauditor: Human-level safety and security evaluation for llm agents.arXiv preprint arXiv:2506.00641
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c1eafb32-4e8a-4616-bc67-c910bcde6c26 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Safety at Scale: A Comprehensive Survey of Large Model and Agent Safety
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e7318f8f-69f2-489e-a78a-6ef3d401c1fe · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Natural Emergent Misalignment from Reward Hacking in Production RL
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 102ae941-79ac-43c2-8907-e5b90e58b0f3 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents McGregor, V
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bc03938c-acf7-4e92-8a54-1d47d8a45a91 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Frontier Models are Capable of In-context Scheming
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation f0295ea1-fa65-467c-be28-0d8d34d5f688 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Unresolved cited work
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation b79c7602-7f7c-4d27-a357-45e108c435fa · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Evaluation and Benchmarking of LLM Agents: A Survey
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation fc1e67d8-bf15-4319-a5e5-da783eb72ca5 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e94a4214-ee0a-462d-b4cb-42256b2cd592 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Colosseum: Auditing Collusion in Cooperative Multi-Agent Systems
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 869c3a57-64a1-4ff9-ae17-3ec45f6b8da3 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents N \"o ther, A
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e8a3e7e2-a852-48ae-8c54-3e7a98dff719 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Do the Rewards Justify the Means? Measuring Trade-Offs Between Rewards and Ethical Behavior in the MACHIAVELLI Benchmark
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 5a9e2519-e0b9-4629-aa8f-4bf3ca85ff2b · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents When AI Agents Collude Online: Financial Fraud Risks by Collaborative LLM Agents on Social Platforms
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 905caca1-94da-40e4-9adc-a3a18b023ce9 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8fc44ee4-99b8-476f-860e-7e8645fc0122 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents BetterBench: Assessing AI Benchmarks, Uncovering Issues, and Establishing Best Practices
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation aa9ecd7e-3111-4372-8f79-b5d03991a1df · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Identifying the Risks of LM Agents with an LM-Emulated Sandbox
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c36f65e9-2496-45e3-bd1c-e9d03e27654a · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Schlatter, B
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 919ac7ed-ea23-413a-b3c7-9eb712fe5cfb · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents PropensityBench: Evaluating propensity under pressure.arXiv preprint arXiv:2511.20703
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6c0c96f1-3b89-4fd2-837e-27a291dffcfe · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation bef34a4a-98b5-41f4-847a-44179fde9bcb · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Vijayvargiya, A
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 061f3076-96ea-4941-bda7-b4a3b59a52aa · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents AgentSpec: Customizable Runtime Enforcement for Safe and Reliable LLM Agents
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9a3344e2-5e67-4e03-bc5a-0669f6456152 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents A Comprehensive Survey in LLM(-Agent) Full Stack Safety: Data, Training and Deployment
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 9465a880-24f8-4a52-90af-93158c92b1b1 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents A Survey on Large Language Model based Autonomous Agents
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c24b1e4b-2cb1-4fa4-a55c-6cb45de64082 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents The Rise and Potential of Large Language Model Based Agents: A Survey
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e1ef2c08-0e70-4678-b327-3718a55b296e · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents SafeToolBench: Pioneering a Prospective Benchmark to Evaluating Tool Utilization Safety in LLMs
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 3dc32d50-682f-4799-860e-6f84d6ceb9db · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents GuardAgent: Safeguard LLM Agents by a Guard Agent via Knowledge-Enabled Reasoning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6792b4dd-5bfc-41f9-95a2-70c74ecb788f · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Unresolved cited work
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 907bf584-306e-46bd-b3cc-2a95e6a32d88 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Survey on Evaluation of LLM-based Agents
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 87a7a443-e182-4e75-b277-cc4a36b9f1e9 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents SafeAgentBench: A benchmark for safe task planning of embodied LLM agents
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 8b25fa8c-80dd-4d2a-96df-b8eab9b2499c · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents How Should AI Safety Benchmarks Benchmark Safety?
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 6afe3eec-e3fa-4324-b7c2-178e295152dc · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents A Survey on Trustworthy LLM Agents: Threats and Countermeasures
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 967c5c0a-2f0d-4fd8-b5e8-bcd54bff5ea2 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents R-Judge: Benchmarking Safety Risk Awareness for LLM Agents
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 273019c1-ba34-48ec-a485-1e7b818939f6 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Nothing humbles you like telling your OpenClaw ``confirm before acting'' and watching it speedrun deleting your inbox
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation c4f743fa-3509-46c5-afde-7f86ab5ef670 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation d6bdd1a1-6a24-4dad-b7be-123f13f5c09c · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Agent Security Bench (ASB): Formalizing and Benchmarking Attacks and Defenses in LLM-based Agents
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 4f527707-ce98-47f7-9f26-9d30dff09837 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Agent-SafetyBench: Evaluating the Safety of LLM Agents
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 57f6ad12-bcd2-428d-a55f-e1bb3dde946e · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents PsySafe: A Comprehensive Framework for Psychological-based Attack, Defense, and Evaluation of Multi-agent System Safety
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation e5c1383d-de45-46e6-9327-ea19121a83e1 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Safepro: Evaluating the safety of professional-level ai agents.arXiv preprint arXiv:2601.06663, 2026
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 27d9d08c-85b0-49c3-8ef9-23467c42abe9 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Mcp-safetybench: A benchmark for safety evaluation of large language models with real-world mcp servers
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation dfac3783-4610-42ff-9440-cb100574a342 · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Universal and Transferable Adversarial Attacks on Aligned Language Models
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 287ada0c-ae60-400a-a549-643165fec13f · outbound
Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents PoisonedRAG: Knowledge Corruption Attacks to Retrieval-Augmented Generation of Large Language Models
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.
Observation 257adf34-3152-4825-9e69-0a95f0fbc0d7 · inbound
Safety, or Just Capability? A Validity Audit of Agent-Safety Benchmarks Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.