Pith. sign in

Paper Citation Record · LEDGER

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components

As of 9 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2506.02357.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.02357 v2

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:28:22.498212Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact1
  • verified fuzzy4
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c731d977-eaff-4ee6-a600-57b1d766d91f · outbound

This paper cites and Vichy, L.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components and Vichy, L

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:23.601617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:28:20.088809Z digest=sha256:0f0f3cc69b0b5f5c929817a3bff1b19f53067a721f641421d090058088f05ca5

Observation e01a18ed-85a8-4d31-90b3-c9742917d96d · outbound

This paper cites AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components AI Governance to Avoid Extinction: The Strategic Landscape and Actionable Research Questions

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-08-07T11:28:22.875538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:28:20.217361Z digest=sha256:168d65f7137ffc934aca231fdf77a65b70203d1d3fb443ed3d85e48b2f0a9eeb

Observation 5c6c9aba-0b6c-4ebc-aeb0-3687fcedd11e · outbound

This paper cites S., and Terry, J.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components S., and Terry, J

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:23.418350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:28:20.369273Z digest=sha256:f55723d93b8314b47dac03ae0ce07f4fe8957c1bbbea9711486739d5deeee756

Observation bfae4b35-f4a7-4ab0-b4be-56ad08290888 · outbound

This paper cites and Jaffer, I.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components and Jaffer, I

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:23.213872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:28:20.485938Z digest=sha256:721d36ef92184ebe2d955bf794eae6ef5afa5b3766359b61340100354619d75d

Observation d92a8644-8493-4c6b-99b3-ccb8f4a3cd73 · outbound

This paper cites an unresolved cited work.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:20.578837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:20.578837Z digest=sha256:e425e04ffda8a84e336ba5210bbe1c899d3614f5488b57a4346eeabc4ed17e3a

Observation d99caefb-5ec6-40dc-bab6-a222343171bf · outbound

This paper cites FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components FollowBench: A Multi-level Fine-grained Constraints Following Benchmark for Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:20.709513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:20.709513Z digest=sha256:2b57b9132cc5f765a3f3a437cf79bb7a3454c88558e8c618525ce7eae97b8cfa

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:20.868221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:20.868221Z digest=sha256:28a2dbdab6ea1385d689f63128dd21bcc94037730af2cdb8fba2e89205bbe09d

Observation 1150a597-0e19-4efa-955e-c9ea5c03a4b7 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components AgentBench: Evaluating LLMs as Agents

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:20.990736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:20.990736Z digest=sha256:54f2a7e30977cd2a89daacdcf940f769a465b8d77e74fae3cc432cbdfbe0f436

Observation 8e0d26c6-80e1-46c1-9551-0cc53537e299 · outbound

This paper cites The Alignment Problem from a Deep Learning Perspective.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components The Alignment Problem from a Deep Learning Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.158123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.158123Z digest=sha256:57c52155a59dc9481aecd06730f07a3d4d7e664738d8a18e424d72b1937ba240

Observation 2cb1cb4b-00b5-4c05-91ac-66e1b4859e42 · outbound

This paper cites S., O'Brien, J., Cai, C.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components S., O'Brien, J., Cai, C

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.294656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.294656Z digest=sha256:a4a08f2f336723713b50bc4af0ce2a68ce9404b37ff6adf86bb09d230a1aabbf

Observation bba966d5-1688-442e-be81-61fac0c0b936 · outbound

This paper cites Open Problems in Technical AI Governance.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Open Problems in Technical AI Governance

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.484300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.484300Z digest=sha256:696a5bcd84bc419d566c44e881293526949e856f0e2340d708f11aa70b193082

Observation 24a86f83-83d2-4712-9042-e6e893aa06df · outbound

This paper cites Model evaluation for extreme risks.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Model evaluation for extreme risks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.649808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.649808Z digest=sha256:e4d760667ddbbf661c52daf4ea77083eb5df6da5ce2165be63b5c13cd139e702

Observation f64e927d-60da-4507-910d-227805fa25bd · outbound

This paper cites The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components The Instruction Hierarchy: Training LLMs to Prioritize Privileged Instructions

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.793229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.793229Z digest=sha256:eef7d6f6dd4d375ed8f04c1c83ef62b24176df08756b2e326d8fec2959a6d447

Observation b9384660-b1cd-4cb1-aed0-a026cb12bb6d · outbound

This paper cites Voyager: An Open-Ended Embodied Agent with Large Language Models.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Voyager: An Open-Ended Embodied Agent with Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:21.929672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:21.929672Z digest=sha256:9aaf43b3997d4b64a21951c250d8fb37ff8ccd175320c005164e61340db1c9cb

Observation 54b5fbf1-be2d-46ee-8c4e-19235f995d69 · outbound

This paper cites Benchmarking Complex Instruction - Following with Multiple Constraints Composition.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Benchmarking Complex Instruction - Following with Multiple Constraints Composition

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:28:23.013534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-07T11:28:22.121414Z digest=sha256:598e356f3b2699d4cf6838dcc323b940cbb4016b275d3a9961864dec60ca7212

Observation 41bb9b5d-f5f4-49f8-af4f-d7e22349bff6 · outbound

This paper cites S., Shah, A., and Tellex, S.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components S., Shah, A., and Tellex, S

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:22.287533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:22.287533Z digest=sha256:37a6af7dc2718f733e5c17a0c20ffd418da763b061b76d737eb437ccda2fcd5a

Observation 0cc63e0e-1775-43ac-bf38-d6870c090c41 · outbound

This paper cites InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components InjecAgent: Benchmarking Indirect Prompt Injections in Tool-Integrated Large Language Model Agents

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:22.414764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:22.414764Z digest=sha256:ab207fc54f42a7b686f8e1bd2f3d0a231b4b82aca482006cce370e8dd8c79eb3

Observation 6b8009c5-20f0-4790-96fa-03d4034b8ec4 · outbound

This paper cites Agent-SafetyBench: Evaluating the Safety of LLM Agents.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components Agent-SafetyBench: Evaluating the Safety of LLM Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:22.489474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:22.489474Z digest=sha256:d524168c308b98b18c11dcd808c96b6e3b22794947405ee306379be7bd141ca8

Observation bc09d1d5-80d4-422a-a37a-b0276b53b150 · outbound

This paper cites write newline.

Evaluating LLM Agent Adherence to Hierarchical Safety Principles: A Lightweight Benchmark for Probing Foundational Controllability Components write newline

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:28:22.498212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:28:22.498212Z digest=sha256:b244309622a7cc11a13092a13aa852443dfbc42f87f8aada3f04d9edfe2e0fa0

Pith citing papers

No inbound Pith citation observations are available.