Pith. sign in

Paper Citation Record · LEDGER

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces

As of 18 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2606.01317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.01317 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-28T16:44:18.994680Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved21
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b0a007ea-a870-4cc8-8a61-7d1dd199d7fc · outbound

This paper cites Red Teaming Language Models with Language Models.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Red Teaming Language Models with Language Models

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:36:15.071498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:1045c99864875c6ad4e4aac2f65b1622b662a4724675c6d202de8714df2738c9

Observation 4c814a72-2663-40e0-ab91-eb609f048e7b · outbound

This paper cites Proceedings of the 16th.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Proceedings of the 16th

Reference 2

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:9b2585e2dd2e84e343222977ff40cecb0d092d07fb5658a7f93e77855e1e41f5

Observation 7b90fe33-20b8-4630-90f4-95fefb774868 · outbound

This paper cites Extracting Training Data from Large Language Models , journal =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Extracting Training Data from Large Language Models , journal =

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:1b10e1e1f7371a4d1e1e73f033fd20b1962fc5cc6b74edcafc0b05974c2e41e2

Observation d7459563-31e4-4fdc-9422-086420ed3837 · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces AgentBench: Evaluating LLMs as Agents

Reference 4

Resolution
metadata mismatch
local_arxiv, observed 2026-07-01T21:36:15.069087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:84bfa48e5bae4f92d740b9f0a4fb1927b90603c74380f016d3e67294631fe61b

Observation 9f57470a-be26-4ef8-97cd-cbd67dd9186b · outbound

This paper cites an unresolved cited work.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:33e61208fcc8f6073955ff75e2f166fe9c5270ea5dd9111585ae360b213c495d

Observation 1e53ebb4-1dec-42d3-bd00-6ed89d372d3c · outbound

This paper cites Forsyth and Dan Hendrycks , title =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Forsyth and Dan Hendrycks , title =

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:feca6181a81fd87f34198d175088a2077412f9133aebfffeb74238cf966f4aaf

Observation 3f34189b-fc60-43d2-b1bb-6419cf2d3b57 · outbound

This paper cites Chasing Shadows: Pitfalls in.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Chasing Shadows: Pitfalls in

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:95da07a4dddd8d9955cc872fdbc0eb6e85faec2d53f1491753d872d8c02fee17

Observation a2f23be2-0ab0-49d4-812b-3b7bf5847f92 · outbound

This paper cites Findings of the Association for Computational Linguistics (.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Findings of the Association for Computational Linguistics (

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:66ba568668eaf50b8defce73e3bda3231d089232822b7a8bef646a80ccaa889c

Observation 8c3ecaf5-4212-4eee-81a3-8fb7f4900011 · outbound

This paper cites Zico Kolter and Matt Fredrikson and Yarin Gal and Xander Davies , title =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Zico Kolter and Matt Fredrikson and Yarin Gal and Xander Davies , title =

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:4be554383b16834becd7388f4b5d11c80bbe027a39185f10bbb523f7c4ad511b

Observation 4a39f42e-d611-4a64-b9ae-ca9df611ec69 · outbound

This paper cites Advances in Neural Information Processing Systems (.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Advances in Neural Information Processing Systems (

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:a68fd4205287744889127a28c59f19ab848b89421d85b4b46172a1c938d905fa

Observation 13bcc762-ad98-4574-9b08-55292b7454a4 · outbound

This paper cites 2026 , eprint =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces 2026 , eprint =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:ddf95b766c69d73cc4bbe8ea1cc9c2c675fac36ec005cb2853294abd26c90388

Observation 32c96666-bce1-412f-a7ea-35e976b6acc0 · outbound

This paper cites Findings of the Association for Computational Linguistics (.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Findings of the Association for Computational Linguistics (

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:0ab671ac0f6dbcdc5d96f9e2df314d0c9b18e45bb1be335281bb1fb56312aed2

Observation d2321bcc-c45c-4a05-8b44-86d74c8aca49 · outbound

This paper cites Qin, Y ., Liang, S., Ye, Y ., Zhu, K., Yan, L., Lu, Y ., Lin, Y ., Cong, X., Tang, X., Qian, B., et al.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Qin, Y ., Liang, S., Ye, Y ., Zhu, K., Yan, L., Lu, Y ., Lin, Y ., Cong, X., Tang, X., Qian, B., et al

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-07-01T21:36:15.074645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:b35358aded2c4c99bc17c38ff4644b218e2f2a23cffa2c4c2cc9d943099f68fb

Observation b2878297-f368-4b5f-ab3b-26ac00db564f · outbound

This paper cites an unresolved cited work.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Unresolved cited work

Reference 14

Resolution
parse uncertain
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:b52fba201a8ab52212bc614708f02fb3d2e4eca946396562790f7dd74baefd1c

Observation 9c7489b0-7248-47a7-8e96-96eb6eb8252b · outbound

This paper cites 2025 , howpublished =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces 2025 , howpublished =

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:da2b82d869626fca41bc740a531f3895dc4fd323872b420f2e135a153a68171e

Observation a553fda0-5d2c-4d78-9968-6217012cc40a · outbound

This paper cites Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks , journal =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Skill-Inject: Measuring Agent Vulnerability to Skill File Attacks , journal =

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:7350992c78e6307a4b2cf17df629c39cd09cb51ba657807ffe3c675d69c9088f

Observation a0a44ef4-f4a6-4b81-9857-e86861ec49e5 · outbound

This paper cites AgentDojo:.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces AgentDojo:

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:feaac6642649021db3f495133ab073dd1a8af088cb12c227e888c8cad48a18cb

Observation ca446d45-f632-4055-8ae5-f677e5036214 · outbound

This paper cites Zico Kolter and Matt Fredrikson , title =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Zico Kolter and Matt Fredrikson , title =

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:81c2d5273fc889bea7cda63d6990c62844b237f440ee2b7e4a283329e13261fe

Observation f1a0435e-8438-4dd0-8844-e0b1829426ff · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces The Thirteenth International Conference on Learning Representations,

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:0dd32fb097c5cf26381970e5a04886b9bf5bf86edbbbc85fc83e40598853535b

Observation 88d93696-3e02-44b6-bf41-15f7cc886619 · outbound

This paper cites OR-Bench: An Over-Refusal Benchmark for Large Language Models , booktitle =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces OR-Bench: An Over-Refusal Benchmark for Large Language Models , booktitle =

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:d0f1289a52c1dff3c7aa1dfc6fb08ed2980dd6a9305213b4049045f6b764d917

Observation 1cc0622c-f20f-4d7e-91c4-cbbbfeea2273 · outbound

This paper cites The Twelfth International Conference on Learning Representations,.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces The Twelfth International Conference on Learning Representations,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:e00639e2f23dee4fa429d9495591a23344df8a666cfcf11a860ebbcf25cc7ae6

Observation ca95514f-71b0-451c-a019-a58b65b88753 · outbound

This paper cites The Thirteenth International Conference on Learning Representations,.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces The Thirteenth International Conference on Learning Representations,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:a51989e41079340964ba8ebd3e6cb6e66509409c056b93779628ce1806ce9aa6

Observation 63957f8f-a573-41e9-90bc-526d4c133cef · outbound

This paper cites Maddison and Tatsunori Hashimoto , title =.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Maddison and Tatsunori Hashimoto , title =

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:b32a99e71d40dffbd10accb77964e1f11fb5c08374548e366897c9054a78a998

Observation bf7be383-ee5b-430c-98b0-968c893b5343 · outbound

This paper cites an unresolved cited work.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:12c1a704b9a357e606f4ea53eb78feaf814dcc7aa4e9ea32e9c7f8040171eac3

Observation 0e47e741-ab83-4a5a-aa3c-4c83b18a6e4d · outbound

This paper cites Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Purple Llama CyberSecEval: A Secure Coding Benchmark for Language Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-07-01T21:36:15.066149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:1065f08c1f485437a993c7d69240065bb1e7403c024b5958ed54c86b6f6c2d9b

Observation 71153747-9fc8-4aaa-b81c-dc3add28371f · outbound

This paper cites Findings of the Association for Computational Linguistics:.

SABER: Benchmarking Operational Safety of LLM Coding Agents in Stateful Project Workspaces Findings of the Association for Computational Linguistics:

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-28T16:44:18.994680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-28T16:44:18.994680Z digest=sha256:50abbaff2f915a565cc06bba95045121c6f4157d6d1170c1f8ebbd6d2292f3ce

Pith citing papers

No inbound Pith citation observations are available.