Pith. sign in

Paper Citation Record · LEDGER

Effective Red-Teaming of Policy-Adherent Agents

As of 10 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 2 inbound Pith citation observations for arXiv:2506.09600.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.09600 v3

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:50:00.253920Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T07:48:36.924294Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T07:54:22.258152Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a9fe4ffe-27a7-4b2d-bac8-777d4d8dd59c · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.119981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.119981Z digest=sha256:39123e38acba4247bf8efcb8ccfb4a7b17267c0841732b2dcc2b0c4539b4a5b0

Observation ed34efe0-0318-4a5e-a5e3-dc1c3f58dbd8 · outbound

This paper cites AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents.

Effective Red-Teaming of Policy-Adherent Agents AgentHarm: A Benchmark for Measuring Harmfulness of LLM Agents

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.124363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.124363Z digest=sha256:3a2d64581dee4a529505cb2c76c640ac608dbaeae9e9cd6c13605ecff702a604

Observation 5bd98dfd-fade-4c92-993d-2104cf826649 · outbound

This paper cites A Framework for Testing and Adapting REST APIs as LLM Tools.

Effective Red-Teaming of Policy-Adherent Agents A Framework for Testing and Adapting REST APIs as LLM Tools

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.129367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.129367Z digest=sha256:8aafc3565c3f3329ba4833153d6af894bcfa544d72626e6bdf4c32fca5b65d8b

Observation 03ed9a8b-08c1-47cf-a062-174d7e1d1cab · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Effective Red-Teaming of Policy-Adherent Agents Evaluating Large Language Models Trained on Code

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.133366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.133366Z digest=sha256:04e6e161f0f5ce393162fa46f949fe4253f9672e5387ca245df6e963288fedd2

Observation a98cffe2-f85d-496d-9947-cd94c9849bc6 · outbound

This paper cites The Llama 3 Herd of Models.

Effective Red-Teaming of Policy-Adherent Agents The Llama 3 Herd of Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.137097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.137097Z digest=sha256:d4c257734461c10df22ef48cb5f278063b3f19580878a37726a106f107ed242f

Observation 1e9eb559-74f2-464a-8fe8-116158765a52 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.987004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.140744Z digest=sha256:f81b430ab8c26320f5a84661693bec7e7f64b860c5b4ec261271ef68f80ac1b9

Observation e6252988-2705-441d-be6b-a23da52a6621 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.976219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.144239Z digest=sha256:abfe2afe269b6c9ae655f8baf90b8bada2cf8c354efcaa333ad360ffb350ddf7

Observation 2d02d8c2-ef73-4a6e-9b5b-998a49c59faf · outbound

This paper cites GPT-4o System Card.

Effective Red-Teaming of Policy-Adherent Agents GPT-4o System Card

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.147390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.147390Z digest=sha256:3886f59266b6b2963744e4405278a7108896aab8564f5c3ffba98447237077ec

Observation 0fc1f222-66e5-48a6-9d06-cccaa70e343d · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.150733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.150733Z digest=sha256:18c444e80e686222461b389a59b54cbd96ba27ed9402b9a7673c869bba70a502

Observation 48730d43-8fe6-4f01-a378-2c588ee6b58a · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 10

Resolution
verified exact
raw_fallback, observed 2026-08-07T04:50:00.672247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.154869Z digest=sha256:63aa017ade5533be43ac701e7ab2c681ffd8463389e03937f9a1affed4c398e3

Observation 986a60bd-ecac-46ec-ba2d-f5e102ed3cfc · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.965469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.158262Z digest=sha256:e28d63e7e25e0993b45b0a77a33d58c82685ff0568ac0c8a9757e136e846c6b8

Observation 7ceccfad-f25e-486e-adc7-2455580e1072 · outbound

This paper cites Unveiling Safety Vulnerabilities of Large Language Models.

Effective Red-Teaming of Policy-Adherent Agents Unveiling Safety Vulnerabilities of Large Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.161709Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.161709Z digest=sha256:d7f6682c07f335e5ac6b0dc560f82a1ef289212a2a3bf00035a7d85389544c13

Observation cc16bd91-aabd-4430-a712-cd7800f2b6f4 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.953787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.165546Z digest=sha256:a2b5b575aea992a5ff452dec7d74edddd3d51b376d8181a12d88d250374e047d

Observation 6c005901-1f7d-42ed-895b-fca8e3ed9781 · outbound

This paper cites ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents.

Effective Red-Teaming of Policy-Adherent Agents ST-WebAgentBench: A Benchmark for Evaluating Safety and Trustworthiness in Web Agents

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.168901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.168901Z digest=sha256:159e845c10641f23792dc2714497a13769903f53a22db6253dff0990e85aa16c

Observation d2052270-87ab-4262-b1ae-ae4e9432038d · outbound

This paper cites SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints.

Effective Red-Teaming of Policy-Adherent Agents SOPBench: Evaluating Language Agents at Following Standard Operating Procedures and Constraints

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.172498Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.172498Z digest=sha256:2ce4ff322746dfb8a0a45dbc34f612c21ab3ace19f569e24b90421bac8d949ab

Observation 1fc9cdf3-3095-4c91-9b47-b3ec89dd02ca · outbound

This paper cites DeepSeek-V3 Technical Report.

Effective Red-Teaming of Policy-Adherent Agents DeepSeek-V3 Technical Report

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.176032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.176032Z digest=sha256:b358b0b8b6cf25dbcd442f3af5cfc5c5ff954ae2db9a1c7b04866e773b0bbcf1

Observation 45a07ed6-ae4e-4bc8-8523-8826db4d0170 · outbound

This paper cites From LLM to Conversational Agent: A Memory Enhanced Architecture with Fine-Tuning of Large Language Models.

Effective Red-Teaming of Policy-Adherent Agents From LLM to Conversational Agent: A Memory Enhanced Architecture with Fine-Tuning of Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.179402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.179402Z digest=sha256:52f436ac0f196ae180658c33a20832521b4a3948ae00af757a1a1486cc894d35

Observation e4aa5648-db3b-4be7-ade6-9fd1c50de94e · outbound

This paper cites AgentBench: Evaluating LLMs as Agents.

Effective Red-Teaming of Policy-Adherent Agents AgentBench: Evaluating LLMs as Agents

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.182946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.182946Z digest=sha256:55a608347a1e3200ffc0e60a0b8fc5624ca71ef4eb785cf7f6f111a6f15f46b5

Observation 50533e89-9e21-42b8-844c-5106ca7e14f1 · outbound

This paper cites Prompt Injection attack against LLM-integrated Applications.

Effective Red-Teaming of Policy-Adherent Agents Prompt Injection attack against LLM-integrated Applications

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.186499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.186499Z digest=sha256:e506d717dc1505ec34e8c69a8b3847056bc04cef8d8aac438ba341d7e44b0076

Observation a358ba86-58cc-4329-87c5-f8f28fee284d · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.943574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.190116Z digest=sha256:4c4d593235a2dce450a5b5814cc975af6dc3894dbef63126dcbd2bbbbbe82292

Observation 0fde2ab8-b1f9-4dd3-afcb-8dacea8d9941 · outbound

This paper cites u ndler, Mark Niklas M \.

Effective Red-Teaming of Policy-Adherent Agents u ndler, Mark Niklas M \

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:50:00.933475Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.193419Z digest=sha256:b43f3de8fb109d65e73a735aa1956c5167ad6299f592fe3b4f65cb31bcadd741

Observation 5b29a17e-92b7-409b-b739-9d0cd5e120c8 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.923174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.196714Z digest=sha256:e68c9cbd6219509b4f7e17afab255186f6cfb8cca8f16e0d34eabc6c25b64e5e

Observation e7683656-a1bd-4537-aaef-7b477760998e · outbound

This paper cites do anything now.

Effective Red-Teaming of Policy-Adherent Agents do anything now

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T04:50:00.912398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.199724Z digest=sha256:b2c78811929c976681fcad990eeef2efa2c4531b320e22b16f8b0aec6c3102a8

Observation 36f055a0-5e8a-40fb-9d6b-71cf64ba9975 · outbound

This paper cites do anything now.

Effective Red-Teaming of Policy-Adherent Agents do anything now

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.202923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.202923Z digest=sha256:aa7ac7ed164530ca9b3b29574f4daf3ddb45823cd69c1049035ec494c4ff686e

Observation f46532e3-00d6-4dd6-99fb-2a643e881f52 · outbound

This paper cites CHOPS: CHat with custOmer Profile Systems for Customer Service with LLMs.

Effective Red-Teaming of Policy-Adherent Agents CHOPS: CHat with custOmer Profile Systems for Customer Service with LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.206194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.206194Z digest=sha256:ded7d438084bb727356f2f6ec6c54fa7675aa828257ae4e015f0a996b9e36ebd

Observation ada563b1-aadb-4908-9872-59d56c02826d · outbound

This paper cites Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers.

Effective Red-Teaming of Policy-Adherent Agents Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.209867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.209867Z digest=sha256:98a1a8bf95ead8518f7488a2a9c2d87fcc8b9ba96cfd8d32bf0624beba8ffb67

Observation 04ecd511-17b3-4848-8819-74ac326f158f · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.214612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.214612Z digest=sha256:2c97a5f246a2f398839529d5d1a126f64e7c0a7ee5b42a6982e09010b6d28e48

Observation c7f4545a-55a6-4feb-8926-256a2478f566 · outbound

This paper cites Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles.

Effective Red-Teaming of Policy-Adherent Agents Multi-Turn Context Jailbreak Attack on Large Language Models From First Principles

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.218144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.218144Z digest=sha256:64ca7edc38db1e714aa771afdc3fefe93be2af91d9cdad535d229111d1f96ecc

Observation 0066ba6d-3785-41fd-96c2-443202205f55 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.895130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.221806Z digest=sha256:d9f5ff811a9444d417614a63ce8078687f9d34174280a65c6b5904c36386af0f

Observation 0c447225-0981-48ad-aafa-07b43f553d88 · outbound

This paper cites Emotional Manipulation Through Prompt Engineering Amplifies Disinformation Generation in AI Large Language Models.

Effective Red-Teaming of Policy-Adherent Agents Emotional Manipulation Through Prompt Engineering Amplifies Disinformation Generation in AI Large Language Models

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-07T04:50:00.289354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.225273Z digest=sha256:be1c0acd5dd2dffc02f5d70e1b6324d25750467695a71fac02199bdbbfef8981

Observation 617a1164-8ca5-4f15-bafa-2f5dfbfabcb7 · outbound

This paper cites Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise.

Effective Red-Teaming of Policy-Adherent Agents Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.228964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.228964Z digest=sha256:c9d347385d62edf01c370aa396f1b96191f5b6f5486b7774957555d79e8313fa

Observation 5f6b1128-03a9-4a3a-885e-aefa408de241 · outbound

This paper cites Qwen3 Technical Report.

Effective Red-Teaming of Policy-Adherent Agents Qwen3 Technical Report

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.232482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.232482Z digest=sha256:a228bf5f4e6f4669c51e5bf873a9016adde64dd2483ca7ea29d796f18a30f17c

Observation f2aa3b30-54c0-4f1e-8b3f-35ebd1d96784 · outbound

This paper cites $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains.

Effective Red-Teaming of Policy-Adherent Agents $\tau$-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.235971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.235971Z digest=sha256:d4592e7569e3d71e037e75eac237127aca65dad3321c87ce2f066f01f29d2653

Observation bdddd043-26b7-436e-9c0a-a2d6f4633c16 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.239482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.239482Z digest=sha256:9ebdbfea78e48491c8196d16ecfb883fc6d132bd20a40593e757d32b5e2bbfe0

Observation a67f23d5-4b4a-419f-8f27-be2a1c2ced3d · outbound

This paper cites Survey on Evaluation of LLM-based Agents.

Effective Red-Teaming of Policy-Adherent Agents Survey on Evaluation of LLM-based Agents

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.243042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.243042Z digest=sha256:c768ded8a52ee2c67ac04434bdb74adf36265e5a05e26e0ba5c1ed3ab2cfa925

Observation 816d0e1b-59bd-4e90-9fb5-a609d7754cc9 · outbound

This paper cites an unresolved cited work.

Effective Red-Teaming of Policy-Adherent Agents Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T04:50:00.877545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-07T04:50:00.246609Z digest=sha256:c116d09338e34e07740c5efdaa10745f65b2686b65dd1dea8ae093f73cffc4c0

Observation 80feea2e-0a2a-4cb2-a8a6-236e3e9bf407 · outbound

This paper cites online" 'onlinestring :=.

Effective Red-Teaming of Policy-Adherent Agents online" 'onlinestring :=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.249996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.249996Z digest=sha256:3f5440f8957c1c1dcf613335ada4f75f6d9212a4922c24718e1552349f5cffa0

Observation 82e72879-d510-4f14-9dad-a2a6be99f225 · outbound

This paper cites write newline.

Effective Red-Teaming of Policy-Adherent Agents write newline

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T04:50:00.253920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:50:00.253920Z digest=sha256:771cea68132a7afd6b14ce38245aa1b09f65a1011f6c1d14f6d7fc5b3dec6f47

Pith citing papers

Observation 29492dc8-2762-4be4-a9ef-17672f85925e · inbound

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation cites this paper.

Don't Pass@k: A Bayesian Framework for Large Language Model Evaluation Effective Red-Teaming of Policy-Adherent Agents

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-05-18T10:06:13.730575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-18T10:04:39.223895Z digest=sha256:df3b6b526e5e21b2f775e47c3fd1610240af6b34136829ac3995d61842f8d9e1

Observation e24b196a-a0fd-4f1c-8ca4-5f9ae0e3084f · inbound

PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents cites this paper.

PolicyGuard: A Dialogue-Grounded Sub-Agent Verifier for Policy Adherence in LLM Agents Effective Red-Teaming of Policy-Adherent Agents

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-06-30T07:54:22.260097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-06-30T07:48:36.924294Z digest=sha256:d3b806d9741841ed170d148f1d60826ae13698923eea23480025cc9014a5044c