Pith. sign in

Paper Citation Record · LEDGER

Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2407.21792.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.21792 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:03.251847Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.315037Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26308fe6-5800-452c-9d4c-c15aaa28442c · inbound

Proof-of-Learning with Incentive Security cites this paper.

Proof-of-Learning with Incentive Security Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:25:55.530098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-24T02:24:17.188754Z digest=sha256:e115bf96fb3c0832045059a695c5a54ee53648ac9b2ad9c1e5787d5e0dd92ee7

Observation f8fb76d6-0c93-4dd7-8984-47370e0825a7 · inbound

Safety Degradation in AI Agents cites this paper.

Safety Degradation in AI Agents Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:03.251847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:03.251847Z digest=sha256:3231d6786965934c20d837cb0d4490d2c3dbc6f2b556d3e9e7f700bebd526cb7

Observation e6b248f2-edd8-40a4-8578-c47c493949fc · inbound

Lessons from Defending Gemini Against Indirect Prompt Injections cites this paper.

Lessons from Defending Gemini Against Indirect Prompt Injections Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:51.682677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:37:51.682677Z digest=sha256:d2e6203ffcf41525f51c078061579943466a24ef7207292b32ea5bc6bd4a1b6f

Observation 46d58363-7caa-4c6c-86aa-ddf5c16ec44d · inbound

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts cites this paper.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.620979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.620979Z digest=sha256:5663808c7f23ffd1dc6b3a9fcc11bb793f1e4a7aebe5311c20c6a23c2630fa1d

Observation a53b041b-1dfd-4607-9626-78256768fe15 · inbound

The bitter lesson of misuse detection cites this paper.

The bitter lesson of misuse detection Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:40.663885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:40.663885Z digest=sha256:a567027e7acc970f3ad20a91c2a35486f118481b4a8ec698ff4fa12aa34c1932

Observation d43b1df5-5003-4882-82af-122502291098 · inbound

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules cites this paper.

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:28:09.978058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T19:24:54.381722Z digest=sha256:22da766221527a8b1863f5541270b2dbb77976649651a7407c9c204675445bf7

Observation 905caca1-94da-40e4-9adc-a3a18b023ce9 · inbound

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents cites this paper.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.649657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:d466efed29f0b8763968b861acc5d5c8be0f5fbe50ee447f8496860ccab9b03d

Observation 34045775-d44e-43ae-bedc-d26990cc8319 · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.268541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:bb9533c533d3f712f321da97b55371a6f16ab6b142f3ac5f8ed7c77ced7d3a73

Observation 855cb524-3d24-4c41-a8e0-c5cf607fc39b · inbound

A Pigouvian Matchmaker Mechanism for De-escalating the AGI Race cites this paper.

A Pigouvian Matchmaker Mechanism for De-escalating the AGI Race Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:07:45.444019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-27T11:03:42.443180Z digest=sha256:3a1832c9d8d47aae904bdaf620aa67902b2855d160e1a88242719f78cfb5a5ab

Observation 158fec31-5471-46d3-b6fe-18feee98accd · inbound

The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems cites this paper.

The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:10:09.316660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T19:01:27.048159Z digest=sha256:ff556d172f12349e0039ceb159e1175611bc965e4b00800aeab02180f1a3537a

Observation 9230e5db-e968-4e5d-9114-e29c6ecd2680 · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:09.409118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:09.409118Z digest=sha256:197dd10daab66fed98ee5db039dc4c33cefc53cde3872423a402faa3d0b0f4f7