Pith. sign in

Paper Citation Record · LEDGER

Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 11 inbound Pith citation observations for arXiv:2407.21792.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.21792 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 11 of 11 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:03.251847Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T21:10:09.315037Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 26308fe6-5800-452c-9d4c-c15aaa28442c · inbound

Proof-of-Learning with Incentive Security cites this paper.

Proof-of-Learning with Incentive Security Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 51

Resolution
verified exact
arxiv_id, observed 2026-05-24T02:25:55.530098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-24T02:24:17.188754Z digest=sha256:ce89cab31e7c16298b29b0da5b7b13fc862972d2e9083c15f88c4e29cff4a3d1

Observation f8fb76d6-0c93-4dd7-8984-47370e0825a7 · inbound

Safety Degradation in AI Agents cites this paper.

Safety Degradation in AI Agents Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:03.251847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:41:03.251847Z digest=sha256:3aa1f80c589186a365089a305c0d12e3ec0b0e80f9437ab9f71aa16385ad1903

Observation e6b248f2-edd8-40a4-8578-c47c493949fc · inbound

Lessons from Defending Gemini Against Indirect Prompt Injections cites this paper.

Lessons from Defending Gemini Against Indirect Prompt Injections Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T15:37:51.682677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:37:51.682677Z digest=sha256:5652745e01a74391c27f24cb0a5dbf465386dac809537c344e74723233b0a42c

Observation 46d58363-7caa-4c6c-86aa-ddf5c16ec44d · inbound

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts cites this paper.

SAGE-Eval: Evaluating LLMs for Systematic Generalizations of Safety Facts Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:28:23.620979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:28:23.620979Z digest=sha256:a842a1ab7ed8a0aff7599e37a00a9d8fe829bcd408b3ae870e247f93381811eb

Observation a53b041b-1dfd-4607-9626-78256768fe15 · inbound

The bitter lesson of misuse detection cites this paper.

The bitter lesson of misuse detection Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:40.663885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:40.663885Z digest=sha256:bad5b8440575f1f0df08708b14d1e8f818f99a9c027fa6d926375c20bfc3f24f

Observation d43b1df5-5003-4882-82af-122502291098 · inbound

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules cites this paper.

Blind Refusal: Language Models Refuse to Help Users Evade Unjust, Absurd, and Illegitimate Rules Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T19:28:09.978058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-13T19:24:54.381722Z digest=sha256:4769ec520095e1c7478c37f82817b0c3d204dacb593d4875a57a6c793cfa556d

Observation 905caca1-94da-40e4-9adc-a3a18b023ce9 · inbound

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents cites this paper.

Taxonomy and Consistency Analysis of Safety Benchmarks for AI Agents Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-21T01:43:56.649657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-21T01:42:55.693115Z digest=sha256:a0f72f3b01a29d2287e230b8bcaec8c445ba37f8f49c49df6ae1ef7a15edc011

Observation 34045775-d44e-43ae-bedc-d26990cc8319 · inbound

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting cites this paper.

Evaluation Cards: An Interpretive Layer for AI Evaluation Reporting Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:07:33.268541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T16:11:36.483820Z digest=sha256:25dfbc8ca1a50bb264830b981b24811354c043810f362f475e8e621bbc067708

Observation 855cb524-3d24-4c41-a8e0-c5cf607fc39b · inbound

A Pigouvian Matchmaker Mechanism for De-escalating the AGI Race cites this paper.

A Pigouvian Matchmaker Mechanism for De-escalating the AGI Race Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T08:07:45.444019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-27T11:03:42.443180Z digest=sha256:38507f033bc5e24e0dc48f7ec178ab0187c6b8f4f72892d2252c32f2e59584e1

Observation 158fec31-5471-46d3-b6fe-18feee98accd · inbound

The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems cites this paper.

The Unfireable Safety Kernel: Execution-Time AI Alignment for AI Agents and Other Escapable AI Systems Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T21:10:09.316660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-25T19:01:27.048159Z digest=sha256:8d74e80960fd16c35842e32957d4289acd8b8c9fbbc7188c61ca7ef6edb3e7b4

Observation 9230e5db-e968-4e5d-9114-e29c6ecd2680 · inbound

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins cites this paper.

Exposure is not manifestation: measurement target and output resolution jointly determine which behavioural-faithfulness evaluator wins Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-02T07:44:09.409118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T07:44:09.409118Z digest=sha256:197dd10daab66fed98ee5db039dc4c33cefc53cde3872423a402faa3d0b0f4f7