Pith. sign in

Paper Citation Record · LEDGER

The bitter lesson of misuse detection

As of 7 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 0 inbound Pith citation observations for arXiv:2507.06282.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.06282 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:18:40.850181Z

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy9
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2c91e3e0-d55e-4ba2-af12-aebd325fadf1 · outbound

This paper cites GuardBench: A Large-Scale Benchmark for Guardrail Models.

The bitter lesson of misuse detection GuardBench: A Large-Scale Benchmark for Guardrail Models

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:18:43.818607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:18:38.633982Z digest=sha256:37d347b50ac14ddb821f2fe23af475ebf59ea30cce00c8b485d52373f90e3b7b

Observation a52d6d52-2bd2-43c6-b438-7005ac68bff6 · outbound

This paper cites Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming,.

The bitter lesson of misuse detection Constitutional Classifiers: Defending against Universal Jailbreaks across Thousands of Hours of Red Teaming,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:18:43.563105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:18:38.761731Z digest=sha256:819291d731ee124ecd140af83cde55f3d2f3bed3e82536517f978cde36894c50

Observation f8c82631-8dd1-4ee7-8b64-f927f3aa1a4a · outbound

This paper cites NeurIPS 2024 Datasets and Benchmarks Track, 2024.

The bitter lesson of misuse detection NeurIPS 2024 Datasets and Benchmarks Track, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:18:43.272682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:18:38.810817Z digest=sha256:120cb2583c89ed924e160ecac896be62e2f6837a942448f7eff3f7e4bac73d34

Observation 039fb256-65d4-4e2c-9e55-e0bbc9b68aee · outbound

This paper cites "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models.

The bitter lesson of misuse detection "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:38.885620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:38.885620Z digest=sha256:8c3a9e8d064bac65f060a96a5fb66dda5b4ecfe790cc5babb27da4d855c1db4a

Observation 52615106-1ebe-4879-9046-de714e9e8211 · outbound

This paper cites SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal.

The bitter lesson of misuse detection SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:39.027078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:39.027078Z digest=sha256:6f7ff3022db238576df2d0a1d306fb7b811c420cbe2f38521863a9b3ecffdad6

Observation df1cfef7-9e07-4e97-95c8-dabb2346ac2e · outbound

This paper cites AdvBench: Universal and Transferable Ad- versarial Attacks on Aligned Language Models,.

The bitter lesson of misuse detection AdvBench: Universal and Transferable Ad- versarial Attacks on Aligned Language Models,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:18:43.058381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:18:39.230798Z digest=sha256:13bc6d3ff503f0722a6dd3dfe7e0b035f3591cc527f6474cdf42e10dafdad3f1

Observation 23662cfe-a460-48e7-96b9-ee0dac2f98c2 · outbound

This paper cites CatQA: A Dataset for Categorizing Questions as Safe or Unsafe,.

The bitter lesson of misuse detection CatQA: A Dataset for Categorizing Questions as Safe or Unsafe,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:18:42.891885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:18:39.311873Z digest=sha256:c5400e9ac183771ec8c8ee7a441036dd758a5d07358d9fa74ebd0579e84e7ca5

Observation f1c918c9-da12-4d60-8c4e-c6abb5eddbd7 · outbound

This paper cites Do Not Answer: Testing AI Refusal to Unsafe Questions,.

The bitter lesson of misuse detection Do Not Answer: Testing AI Refusal to Unsafe Questions,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:18:42.701691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:18:39.365817Z digest=sha256:fd5989f17da095d5bec7e614ef8dcf1a54b3b1b37166ec29013d8e492c8405e7

Observation c0105b1f-cfac-43dd-8207-35d7d779c5c9 · outbound

This paper cites HH-RLHF: Helpful and Harmless Reinforcement Learning from Hu- man Feedback,.

The bitter lesson of misuse detection HH-RLHF: Helpful and Harmless Reinforcement Learning from Hu- man Feedback,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:18:42.506895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:18:39.472725Z digest=sha256:e4fae61a58ccb725a77aa6ee812501c1f5d0d390d4889544a89a0739636fca6c

Observation 8543b09b-adba-4383-be52-c9d1d4dff676 · outbound

This paper cites HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal.

The bitter lesson of misuse detection HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust Refusal

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:39.588048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:39.588048Z digest=sha256:6bf3a7e2775df21251a7128e2ae1cc89e43db0ba958f8e13da17e4887fd34f6d

Observation 9a3352f7-c722-4e0a-92c8-3a166a2de097 · outbound

This paper cites A StrongREJECT for Empty Jailbreaks.

The bitter lesson of misuse detection A StrongREJECT for Empty Jailbreaks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:39.703539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:39.703539Z digest=sha256:7b8340d302e3a991501a7b7ed685e9fb450046ff4198411d7d3bd9ae6ba59fc6

Observation f6b9e1a3-e909-4b4a-9aab-9f0e404c9aac · outbound

This paper cites The Twelfth International Conference on Learning Representations, 2024.

The bitter lesson of misuse detection The Twelfth International Conference on Learning Representations, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:18:42.239110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:18:39.798982Z digest=sha256:f21e53bfd1818537ebc9948c39b47bc0cc036fb09ebc87a0943987287082ef0a

Observation 82e96a74-af9e-44c2-99d9-b2a6927de0eb · outbound

This paper cites an unresolved cited work.

The bitter lesson of misuse detection Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:42.037851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:18:39.887601Z digest=sha256:ab8565be663dc15509bf9f213354d98b4b1df84f51b579aa38fc001297f3153c

Observation 219f32df-033c-482f-bf49-dc911055af8f · outbound

This paper cites an unresolved cited work.

The bitter lesson of misuse detection Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T19:18:41.743137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:18:40.022590Z digest=sha256:44eeac1179e7118664851a1672073e86b47bab0c53f9cda37cfb1568cd452c0b

Observation 7efef56f-9ccc-4f3e-b0c4-589bac84fd58 · outbound

This paper cites DeepInception: Hypnotize Large Language Model to Be Jailbreaker.

The bitter lesson of misuse detection DeepInception: Hypnotize Large Language Model to Be Jailbreaker

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:40.154024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:40.154024Z digest=sha256:f79b3506a78f37246140c6e319577eb3c62c451560fc5c0f8eb8fd87d45ad6e6

Observation 05f8fcb2-eb3f-441e-99f1-8a526aa9d2fe · outbound

This paper cites Universal and Transferable Adversarial Attacks on Aligned Language Models.

The bitter lesson of misuse detection Universal and Transferable Adversarial Attacks on Aligned Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:40.283148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:40.283148Z digest=sha256:4c339a71632aa78310109790d773e6f814fd86b6c3b1cf344dfb182e5365fda1

Observation 25f66079-b23c-41d8-b039-4b95459d3b93 · outbound

This paper cites Jailbreaking Black Box Large Language Models in Twenty Queries.

The bitter lesson of misuse detection Jailbreaking Black Box Large Language Models in Twenty Queries

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:40.429244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:40.429244Z digest=sha256:80ecbd0a49984529761a32babdc6ffaeb96d6e97923b71ed1ada173c76748b7f

Observation cdba531b-0cde-4c13-b6f1-c526f77578c9 · outbound

This paper cites AI Control: Improving Safety Despite Intentional Subversion.

The bitter lesson of misuse detection AI Control: Improving Safety Despite Intentional Subversion

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:40.580182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:40.580182Z digest=sha256:988efe6d767e20ed66913462a8192671995b35c72c62fa09920584de682a1070

Observation a53b041b-1dfd-4607-9626-78256768fe15 · outbound

This paper cites Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?.

The bitter lesson of misuse detection Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:18:40.663885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:18:40.663885Z digest=sha256:a567027e7acc970f3ad20a91c2a35486f118481b4a8ec698ff4fa12aa34c1932

Observation 15b373f1-d82b-4fc5-b86b-a2de051580f3 · outbound

This paper cites Is this prompt harmful or not?.

The bitter lesson of misuse detection Is this prompt harmful or not?

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:18:41.442675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T19:18:40.850181Z digest=sha256:05048a48d3b555232785e230e04a6dbd073cd5e49c058701b3f3a5dbd64ba062

Pith citing papers

No inbound Pith citation observations are available.