Pith. sign in

Paper Citation Record · LEDGER

Open Sesame! Universal Black Box Jailbreaking of Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2309.01446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.01446 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:22:00.465363Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T19:00:30.340656Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a9e0449a-d2cf-4d77-a059-af389bc5cba1 · inbound

Low-Resource Languages Jailbreak GPT-4 cites this paper.

Low-Resource Languages Jailbreak GPT-4 Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:24:14.027167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T09:24:13.911401Z digest=sha256:a773cc3b268ab9c62f0e8c9361b3997e39b988aac4ea8cca98b64132e2253bd5

Observation 391b345f-5beb-46e7-9578-5914dc60264a · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 245

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:17:08.691983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:7eb89c042fd468e1b141a50834a8a08e8869b77b36f567059dd00d1ceee8876b

Observation 6755de70-5077-48ad-bdfc-2a14518faf4b · inbound

A StrongREJECT for Empty Jailbreaks cites this paper.

A StrongREJECT for Empty Jailbreaks Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:02.830434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T21:28:02.745230Z digest=sha256:6c586721ef934e48bcd63f378932112c439777519abfb41300455fe65dbd1bbe

Observation fb3e3ded-af34-448d-8780-349c26f73a04 · inbound

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models cites this paper.

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:08:05.637785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-15T06:08:05.386345Z digest=sha256:c969d82030187d932c7cc14b76d4c503f03725720709ecce34f3fc0b6ae0fb1d

Observation 45ced3d8-724b-4da4-a873-530e46cf695c · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.705732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:511bcb32de01f8b0ca54da9ca823e77b90e10a587a43449355e2b4ff9ac4b492

Observation d9172f6b-7dbe-4588-9038-23b5edfb9671 · inbound

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs cites this paper.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.465363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.465363Z digest=sha256:7d8eda04f71043123dbdb7ce5071d8b704e97404837eba2df01802b125fd2809

Observation 5d779aee-93d4-485b-b0e5-2ed58e4f5896 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:02:44.303435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:e8049cadb802c1780cf014b218cd022bbe808a4e452c12935fedb0b1a9d82dac

Observation 5df0e263-d0cd-4aa3-ae20-573e45d7e151 · inbound

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval cites this paper.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.657664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.657664Z digest=sha256:75acf65f6953c077dff91d58e83323e5c44ac0d2da9255165569e98fc8e58f49

Observation 9820e81e-1a97-4abd-9d41-c424982aea51 · inbound

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation cites this paper.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.722321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.722321Z digest=sha256:588320f5dcaeebcb924778461b1380e5979ab6101a0034871d20292a8215455d

Observation 6a40cea9-ec0b-4f07-8e47-c32fc995deac · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:42:13.982380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:cf9c68142f1391f68ae5b40aca45949503cb29c136df2bf66e44933cddf0e85e

Observation fdeb4338-48e7-4788-a7ae-2b939a1f1ff5 · inbound

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem cites this paper.

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:42:12.843952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T08:40:56.186349Z digest=sha256:6189477d39aa6452d0ceed20dccb0c9c23c36ef75ee6d97c0a277bfa756ff524

Observation fd629e92-463b-4d93-bd3f-e286209cb7df · inbound

VERA: Variational Inference Framework for Jailbreaking Large Language Models cites this paper.

VERA: Variational Inference Framework for Jailbreaking Large Language Models Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:07:03.747430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:07:03.747430Z digest=sha256:d1d34e854099049d03fecd7c8bac38611ec2d3e360bf32c46fd4889b8f93c8bf

Observation 8e39d5a0-df53-4a41-813c-fd6d04b2881a · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.251121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.251121Z digest=sha256:05457a8f3047abe7b31b4d98c9173629f57272628b3847a82fd9db5a27ec763a

Observation a3f44fa2-acd8-4f49-8551-f7f32084e4ed · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:42.141679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:42.141679Z digest=sha256:3ff6120cc6fe3173844d31aa83abd6e0faca743363830fbba7b20afbf162d400

Observation fa71d631-b006-417e-b74d-85248c6a3edb · inbound

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment cites this paper.

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T12:27:28.685472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:27:28.685472Z digest=sha256:73c87b7278abd79dc8d0527d6ce944a4d97e3f93003cf972741c966e526a3e51

Observation 8ab876c1-6575-4d55-b711-3fb3217f109b · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:45.557792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:45.557792Z digest=sha256:815a364b3af16a1faff185000e35889604918becc6367c609c1b138f1fde9c35

Observation 73840eff-dd60-4c1d-a220-9fdc976e666e · inbound

Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs cites this paper.

Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:00:30.342867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-21T18:58:53.183734Z digest=sha256:2eceeda636db19c943e5935000a871bc857de2a4b757df70c9218e58f3af745d

Observation 68043072-bda3-480b-a46b-c56692804f02 · inbound

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs cites this paper.

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:23.277446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T05:22:59.050348Z digest=sha256:d2833d6dfa555b61c9e2f728decd5cb78b7112f1099b869f2b3a8f2267c590e4

Observation 036a0ae2-1003-4bc5-a3b7-752b02ea33b5 · inbound

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems cites this paper.

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:04.331000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T16:07:31.602378Z digest=sha256:4d53a88e3a2aaff82c443cc970b7f7b37d87a3e9649da13f184c1bc6c2746d0c

Observation 2bd2d478-cd8b-427f-b55e-49ff8614630b · inbound

From Concept-Aligned Tokens to Vulnerable Features: Mechanistic Localization of Jailbreaks cites this paper.

From Concept-Aligned Tokens to Vulnerable Features: Mechanistic Localization of Jailbreaks Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:11.253480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-08T08:27:28.848271Z digest=sha256:6ba134d2cb93c6f92dd043efaa238791f1d33577ef88dd2c988c4b29041c0886