Pith. sign in

Paper Citation Record · LEDGER

Open Sesame! Universal Black Box Jailbreaking of Large Language Models

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2309.01446.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.01446 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T04:22:00.465363Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-21T19:00:30.340656Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a9e0449a-d2cf-4d77-a059-af389bc5cba1 · inbound

Low-Resource Languages Jailbreak GPT-4 cites this paper.

Low-Resource Languages Jailbreak GPT-4 Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T09:24:14.027167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T09:24:13.911401Z digest=sha256:6090c0c61e4238dbbf54396404b31c628ea6b9203cc249acd1253f9c902cffa1

Observation 391b345f-5beb-46e7-9578-5914dc60264a · inbound

TrustLLM: Trustworthiness in Large Language Models cites this paper.

TrustLLM: Trustworthiness in Large Language Models Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 245

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:17:08.691983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-18T11:17:08.108565Z digest=sha256:574dbcc4c19dbb4467454b3e29a593775824c8ae0a4680538e3ddb8c6fe852f2

Observation 6755de70-5077-48ad-bdfc-2a14518faf4b · inbound

A StrongREJECT for Empty Jailbreaks cites this paper.

A StrongREJECT for Empty Jailbreaks Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-16T21:28:02.830434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T21:28:02.745230Z digest=sha256:c4015b6eeca26799bed97c22cbf9c0044d7d815a5a676c0b4d7f6fd8e3e98839

Observation fb3e3ded-af34-448d-8780-349c26f73a04 · inbound

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models cites this paper.

JailbreakBench: An Open Robustness Benchmark for Jailbreaking Large Language Models Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:08:05.637785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-15T06:08:05.386345Z digest=sha256:c96f5c309dffcf7d50958aef2f27a2ceb7324926eeed0680131adb310c99986f

Observation 45ced3d8-724b-4da4-a873-530e46cf695c · inbound

Jailbreak Attacks and Defenses Against Large Language Models: A Survey cites this paper.

Jailbreak Attacks and Defenses Against Large Language Models: A Survey Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:20:44.705732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-15T02:20:44.368219Z digest=sha256:de52e0714771c79962376f3ff1aba098289b9ce807852587310c9cd217b29806

Observation d9172f6b-7dbe-4588-9038-23b5edfb9671 · inbound

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs cites this paper.

KDA: A Knowledge-Distilled Attacker for Generating Diverse Prompts to Jailbreak LLMs Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-09T04:22:00.465363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T04:22:00.465363Z digest=sha256:7d8eda04f71043123dbdb7ce5071d8b704e97404837eba2df01802b125fd2809

Observation 5d779aee-93d4-485b-b0e5-2ed58e4f5896 · inbound

Towards an AI co-scientist cites this paper.

Towards an AI co-scientist Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T13:02:44.303435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T13:02:43.571234Z digest=sha256:c2fe57efebd5e6bcc9dd307cf823a77d2e3df73208da7994d2c74ba8292e4948

Observation 5df0e263-d0cd-4aa3-ae20-573e45d7e151 · inbound

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval cites this paper.

Scalable Defense against In-the-wild Jailbreaking Attacks with Safety Context Retrieval Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T15:15:48.657664Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:15:48.657664Z digest=sha256:75acf65f6953c077dff91d58e83323e5c44ac0d2da9255165569e98fc8e58f49

Observation 9820e81e-1a97-4abd-9d41-c424982aea51 · inbound

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation cites this paper.

Exploring the Vulnerability of the Content Moderation Guardrail in Large Language Models via Intent Manipulation Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T14:32:29.722321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T14:32:29.722321Z digest=sha256:588320f5dcaeebcb924778461b1380e5979ab6101a0034871d20292a8215455d

Observation 6a40cea9-ec0b-4f07-8e47-c32fc995deac · inbound

Exploring the Secondary Risks of Large Language Models cites this paper.

Exploring the Secondary Risks of Large Language Models Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-19T09:42:13.982380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T09:40:58.067398Z digest=sha256:260f0ce6b5c44fd9074c30726b05e51342bf3ea773beac369e69372be0941908

Observation fdeb4338-48e7-4788-a7ae-2b939a1f1ff5 · inbound

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem cites this paper.

Toward Principled LLM Safety Testing: Solving the Jailbreak Oracle Problem Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-19T08:42:12.843952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T08:40:56.186349Z digest=sha256:bddbb7472019f2eb2d793a63f931c552aaef94e4f3edc3b08481c0c37348913a

Observation fd629e92-463b-4d93-bd3f-e286209cb7df · inbound

VERA: Variational Inference Framework for Jailbreaking Large Language Models cites this paper.

VERA: Variational Inference Framework for Jailbreaking Large Language Models Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T22:07:03.747430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:07:03.747430Z digest=sha256:d1d34e854099049d03fecd7c8bac38611ec2d3e360bf32c46fd4889b8f93c8bf

Observation 8e39d5a0-df53-4a41-813c-fd6d04b2881a · inbound

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training cites this paper.

PRM-Free Security Alignment of Large Models via Red Teaming and Adversarial Training Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T17:35:48.251121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:35:48.251121Z digest=sha256:05457a8f3047abe7b31b4d98c9173629f57272628b3847a82fd9db5a27ec763a

Observation a3f44fa2-acd8-4f49-8551-f7f32084e4ed · inbound

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation cites this paper.

Layer-Wise Perturbations via Sparse Autoencoders for Adversarial Text Generation Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-05T20:31:42.141679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T20:31:42.141679Z digest=sha256:3ff6120cc6fe3173844d31aa83abd6e0faca743363830fbba7b20afbf162d400

Observation fa71d631-b006-417e-b74d-85248c6a3edb · inbound

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment cites this paper.

Certifiable Safe RLHF: Semantic Grounding and Fixed Penalty Constraint Optimization for Safer LLM Alignment Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-04T12:27:28.685472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T12:27:28.685472Z digest=sha256:73c87b7278abd79dc8d0527d6ce944a4d97e3f93003cf972741c966e526a3e51

Observation 8ab876c1-6575-4d55-b711-3fb3217f109b · inbound

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses cites this paper.

SoK: Systematizing LLM Prompt Security: Taxonomies, Datasets, and Unified Evaluation of Attacks and Defenses Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-04T09:25:45.557792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T09:25:45.557792Z digest=sha256:815a364b3af16a1faff185000e35889604918becc6367c609c1b138f1fde9c35

Observation 73840eff-dd60-4c1d-a220-9fdc976e666e · inbound

Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs cites this paper.

Evolve the Method, Not the Prompts: Evolutionary Synthesis of Jailbreak Attacks on LLMs Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-21T19:00:30.342867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T18:58:53.183734Z digest=sha256:a21fec0e794a558eee3c163c9e5ddc037e69771e1af501ccb683608b406ec490

Observation 68043072-bda3-480b-a46b-c56692804f02 · inbound

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs cites this paper.

Stop Tracking Me! Proactive Defense Against Attribute Inference Attack in LLMs Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-16T05:27:23.277446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T05:22:59.050348Z digest=sha256:1b763a558e362809cee4240382f15afaafefae5d76ff995e7cdd8ee6595933e0

Observation 036a0ae2-1003-4bc5-a3b7-752b02ea33b5 · inbound

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems cites this paper.

The Salami Slicing Threat: Exploiting Cumulative Risks in LLM Systems Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:16:04.331000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:07:31.602378Z digest=sha256:445cef2c05bdc5baf1a231766dd6a9009689db60b42b58b882ff183a7a25d118

Observation 2bd2d478-cd8b-427f-b55e-49ff8614630b · inbound

From Concept-Aligned Tokens to Vulnerable Features: Mechanistic Localization of Jailbreaks cites this paper.

From Concept-Aligned Tokens to Vulnerable Features: Mechanistic Localization of Jailbreaks Open Sesame! Universal Black Box Jailbreaking of Large Language Models

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-11T20:36:11.253480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-08T08:27:28.848271Z digest=sha256:f147ce722f36fe01a61bd656ef4cc8dd3ec5768c60b2245301d76dbc0554ce11